Yandex has added an increasingly important capability to Alice AI: the assistant now decides for itself how much computational effort a request deserves. Instead of requiring users to choose between a fast model and a more capable reasoning mode, a proprietary routing system analyzes the request, selects the appropriate model and can automatically activate Expert mode when the task demands deeper analysis.
The change was announced by Yandex on September 16. For everyday requests, Alice AI prioritizes models that can respond quickly. When a prompt requires comparing multiple constraints, analyzing several documents, performing calculations or constructing a multi-step plan, the system can escalate the task automatically to Expert mode. Users can still select Expert mode manually, but Yandex says automatic routing is now the default behavior.
The larger significance is that model selection is only the first decision. Yandex says the assistant can also determine that a specialized request needs fresh information, uploaded files, connected services or other tools. Its agent architecture — which Yandex calls a “harness” — coordinates those resources, maintains context and plans the task step by step. In other words, the AI is increasingly deciding not just what to answer, but what process is required to produce the answer.
The query now determines the amount of intelligence allocated to it
Most AI interfaces historically exposed model choice to the user. Someone wanting a faster answer selected a lightweight model; someone needing deeper reasoning switched to a more capable option. Yandex is moving that decision into the system itself.
Its automated router evaluates each request to estimate the resources required and which model is best suited to handle it. Simple tasks can stay on a faster path, preserving latency and computational efficiency. Complex requests can receive additional reasoning resources without the user needing to understand the differences between the underlying models.
This is a meaningful product shift because model names are implementation details for most people. Users generally care about whether the answer arrives quickly and whether the system recognizes when a problem is too complicated for the cheapest path. Automatic routing turns that judgment into part of the AI product rather than part of prompt engineering.
It also mirrors a broader trend across AI systems: intelligence is becoming dynamically allocated. Instead of every query consuming the same model and inference budget, orchestration layers can classify the task and spend more compute only where additional reasoning is likely to improve the result.
Expert mode activates automatically for complex tasks
Yandex gives several examples of requests that can trigger Expert mode: weighing multiple conditions, reviewing several documents simultaneously, running calculations and producing detailed step-by-step plans. The company says the mode is particularly useful in areas such as shopping, finance, law, healthcare and travel.
One example is trip planning under a fixed budget. A user can ask Alice AI to plan a three-day trip for two people, find less expensive dates and calculate the total cost of transportation and accommodation. That is not a single retrieval problem. It requires breaking the request into subproblems, gathering current information, applying constraints and synthesizing the pieces into a coherent plan.
Yandex says users in Expert mode can follow how the system works through the problem and which information it draws on. Expert mode is currently available to all Alice AI users.
That does not mean every difficult query necessarily follows an identical reasoning trace or uses the same set of tools. The important architectural point is that complexity itself can now trigger a different execution path.
Web search becomes a tool the AI can choose
The routing announcement is particularly relevant to search because Yandex explicitly says models are only one component of the system. To answer a specialized request, Alice AI may search for information, process files, connect to services and invoke other tools.
This turns web search from the entire product into one capability inside a broader reasoning workflow. The AI can potentially decide that some tasks can be answered directly, while others require fresh external information before generation begins.
Yandex's existing Search with Yandex AI documentation already describes a related workflow for complex web questions: the system breaks a question into parts, finds relevant sources, analyzes them and synthesizes the information into a single answer with links. The new Alice AI routing architecture extends the same general principle beyond retrieval by allowing the assistant to combine models, files, services and tools according to the task.
For publishers and SEOs, that distinction matters. The opportunity for a web page to influence an AI answer increasingly depends on whether the system decides external retrieval is necessary in the first place. Retrieval is becoming conditional rather than assumed.
The harness coordinates models, context and actions
Yandex calls its orchestration layer a harness. According to the company, this technology helps the model plan actions, preserve context and execute tasks step by step. That makes it useful to think of the model as one component inside a larger runtime rather than as the entire assistant.
A complex request can therefore involve several decisions before the final answer appears. The router can estimate task difficulty and choose a model. The harness can determine which additional resources are required. Search can supply current information, file-processing systems can analyze uploaded documents, and connected services or tools can provide specialized capabilities. The model then has to integrate those outputs into a useful response.
This layered architecture is important for understanding modern AI search. A generative answer is not necessarily the direct output of one language model given one prompt. It can be the final result of an orchestrated process in which retrieval, reasoning and actions happen at different stages.
Yandex is exposing two different layers of its AI search architecture
The September 16 routing announcement comes just two days after Yandex open-sourced Alice AI Search Pretrain, the pretrained base model behind the system used to generate AI answers in Yandex Search. That model uses a sparse Mixture of Experts architecture with 35 billion parameters in total but only around 600 million active for each token.
The two announcements illuminate different layers of the same broader technological shift. The open-source release shows how Yandex can make an individual model computationally efficient. The new routing system shows how the product can decide when a request needs a different model or additional resources altogether.
Efficiency therefore exists at two levels. Inside a model, sparse routing can activate only the parameters needed for a token. Above the model, system-level routing can allocate different models, reasoning budgets and tools according to the request. Production AI products increasingly depend on both forms of optimization.
It is important, however, not to collapse Alice AI and every Yandex Search request into one identical stack. Yandex's September 16 announcement concerns the Alice AI assistant and its automated model selection. Yandex Search has its own documented AI-answer pipeline. The products are closely related and share Yandex's AI technology, but the company does not say that every conventional Search query automatically invokes Expert mode or the complete Alice AI harness.
SEO visibility now depends on orchestration decisions
For search marketers, adaptive routing adds another layer to the visibility problem. Traditional SEO asks whether a crawler can access a page and whether the ranking system considers it relevant. Generative search adds questions about retrieval and synthesis. Agentic systems add another question before those: what tools does the system decide to use for this particular request?
A query that triggers web research can create opportunities for external sources to enter the answer. A request handled entirely from model knowledge may not. A task that invokes a connected service could route the user toward structured data or an API rather than a conventional webpage. A multi-step agent may retrieve information at several points as its plan changes.
This means AI visibility cannot be understood solely by measuring citations after the fact. The system's orchestration policy influences whether the open web is consulted, which type of source is useful and at what stage of the task that information enters the process.
Yandex's Webmaster documentation says pages used in Yandex AI search answers are drawn from content indexed by Yandex and that well-structured, informative and grammatically correct pages are more likely to be used. Adaptive agents make that foundation even more important: when the harness decides web evidence is necessary, the retrieval layer still needs machine-readable sources it can confidently select and process.
Complex queries can become workflows instead of searches
The travel example illustrates how the unit of interaction is changing. “Find hotels in Tokyo” is recognizably a search query. “Plan a three-day trip for two under this budget, find cheaper dates, compare transportation and accommodation and calculate the total” is closer to a project.
Once the request becomes a project, search is only one step. The system may need to discover options, compare them, calculate costs, preserve constraints across several steps and revise its plan when one condition changes. That is why an agentic harness becomes more important than simply choosing a stronger language model.
For businesses, this could eventually change how visibility is won. A hotel may need to be discoverable through search, but it may also need accurate prices and availability in the service an agent consults. A retailer may need strong product pages, structured catalog data and a transaction path compatible with agentic tools. A publisher may need passages that can be retrieved as evidence during one stage of a longer research process.
The competitive surface therefore expands from ranking pages to being usable inside workflows.
Automatic reasoning is also an economics decision
There is a practical reason not to run the most expensive reasoning process for every request. Deep inference, web research and tool calls consume time and computing resources. A system that automatically distinguishes between “What is the capital of Japan?” and a constrained two-week travel plan can allocate those resources more efficiently.
That makes the router economically important as well as technically interesting. The ideal orchestration system spends enough compute to solve the task reliably without imposing maximum latency and cost on simple requests. The user sees one assistant, but behind the interface the platform can operate a portfolio of models and capabilities.
This architecture also creates a new quality challenge. The router can make mistakes. A seemingly simple request may require fresh information, while an apparently complex prompt may not benefit from expensive tools. The quality of the AI experience therefore depends not only on model intelligence but on whether the orchestration layer correctly recognizes what the task requires.
AI search is becoming a decision system about how to answer
Yandex's latest Alice AI update is easy to describe as automatic Expert mode, but that understates the change. The assistant is gaining responsibility for deciding the computational path between the user's request and the final response.
Sometimes that path can be short: choose a fast model and answer. Sometimes it can expand into deeper reasoning, web research, multi-file analysis, service connections and tool use coordinated by an agentic harness. The user no longer needs to know which architecture is appropriate before asking the question.
For the search industry, this is another sign that AI search should not be modeled as a simple replacement for ten blue links. The emerging system first interprets the task, then decides how much reasoning to allocate, whether fresh retrieval is necessary and which tools or agents should participate. Only after those choices does it produce the answer. In that environment, understanding how an AI decides to search may become almost as important as understanding what it finds once it does.