As soon as a task goes to the agent, a model choice arises, and the menu in Devin Desktop is long. The temptation is to choose by brand or by the principle "take the strongest, you won't go wrong". Both feel safe: a strong model will surely handle anything. But choosing a model is not choosing "the best overall" - it is matching to the uncertainty of a specific task.
The naive strategy is simple: pin one flagship for everything. It saves the decision now - no need to think before each task - and is therefore sticky. The price comes later and unnoticed: a strong model on a trivial search or a one-line diff spends credits and time where a fast one would have done, and over a stream of such tasks the difference adds up into real cost.
"Always the strongest" breaks on two sides. From above - overpayment and extra latency on simple tasks. From below - if out of habit you pick a fast model for a genuinely complex task with ambiguous architecture, it saves pennies and fails the essence, and the redo costs more than any saving. No single point on the scale is optimal for the whole range of tasks.
The professional move is to shift the choice onto uncertainty, not brand. A fast option - for low-uncertainty tasks: code search, retrieval, a simple predictable diff. A strong one - where uncertainty is high: ambiguous architecture, complex debugging, a plan touching several systems at once. The question is not "which model is better" but "how risky and ambiguous is this task".
It is exactly this choice that Devin offers not to make by hand by default. Adaptive is the recommended router: it analyzes the task and itself routes the simple to fast and efficient models and the complex to more capable ones, balancing quality and cost. The documentation calls it outright the best default for most users. The point of delegation here is the same as elsewhere in the product: a routine decision is handed to a mechanism, and human attention is saved for the cases where it is actually needed. A manual choice makes sense when you have a reason not to trust the routing on a specific class of tasks.
The catalog beyond Adaptive should be read as a snapshot, not a list for the future. At the slice it held specialized models: variants for fast and for careful work, separate ones for Tab, for retrieval, for Quick Review, and some variants marked as available only in Devin Local. The specific names change from release to release, and there is little point in memorizing them.
Why the book is especially cautious with names and numbers here. The catalog, availability and prices change faster than any text is printed. At the date of this snapshot the documentation still carried an already-expired promotion on one of the models "through August 8" - a live example of how a page lags behind fact. So the source of truth here is the actual model picker and the usage page, not a remembered line.
The cost of inattention is two-sided and therefore treacherous. Overpayment on the simple is unnoticeable one by one and painful over a stream. A shortfall of power on the complex is unnoticeable at launch and painful in the result - when a cheap model confidently produced a plausibly wrong plan, and it still has to be recognized and rolled back. A plausibly wrong plan is more dangerous than an obviously bad one precisely because it passes the first glance and wastes time already in implementation, not at the start. Both errors are cheaper to prevent by choosing on risk than to treat afterward.
A manual choice is justified in two situations. The first - when you know the class of task better than the router: for example, a knowingly heavy architectural analysis where it is worth taking a strong model at once. The second - when you honestly compare models against each other. But a comparison makes sense only on equal criteria: the same task, the same context, the same way of verifying the result.
A model choice must be checked by the same measure as any agent work - proof of result, not the feeling that "this one is smarter". Compare the output on the same input and by the same criterion: did the tests pass, is the task solved, how much did it cost. Different tasks fed to different models say nothing about the models - only about the tasks.
The typical failure is choosing by brand and comparing on unequal terms: one model got a simple case, another a complex one, and a conclusion is drawn about the models. The sign is simple: a model decision is explained by its name rather than by the task's risk and a measurement on equal input. Ask first how uncertain the task is - and in most cases the answer will be "leave Adaptive", while a manual choice is saved for the cases where you have a measured reason.