Citi has reversed its reading of the AI model market. In a note on 1 October, the bank told clients that the capability gap between proprietary frontier models and open-weight alternatives is widening again, after a period in which the two had been converging.
Citi's measure is the Artificial Analysis Intelligence Index, an independent score that combines ten benchmarks. The best proprietary models have moved up to 58 from 53, while the strongest open-weight models have gone to 46 from 44, taking the gap from 9 points to 12. The bank pointed to Anthropic's Claude Opus 5.5, which it said is 9% more capable than Opus 5 while costing 40% less on typical workloads, alongside successive improvements in Google's Gemini and xAI's Grok.
Open-weight models lag furthest behind on long-horizon cyber tasks and on production reliability, according to the note. Both matter more than raw benchmark scores for companies putting models into live systems, where a model has to hold up across long chains of work rather than single questions.
A Month Ago, Citi Said The Opposite
The reversal is sharp. On 1 September, the same bank told clients that open-weight models had overtaken proprietary ones in developer usage, reaching 53% of token volume on Vercel's AI Gateway by late August, a rise of 24 percentage points since late June. It put the capability gap at just 3 points at that stage, down from 9.
That note cited DeepSeek earning about $71 million in revenue through July at an 83% gross margin, evidence that open-weight developers could build real businesses, and Nvidia's $12.9 billion agreement to buy Hugging Face, the main hub for open models, as a sign of where the industry's money was moving.
Citi had made a similar argument in June, when it said demand for open-weight models had risen sharply as companies faced higher costs and tighter restrictions on access to leading proprietary systems. Open-source models' share of tokens on OpenRouter had climbed to 65% in June from 34% in January.
What Changed In Between
Two things explain much of the swing. The first is that the frontier labs shipped. Claude Opus 5.5, Gemini and Grok all improved over the period, and Google released Gemini 4 Argon this week, so the top of the index moved up faster than the open-weight field.
The second is that Citi's earlier note credited part of the convergence to conditions that have since changed, including the distillation of frontier models by open-weight developers and their access to computing power offshore. Both have come under pressure: US agencies named six Chinese AI companies in a September advisory over distillation, and the frontier labs have tightened their defences against it.
Two Different Questions
The apparent contradiction between Citi's notes is less severe than it looks, because they measure different things. Usage share counts which models developers actually send work to, and that is driven mainly by cost. Index scores measure what the best models can do, and that is driven by the leading labs.
Both can rise at once. Open-weight models can take a growing share of routine traffic on price while the frontier pulls ahead on the hardest tasks, and the Artificial Analysis data shows why the cost argument is so strong: in its index methodology, the open-weight leader ran a typical task for a few cents against roughly a dollar or more for the top proprietary models.
For most business workloads, which involve summarising, classifying, drafting and answering questions, a model scoring in the mid-40s at a fraction of the price is the sensible choice. The gap matters where tasks are long, agentic or safety-critical.
Capability Is Not The Only Difference
Open weights bring a trade-off that no index captures. Because anyone can download and modify the parameters, safety training can be stripped out, and Anthropic showed this week that a technique called abliteration cut the refusal rate of Z.ai's GLM-5.3 model from above 90% to around 2 to 3%, while its own models kept their safeguards because the weights are not available.
That cuts both ways for a business. Full control over the weights is exactly what appeals to companies that want to run models on their own hardware, keep data in house, fine-tune on private information or deploy at the edge, as the edge AI benchmarks we examined showed. The same openness means the buyer, not the developer, carries responsibility for the guardrails.
What Citi Expects Companies To Do
The bank's conclusion is practical: enterprises will run mixed portfolios of models as a way to spread risk, rather than committing to one provider or one approach. That means routing cheap, high-volume work to open-weight models and sending the hardest tasks to frontier APIs.
Citi also expects the labs to respond by building routing and orchestration into the models themselves, so that a single API handles the decision about which model should answer a given request. If that happens, the choice businesses make shifts from picking a model to picking a platform that picks models, which would make the frontier labs harder to displace on price alone.
How To Read Bank Research On AI
Citi's reversal is a useful reminder about this kind of analysis. Investment bank notes track quarterly shifts for investors deciding where to put money, and in a market where leading models change every few weeks, a snapshot ages quickly. A company choosing an AI platform works on a longer horizon than a trading desk.
The firmer signals in these notes are the structural ones: cost per task, usage share on real developer platforms, and the kinds of work where each class of model falls short. Those hold up better than a gap measured in index points, which the next model release can change.
The Gap Will Keep Moving Both Ways
A 12-point lead on a benchmark index tells businesses less than it appears to. What it does confirm is that the frontier labs still set the ceiling on capability, and that their advantage shows up most clearly in long, multi-step and security-sensitive work, which is exactly where companies are now trying to deploy AI agents.
For buyers, the useful response is not to pick a side in the open-versus-closed debate but to match each workload to what it actually needs, keep the ability to switch providers, and re-test the assumptions every few months. On Citi's own evidence, the answer changed twice in a quarter.