As at 30 September 2026 · next review: Monday, 5 October 2026. This post is a ‘living post’: models, benchmarks, prices and limits are re-verified every week. Any changes are listed, with dates, in the changelog at the end.
The benchmark table you’re currently reading somewhere in the newsletter is already out of date. Over the last four weeks, the top spots have been reshuffled five times. And yet the question that’s really burning in the developer forums is a different one. A ChatGPT Pro user writes: “Weekly limit reset yesterday, still 20 per cent left today.” The model question has become a side issue. Your productivity is determined by the cost per task and your weekly limit; that last percentage point on the leaderboard is just noise.
This post therefore clarifies four things anew every week: which models are leading the pack, which benchmarks can actually still measure this, what the 20, 100 and 200 euros, and what developers experience with Claude Max, Codex and Kimi in their day-to-day work.
AImodels compared: who’s currently in the lead
The real news this summer comes from the open-source camp. Since July, four Chinese laboratories have released models with between one and three trillion parameters and one million tokens context under an open licence: Kimi K3 from Moonshot (2.8 trillion parameters, 104 billion active, weights available since 27 July), Alibaba’s Qwen3.8-Max, DeepSeek V4-Pro under the MIT licence and, most recently, Xiaomi’s MiMo-V2.6-Pro, with an index score of 46 the most powerful open-source model in the Artificial Analysis ranking.
How big is the gap? The US institute CAISI estimates it for DeepSeek V4 Pro at around eight months behind the US frontrunners; for GLM-5.3, it measures on cyber benchmarks around four months. Nathan Lambert from Interconnects sees the gap narrowed by K3 from six to nine months to three to five months. Three to eight months: That’s the head start you gain today if you pay for a closed-source top-of-the-range model.
And the leaders? A duopoly. In June, Anthropic moved the ‘Fable/Mythos 5’ class above Opus; Fable 5.1 and Mythos 5.1 were launched on 1 September; Mythos remains reserved for programmes with security clearance. These were followed by Claude Opus 5.5 at 4/20 US dollars per million tokens, which, according to Anthropic, is “on a par with Fable 5.1 for most tasks”, and Sonnet 5.5 at the unchanged price of 2/10.
OpenAI has simultaneously replaced its entire range. GPT-6 Astra is the company’s largest training run (over 100,000 GPUs), with 1.05 million tokens of context and a price of 10/50 US dollars per million. Among these are GPT-6 Sol and Luna, at half the price of the 5.6 generation, and on 29 September, GPT-6.1 Sol, which, according to OpenAI, delivers Astra-level performance on DeepSWE at a fifth of the cost.
Google has spent the year without a new Pro model. Gemini 3.5 Pro was announced at I/O and never launched; instead, three Flash models were released in six weeks, most recently Gemini 3.8 Flash on 2 September. According to Google, Gemini 4 is set to be released well before the end of the year. xAI releases new models monthly (Grok 4.7 on 21 September, US$2/6). Meta has effectively discontinued the Llama line: Since April, Muse Spark has been the only proprietary model available, currently in version 1.3.
| Model | Provider | Release | Context | API price (1M tokens in/out) | Open weights |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 1 September 2026 | 1M | 10 $ / 50 $ | no |
| Claude Opus 5.5 | Anthropic | 22 September 2026 | 1M | 4 $ / 20 $ | no |
| Claude Sonnet 5.5 | Anthropic | 28 September 2026 | 1M | 2 $ / 10 $ | no |
| GPT-6 Astra | OpenAI | 4 September 2026 | 1.05M | 10 $ / 50 $ | no |
| GPT-6.1 Sol | OpenAI | 29 September 2026 | 1.05M | 2 $ / 10 $ | no |
| GPT-6 Luna | OpenAI | 22 September 2026 | 1.05M | $0.10 / $0.50 | no |
| Gemini 3.8 Flash | 2 September 2026 | N/A | $0.75 / $3.75 (introductory price until 31 December) | no | |
| Grok 4.7 | xAI | 21 September 2026 | 500K | $2 / $6 | No |
| Kimi K3 | Moonshot | 16 July 2026 (weigh-in 27 July) | 1M | 3 $ / 15 $ | yes (revenue-based licence) |
| DeepSeek V4-Pro (0813) | DeepSeek | 13 August 2026 | 1M | N/A | yes (MIT) |
| Qwen3.8-Max | Alibaba | 3 August 2026 | 1M | N/A | yes (own licence; Apache 2.0 only for Qwen3.8-27B) |
| MiMo-V2.6-Pro | Xiaomi | 21 September 2026 | 1M | $0.43 / $0.87 | Yes (MIT) |
Vendors’ list prices, as at 30 September 2026; batch and cachediscounts are on top of that. And that is precisely where the most important figure of the month lies: Anthropic has reduced the cache read prices by 75 per cent for Fable 5.1 (to $0.25 per million) and by 60 per cent (to 0.20). In an agent run that re-reads the same 200,000-token context fifty times, that amounts to ten million cache tokens: with Opus 5.5, the cost is two US dollars instead of five. The list price for fresh tokens is almost irrelevant for such workloads.
Benchmarks: what still has room for growth and what is saturated
Anyone who selects models based on tables has a problem: the benchmarks you know are obsolete. OpenAI removed SWE-bench Verified from its own reporting in February because 59 per cent of the tested tasks were flawed and top models were able to reproduce the human reference solution verbatim. GPQA Diamond was removed from the index by Artificial Analysis in September because all top models ‘simply solve’ it. AIME 2026 ranges from 97 to 99 per cent for the top 3. At FrontierMath, Epoch AI had to re-release in June after 42 per cent of the tasks contained errors.
What counts instead are interactive and economically sound tests where contamination is difficult. Terminal-Bench 4.0 comprises 66 curated tasks; eight were removed (saturated, publicly solved or with quality issues), 19 were revised. ARC-AGI-3 is interactive; the model must learn the rules of an environment through trial and error, so there is no solution that could be found in the training dataset. GDPval-AA measures real-world professional tasks from 44 occupations in blind pair comparisons. And Artificial Analysis weights private, never-before-published test sets at 45 per cent in the Index v4.3, compared with 40 per cent in v4.2 and 20 per cent previously.
| Benchmark | What it measures | Top 3 (as at 30 September) | Status |
|---|---|---|---|
| AA Intelligence Index v4.3 | 10 evaluations, 45% private sets, independent | Opus 5.5 (58), Opus 5.5 xhigh / Sonnet 5.5 (56), Fable 5.1 / GPT-6 Astra (53) | current |
| Terminal Bench 4.0 | agent-based terminal tasks, independent | Sonnet 5.5 (63.6%), Opus 5.5 max (59.6 per cent), Opus 5.5 xhigh (59.6 per cent) | current |
| Arena WebDev | Web apps in a blind comparison, 795,000 votes | Opus 5.5 (1,827), GPT-6 Astra (1,792), Fable 5.1 (1,751) | current |
| GDPval-AA v2.1 | 220 professional tasks, 44 professions, Elo, independent | Opus 5.5 (1846), Sonnet 5.5 (1844), Opus 5.5 xhigh (1820) | current |
| Humanity’s Last Exam | Expert knowledge, no tools, independent (Scale) | GPT-6 Astra (54.2%), Gemini 3.1 Pro (47.3%), Fable 5.1 (46.8%) | current |
| ARC-AGI-3 | Interactive reasoning, independent (ARC Prize) | GPT-6 Astra 62.7% on the standard harness (99.9% on the provider’s harness), Opus 5 30.2%; Opus 5.5 not yet measured | current, harness dispute |
| SWE-bench Verified | Coding tickets | Top score at 96 per cent; positions 11 to 15 within half a point at 80 per cent | Saturated; discontinued by OpenAI |
| GPQA Diamond | Natural Sciences | Provider figures: Astra 96.0 per cent, GPT-5.6 Sol 94.6 per cent | saturated, removed from the AA Index |
| AIME 2026 | Maths Olympiad | Top 3 within 2.1 points at 97 to 99 per cent | saturated |
‘xhigh’/“max” denote the reasoning effort used in the measurement. “Harness” refers to the test environment: which tools the model is given, how many attempts, and what computational budget.
The figures from providers’ blogs are almost universally higher than the independent measurements. Anthropic reports 70.6 per cent for Sonnet 5.5 on Terminal Bench 4.0, whilst Artificial Analysis measures 63.6 per cent. OpenAI reports 99.9 per cent for Astra on ARC-AGI-3, whilst the standard ARC Prize harness yields 62.7 per cent, with computing costs of around 26,000 US dollars for the test run. Neither figure is a lie; they simply represent different harnesses and levels of reasoning. It’s just not comparable. In short: trust the figure measured by someone who isn’t selling the model.
What’s been happening in recent weeks: Since the release of Opus and Sonnet 5.5, Anthropic has been leading the way in agentic coding tests and professional task indices, whilst OpenAI leads in knowledge, maths and ARC. Both providers are now promoting the same metric. Anthropic calculates that Opus 5.5 outperforms Astra at the maximum level on the intermediate thinking level, at around a fifth of the cost per task. OpenAI countered on 29 September with GPT-6.1 Sol and exactly the same formula.
AI subscriptions compared: what you get for 20, 100 and 200 euros
Comparing AI subscriptions is easier than it looks. All providers have settled on three tiers (around 20, 100 and 200 US dollars), and all control the price via limits. Prices in euros include VAT where the provider specifies this:
| Provider | Entry level | Mid-range | Top tier | What distinguishes the tiers |
|---|---|---|---|---|
| Claude | Pro €21.42 (Opus, Sonnet, Claude Code) | Max 5× €107.10 | Max 20× €214.20 | 5-hour window plus weekly limit; Fable 5.1 available only on Max and Team Premium, up to 50 per cent of the weekly limit |
| ChatGPT | Plus €23 (15 to 150 Sol messages per 5 hours) | Pro 5× €103 | Pro 20× €229 (new sign-ups since 29 September with halved quota) | Weekly quotas at all tiers; Astra Ultrafast only on Pro for $500 |
| Gemini | AI Pro $19.99 (3.1 Pro, Antigravity entry-level) | AI Ultra $99.99 | AI Ultra $199.99 | Gemini CLI disabled for subscriptions since 18 June; replaced by the closed Antigravity CLI |
| Kimi Code | Andante ¥49 (K2.7 Code) | Moderato ¥99 / Allegretto ¥199 (K3, 1M) | Allegro ¥699 | Monthly pool plus weekly allowance plus 5-hour rate; processed by Moonshot AI Pte. Ltd. in Singapore |
| Cursor | Pro 20 $ (20 $ credit) | Pro+ $60 (credit balance of $70) | Ultra $200 (credit balance of $400) | From 2025, credits will replace requests; Auto mode will also use up credits |
| GitHub Copilot | Pro $10 ( $10 credit) | Pro+ $39 | Business $19 / Enterprise $39 per user | From 1 June 2026: usage-based pricing according to API rates |
Euro prices for Claude according to ssdnodes, for ChatGPT according to heise; Google and Kimi do not quote euro prices for Germany. For teams, the number of user seats is relevant: Claude Team Premium with Claude Code costs 100 US dollars per seat on an annual subscription, ChatGPT Business Premium also 100; with a VAT number, Claude Max 20× comes to 180 euros net.
The difference between the tiers is not a model upgrade. With Claude, even the Pro tier at 21 euros gets Opus 5.5 and Claude Code. With OpenAI, the Plus tier gets GPT-6 Astra, but only between five and 45 Astra messages every five hours. You pay by volume. The more interesting question is therefore: How quickly will I reach the limit?
In practice: what developers really experience with Claude Max, Codex and Kimi
The subscription page states “5× Pro”. In recent weeks, both major providers have revised downwards what this actually amounts to.
On 14 September, Anthropic replaced the “+50 per cent weekly limit” promotion, launched in May, with a permanent ‘+25 per cent compared to the old base value’. This sounds like an improvement, but compared to the situation over the last four months a 17 per cent reduction, a figure that Anthropic also quantified. The 5-hour windows, which were doubled in May, remain in place.
At the same time, a class action lawsuit has been running in the US since June, which alleges that Max 20× actually delivers six to eight times the data allowance of Pro, not twenty times. A Max 20× user who has been logging his usage for months reports 20 per cent weekly usage on the first day after the reset. His conclusion: For $200, yes, “but I wouldn’t bet on it in a year’s time”.
An effect that hardly anyone takes into account: Fable 5.1 is capped at 50 per cent of the weekly limit and costs 2.5 times as much per token as Opus 5.5. Anyone setting Opus 5.5 as their default has, according to a circulating estimate by Theo Browne, will have around four times the effective quota.
OpenAI has solved the problem from the other side. On 10 September, new registrations for Pro 20× were paused because, according to Codex head Tibo Sottiaux, this segment places the greatest load on the systems; on 29 September, sign-ups reopened with halved quota; existing customers are being spared until 29 October. The quote from the hook comes from this phase: “Weekly limit reset yesterday, 20 per cent left today”, writes a Pro-5× user on the OpenAI forum.
The common practice is to divide the work: Codex for the bulk of small changes and cloud parallelisation, Claude Code for long, continuous sessions. In a analysis of 500 Reddit comments, 65 per cent preferred Codex, mainly because of the limits, whilst Claude Code won eight out of twelve blind tests, which is not a contradiction.
Kimi is the third contender, and she’s one to be taken seriously. Moonshots’ Kimi K2.7 Code runs as a drop-in replacement for Claude Code (ANTHROPIC_BASE_URL to https://api.moonshot.ai/anthropic – that’s it) and, at 0.95/4 US dollars per million tokens, costs around a quarter to a fifth of Opus 5.5.
A sample calculation compared to Opus 4.8 at the time came to 0.69 US dollars instead of 4.00 US dollars for a 50-turn agent run. A practitioner who ran K2.7 in Claude Code for around two weeks, describes it as disciplined in following instructions, with no scope creep, but lacking in extended thinking and with less community knowledge regarding the best prompts.
The catch lies elsewhere. The international Kimi platform is operated by Moonshot AI Pte. Ltd. in Singapore, where the data is processed; a data processing agreement compliant with GDPR standards is not publicly available. Singapore is a third country without an adequacy decision from the EU, and the parent company is based in Beijing. For a German SME with customer data in the repository, this is a deal-breaker as long as there is no EU entity with a data processing agreement.
For code without personal data, it is a cost-effective secondary option. Ramp reports that 6.1 per cent of US companies paying for AI are now using platforms that provide open-source or Chinese models.
Two figures to put into context just how far removed this practice is from the forums. According to Menlo Ventures, around 54 per cent of corporate spending on AI coding in 2025 was accounted for by Anthropic, 21 per cent to OpenAI; in the Ramp AI Index from August, 43.5 per cent of US companies pay Anthropic and around 40 per cent pay OpenAI.
User figures tell a different story to the budgets. The Stack Overflow survey 2026, which received over 49,000 responses, reveals that 84 per cent use AI tools, 29 per cent trust their accuracy – 3 per cent ‘strongly’ – and 66 per cent say the code is “almost right, but not quite”. Among the tools, Copilot leads the way with 68 per cent of AI users, Cursor stands at 18 per cent, and Claude Code at just under 10 per cent.
What this means for your setup
I think it’s a mistake to rely on a single model for a team. The reason is economic: the cost per task depends more on the reasoning level than on the model. Opus 5.5 delivers an identical 59.6 per cent on Terminal Bench 4.0 at ‘max’ and ‘xhigh’, the higher level simply costs more tokens. My recommendation, as things stand today:
The default is a mid-range model at a medium reasoning level. Today, that means Opus 5.5 medium or GPT-6.1 Sol; both deliver top results at a fifth of the cost. Frontier models – that is, the most expensive top-tier offering from each provider (Fable 5.1, Astra) – are reserved for scenarios where an error proves costly: security reviews, architectural decisions, and the final check before the merge.
For high-volume tasks with clear instructions (generating tests, writing migration scripts, producing documentation), a third, low-cost option is worthwhile, provided that data protection and data processing on behalf of a client have been clarified. This could be Luna, Gemini Flash or an open model on European infrastructure.
As I described in the post “Agent = Harness + Model”, the harness – comprising context, memory, tools and evaluations – determines whether an agent is productive. In a properly built setup, changing the model is a matter of a single configuration line; ANTHROPIC_BASE_URL is an example of this above.
If switching from Opus 5 to 5.5 takes days, the dependency is in the wrong place. And anyone sending customer data through a model outside the EU because it’s five times cheaper should read our article on the European AI stack first.
When deciding on a subscription: a team of five developers is better off with Claude Team Premium (US$100 per seat, including Claude Code) or ChatGPT Business Premium than with five individual Max subscriptions. SSO and audit logs are one reason.
The other: on the Team and Business plans, chats are not used for training by default and belong to the company account; with private subscriptions, they are linked to the employee and move with them. Anyone keeping Cursor or Copilot should be aware that both now charge according to API rates and that the $20 subscription is a credit balance.
What I’m keeping an eye on until the next review
Four open threads that will shape this post over the coming weeks.
Anthropic announced Haiku 5.5 on 22 September, saying it would be released ‘in the coming weeks’, but it isn’t here yet. At OpenAI, the grace period for existing Pro-20× customers ends on 29 October; after that, the halved quota will apply to everyone. Gemini 4 is expected to launch well before the end of the year and would be Google’s first new Pro model since February. And the class action lawsuit against Anthropic’s Max limits is pending; a judgement or settlement would directly affect the subscription table above.
Frequently Asked Questions
Which AI model is currently the best? For agentic coding and professional tasks, Claude Opus 5.5 leads the independent indices (Artificial Analysis, Terminal-Bench 4.0, GDPval-AA); for knowledge and interactive reasoning, GPT-6 Astra (Humanity’s Last Exam, ARC-AGI-3). The performance gap is smaller than the price difference.
Is Claude Max 20× still worth it? If you work in Claude Code for several hours a day and set Opus 5.5 as your default instead of Fable, then yes. Expect to pay 214 euros gross and bear in mind that a single large agent run will make a noticeable dent in your weekly limit. For teams, Team Premium at 100 US dollars per seat is usually the better option.
Can I use Kimi K3 for business purposes in Germany? Technically yes, via API or as a drop-in in Claude Code. Legally, it depends on what data you send: Processing is carried out by Moonshot AI Pte. Ltd. in Singapore, a third country without an EU adequacy decision, and a data processing agreement is not publicly available. For code without personal data, it’s a cost-effective alternative; for customer data, it isn’t.
Why do provider blogs show different benchmark figures to those here? Because providers measure using their own test suite, at the highest difficulty level and, in some cases, with multiple attempts. Independent bodies such as Artificial Analysis, ARC Prize or Scale measure using a fixed test suite and budget. This post gives preference to the independent figures and marks vendor figures as such.
Changelog
From 5 October, every change to models, prices, limits and top-3 rankings will be listed here, with the date. Anything not listed here has not changed.
- 30 September 2026 — First published. Status: Opus 5.5, Sonnet 5.5, GPT-6 Astra, GPT-6.1 Sol, Gemini 3.8 Flash, Grok 4.7, Kimi K3, MiMo-V2.6-Pro. Claude weekly limit −17% since 14 September; ChatGPT Pro 20× has been open again (halved) since 29 September.