← Blog

September 2026: 47 new models and a different yardstick

24 set 2026

⏱️ Reading time: ~5 minutes

Introduction

Forty-seven models in thirty days. That's the catalogued count of releases in September 2026 — and the number alone says something about the state of the market: tracking launch by launch has stopped making sense.

But the most interesting data point of the month isn't the count. It's the sales pitch. For the first time in a recent cycle, the most-discussed release didn't arrive promising to be smarter. It arrived promising to be cheaper per completed task. The yardstick moved.

The top of the frontier

GPT-6 Astra (OpenAI, September 3). The heaviest launch of the month. A context window of roughly 1.05 million tokens, 128K output, priced at US$ 10 per million on input and US$ 50 on output. OpenAI positions it as state of the art in computer use, browsing, software engineering, cybersecurity, science, and professional work. A detail that matters to anyone managing risk: it was the company's first model to cross a critical cybersecurity threshold, which is why those capabilities shipped under gated access. On September 22 the family gained two more members, GPT-6 Sol and GPT-6 Luna.

Claude Opus 5.5 (Anthropic, September 22). The move that shifted the math. First model of the 5.5 family, at US$ 4 on input and US$ 20 on output per million tokens — 20% below Opus 5. The company claims, however, that typical workload cost drops around 40%, because the model burns fewer tokens to finish the same task. Cache pricing followed: writes from US$ 6.25 to US$ 5, reads from US$ 0.50 to US$ 0.20. Output got more than 30% faster.

It's worth separating what's measured from what's vendor claim. The list price is fact; the 40% saving is Anthropic's assertion, and it depends on the workload profile. The published example is concrete: porting a legacy HAProxy from C to Rust took 9.5 hours, against 12 hours for Fable 5.1 — a considerably more expensive model per token.

Claude Fable 5.1 and Mythos 5.1 (September 1). Updates to the previous line that held the spotlight for three weeks before being overshadowed by their own vendor.

Grok 4.7 (xAI, September 21). An upgrade over 4.6 at the same price, with gains spread across benchmarks.

Gemini 3.8 Flash and Flash Cyber (September 2) and Gemini 3.8 Live (September 15), from Google.

What, once again, didn't ship

Two absences deserve the record, because they contradict what was announced with fanfare months ago.

Gemini 4 is still in pre-training. There is no model page, no API identifier, no pricing, no benchmark table. Estimates point to November or December, and the September launch rumor never materialized. The flagship remains Gemini 3.8 Flash.

Grok 5 didn't show up either. The stable line is still 4.6/4.7 — adding another quarter to the list of missed deadlines.

The flood on the other side

If the top moved on price, the base moved on volume — and most of it is open.

K2 Horizon arrived on September 3 as an entire family, all at once: 0.9B, 3.7B, 7B, 32B, a 375B A23B, and a MoVA variant at 36B A4B, all with reasoning and open weights. DeepSeek released V4.1 Flash on September 10, open, with reasoning and vision. Alibaba kept an almost weekly cadence with Qwen3.8-Max, Omni-Flash, LiveTranslate, and Qwen-Image 2.1. Xiaomi delivered MiMo V2.6 in Pro and Flash versions. Add MiniMax H3 Max, the Nex-N2.5 family, the sector-specific variants of Ling 3.0 Flash — one for finance, one for healthcare — and the LLaDA line.

The pattern is what stands out: these aren't isolated models, they're complete families, covering sub-1B up to hundreds of billions of parameters, released on the same day.

Conclusion

For anyone running infrastructure, the practical reading of September is direct. Competition at the top stopped being fought over percentage points on a benchmark and started being fought over total cost to complete the work — which involves price per token, number of tokens burned, output speed, and cache efficiency. That's a metric far closer to the spreadsheet than to the leaderboard.

Meanwhile, the sheer volume of open models makes hybrid architectures increasingly defensible: an open model where the task is predictable and cheap, a frontier model where the cost of being wrong is high.

Forty-seven launches in a month isn't a sign of maturity. But the shift in the argument — from "I'm smarter" to "I cost less to deliver" — is. A market that starts competing on price is a market that is ceasing to be a novelty and becoming infrastructure.

Recibe las publicaciones

Nuevos artículos sobre IA, Vibe Code y Builder Code — por correo o Telegram.

o
Recibir en Telegram

Al suscribirte, aceptas recibir correos/mensajes y la Política de Privacidad. Puedes cancelar cuando quieras. Sin spam.