Gemini 3.1 Pro
“Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.” — Google, Sep 30, 2026 · Google
Next AI Model follows every release from OpenAI, Google, Anthropic, xAI, Meta, Mistral, Alibaba, DeepSeek and Microsoft: what each new model changed, in the maker's own numbers, and which lines are due next by their own release rhythm.
Every AI model line we track: first the ones with a model officially announced by the maker, then the ones reaching their usual gap between releases, then mid-cycle, then those already past it, then the ones just released. Each card shows how long it has been since the last release against that line's usual gap and its recent ones. Closed models and open-weight models are listed separately.
“Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program.” — Google, Sep 30, 2026 · Google
“Grok 4.8, which is a 2.5T model trained with our new C++ software stack, will finish training this week and start RL” — xAI, Sep 14, 2026 · Elon Musk on X
24 lines closed by the maker or without a release in over a year live in the Archive.
Every Monday, only when something shipped: the AI models that came out, what each one changed, and which lines are due next.
We email you a link to confirm first · unsubscribe anytime · read past issues · about the newsletter
What shipped recently and what actually changed — compared with the model it replaces (or the rival its maker chose), numbers included.
Costs a tenth of Claude Haiku 4.5 for prompts under 100K tokens while raising its computer-use score from 15.7% to 72.4%, with a 1M-token context window.
Compared with Claude Haiku 4.5 · the model it replaces
Finishes 7 in 10 tasks when operating a computer on its own
OSWorld 2.1 — completing office tasks on a real computer desktop (offline subset)
72.4%
▲ up from 15.7% · +56.7 points
Scores far higher on the all-round intelligence index
Artificial Analysis Intelligence Index v4.3.2 — one score averaging 10 hard tests
43
▲ up from 17 for Claude Haiku 4.5 (reasoning) · +26 points · Haiku 5.5 measured at max effort
Costs developers a tenth of what Claude Haiku 4.5 cost
Price per 1M tokens — what developers pay, input / output
$0.10 / $0.50
▲ down from $1 / $5 · prompts under 100K tokens; $0.50 / $2.50 above that
An update to Nano Banana 2 built on Gemini 3.6 Flash that scores 60 Elo higher in Google's blind text-to-image tests at about half the price per image.
Compared with Nano Banana 2 (Gemini 3.1 Flash Image) · the model it replaces
Preferred more often than Nano Banana 2 in blind image comparisons
Google text-to-image Elo — people pick the better image blind
1050
▲ up from 990 for Nano Banana 2 · +60 Elo, margin about 14 · Google's own test, thinking mode
Costs about half as much per image as Nano Banana 2
Gemini API price per generated image — 1K resolution
$0.0336 / image
▲ down from $0.067 for Nano Banana 2 · 2K image $0.0504
Mistral's first large reasoning model, natively multimodal with 1 trillion parameters (49 billion active), lifts the Intelligence Index from 9 to 38 but costs almost three times as much as Mistral Large 3.
Compared with Mistral Large 3 · the model it replaces
Solves about 6 in 10 hard real-world software engineering tasks
DeepSWE 1.1 — fixing real bugs in large code projects
61.7%
✦ first score published — no Mistral Large 3 result to compare
Scores four times higher on the all-round intelligence index
Artificial Analysis Intelligence Index v4.3.2 — one score averaging 10 hard tests
38
▲ up from 9 for Mistral Large 3 · +29 points · preview with reasoning, Mistral Large 3 has no reasoning mode
Costs developers almost three times as much to run
API price per 1M tokens — input / output
$1.36 / $4.18
preview launch price — was $0.50 / $1.50 for Mistral Large 3
One week after GPT-6 Sol, it scores 7 points higher on computer-use tasks and 4 points higher on the all-round intelligence index at the same $2 / $10 price.
Compared with GPT-6 Sol · the model it replaces
Completes 7 in 10 tasks when operating a computer on its own
OSWorld 2.0 — using real desktop apps to finish tasks
71.4%
▲ up from 64.4% · +7 points · both at max effort
Scores 4 points higher on the all-round intelligence index
Artificial Analysis Intelligence Index v4.3.2 — one score averaging 10 hard tests
52
▲ up from 48 · +4 points · both at max effort
Costs the same as the model it replaces
Price per 1M tokens — what developers pay, input / output
$2 / $10
same as GPT-6 Sol · cached input halved to $0.10
Scores 18 points higher than Sonnet 5 on the all-round intelligence index and runs over 30% faster, at the same $2 / $10 price per 1M tokens.
Compared with Claude Sonnet 5 · the model it replaces
Makes the right code edits more than half the time inside an editor
CursorBench 4.0 — editing code the way developers do in an IDE
55.5%
▲ up from 34.1% · +21.4 points
Scores 18 points higher on the all-round intelligence index
Artificial Analysis Intelligence Index v4.3.2 — one score averaging 10 hard tests
56
▲ max effort (Artificial Analysis: max with fallback) · up from 38 · +18 points
Costs the same as the model it replaces
Price per 1M tokens — what developers pay, input / output
$2 / $10
same as Claude Sonnet 5
Half the price of GPT-5.6 Sol with about half as many factual mistakes, while its all-round score rises only one point.
Compared with GPT-5.6 Sol · the model it replaces
Makes about half as many factual mistakes
OpenAI factuality test — errors in real-world conversations
about half
▲ vs GPT-5.6 Sol · OpenAI internal test
Scores 1 point higher on the all-round intelligence index
Artificial Analysis Intelligence Index v4.3.2 — one score averaging 10 hard tests
48
▲ up from 47 · +1 point
Costs half as much as the model it replaces
Price per 1M tokens — what developers pay, input / output
$2 / $10
▲ down from $4 / $20 · 50% cheaper
OpenAI's most recent flagship release is GPT-6.1 Sol (GPT line, September 29, 2026); its other flagship line's latest model is GPT-6 Astra (GPT Astra line, September 4, 2026). It shipped 9 days ago, and the line's usual gap is about 75 days. We only report official announcements, never rumored release dates.
Anthropic's most recent flagship release is Claude Opus 5.5 (Claude Opus line, September 22, 2026); its other flagship line's latest model is Claude Fable 5.1 (Claude Fable line, September 1, 2026). It shipped 16 days ago, and the line's usual gap is about 60 days. We only report official announcements, never rumored release dates.
Google's newest flagship model is Gemini 3.1 Pro, released on February 19, 2026. It has been 231 days since then, longer than the line's usual gap of 125 days. Gemini 4 Argon is open only to selected partners so far; Google said on Sep 30, 2026: “Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program”. We only report official announcements, never rumored release dates.
The most recent releases we track are Claude Haiku 5.5 by Anthropic (Oct 7, 2026), Mistral Large 4 by Mistral (Oct 6, 2026) and GPT-6.1 Sol by OpenAI (Sep 29, 2026). The archive lists every release, newest first.
Each line's usual gap is the median of its last five gaps between releases, counting launches less than a week apart as one. A line is Due from 85% to 100% of that gap, Past its usual gap after that, Mid-cycle in between, and Just released during the first third. It is a reading of the line's own rhythm, not insider information.
We tested a simple forecast, the usual gap plus or minus 15 days, against every release in our archive: it would have been right fewer than 1 time in 5. Release rhythms are too irregular for a date to be honest, so we show where each line stands instead.
How the status works. Each line's usual gap is the median of its last five gaps (releases less than a week apart count as one launch). The label says where the line stands in its own rhythm: Due from 85% to 100% of the usual gap, Past its usual gap after that, Mid-cycle in between, Just released for the first third. Lines are listed in that order, bigger makers first.
Why no dates. We tested "usual gap ± 15 days" against every release in our archive: it would have caught fewer than 1 in 5 of them. AI release rhythms are too irregular for a date to be honest — so we show the rhythm, not a prediction. In our archive, a new model came within a month about 1 time in 3 when a line was reaching its usual gap, and less often once it was past it: that is why we never count "days late".
A line whose wait, measured in its own usual gaps, is longer than almost every past gap in our archive (fewer than 5 lasted that long) is marked Quiet; after a year of silence it becomes Inactive and moves to the Archive. "Discontinued" is only for lines the maker has officially closed.
These are statistical estimates, not insider information. Actual release dates depend on development progress, strategic decisions, and competitive dynamics.
Disclaimer: Next AI Model is an independent project and is not affiliated with, endorsed by, or connected to any AI company mentioned on this site. Status labels are statistical readings of publicly available release dates. This content does not constitute financial or investment advice. All trademarks and product names belong to their respective owners.