Artificial Analysis published its evaluation of Google’s Gemini 4 Argon on September 30, 2026, reporting that Gemini 4 Argon (high) scores 53 on the Artificial Analysis Intelligence Index, level with GPT-6 Astra (max). In its Gemini 4 Argon analysis, the firm writes that “Google returns as one of the top three labs on intelligence.”
On this page
Intelligence Index scores
According to Artificial Analysis, “Gemini 4 Argon (high) scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52).” The same analysis gives Gemini 3.8 Flash (high) a score of 41.
| Model (setting) | Intelligence Index (Artificial Analysis, Sept. 30, 2026) |
|---|---|
| Gemini 4 Argon (high) | 53 |
| GPT-6 Astra (max) | 53 |
| GPT-6.1 Sol (max) | 52 |
| Gemini 3.8 Flash (high) | 41 |
The table lists only the models and scores given in the September 30 analysis.
Hallucination rate on AA-Omniscience
The widest gap in the analysis is in hallucination rate on AA-Omniscience. The firm writes: “On AA-Omniscience, Gemini 4 Argon has a 15% hallucination rate, the lowest of any model scoring 45+ on the Intelligence Index, compared with 51% for GPT-6 Astra (max) and 54% for GPT-6.1 Sol (max).”
| Model | AA-Omniscience hallucination rate |
|---|---|
| Gemini 4 Argon | 15% |
| GPT-6 Astra (max) | 51% |
| GPT-6.1 Sol (max) | 54% |
Artificial Analysis adds that “Argon is much more likely to acknowledge when it does not know an answer rather than guess incorrectly.”
Agentic evaluations
Artificial Analysis also reports two agentic results. Gemini 4 Argon “ranks #1 on AutomationBench-AA at 78%, 7 points ahead of Claude Sonnet 5.5 (max, 71%).” On Terminal Bench 4, it “achieves 57%, a + 53 point improvement from Gemini 3.1 Pro Preview, only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%).”
What it changes
On Artificial Analysis’s measurements, Gemini 4 Argon ties the top OpenAI model in this comparison on the composite index while giving wrong answers far less often on AA-Omniscience. The firm’s own framing is that the result places Google back among the top three labs on intelligence. On terminal-based agent work, Anthropic and OpenAI models still score higher in the same analysis.
All scores in this piece are Artificial Analysis’s own runs. Gemini 4 Argon’s results are recorded with their source and date on our benchmark leaderboard, and the model’s release is in the AI model release timeline.
Sources
- Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved — Artificial Analysis




