❯ Artificial Analysis: Gemini 4 Argon matches GPT-6 Astra on intelligence with a hallucination rate of just 15%
Same score, a third of the hallucinationsIndependent benchmarking firm Artificial Analysis said Google’s new Gemini 4 Argon (high reasoning) matches OpenAI’s GPT-6 Astra (max reasoning) on its Intelligence Index. The gap is in hallucination: Gemini 4 Argon’s rate is 15%, versus 51% for GPT-6 Astra.
How hallucination is measuredThe Intelligence Index aggregates ten tests covering coding, agentic tasks and scientific reasoning. The hallucination rate measures how often a model makes things up when it does not know the answer; lower is more reliable. When it tested GPT-6 Astra in September, Artificial Analysis noted its rate had already fallen by nearly half from GPT-5.6 Sol’s 92%.
Another factor in enterprise choicesFor companies using models to write reports, research or handle legal and financial documents, a model that fabricates less at equal capability saves a lot of manual checking. Still, these are third-party results on specific test sets, and real performance depends on the task. Gemini 4 Argon launched this week alongside Claude Sonnet 5.5 and GPT-6.1 Sol, and also ranks near the top of the coding agent leaderboard.
▪ SIGNALAs frontier models’ intelligence scores converge, knowing what they don’t know is becoming the metric that separates them.