❯ GPT-6 Astra’s blog was pulled and reposted as hallucination and other scores kept changing
a rocky launchOpenAI published its GPT-6 Astra announcement on September 3 local time, and the process went sideways: the blog went up, came down, and the page stayed unreachable for a long stretch. The company blamed a content management system failure and then an internet outage, and stressed the takedown had nothing to do with benchmark scores. After the post returned, several evaluation figures were revised repeatedly.
the revisionsComparing archived snapshots, Fortune found Astra’s hallucination rate first halved from 4.2% to 2%, then went back to 4.2%; the same metric for GPT-5.6 Sol fell from 12.2% to 9.4% before returning to 12.2%. The largest move was Sol’s score on OpenAI’s internal ExploitBench, raised from 5.5% in the first version to 11.5% later; OpenAI said afterward it was considering reverting to 5.5% because the 11.5% result reflects a reasoning level not commercially available. Some adjustments happened before the blog was first published.
the disputeOpenAI’s explanation is that the changes ensure the numbers represent the best estimate of usable model performance, and that test conditions affect results. But the system card does not explain the methods — for the internal hallucination benchmark it gives “barely any details about the evaluation.” That pushes the industry’s long-running fight over benchmark credibility back into view, a fight that has previously involved Meta among others. The call from practitioners is for a norm: when a score changes, state what changed in the evaluation conditions.
two readingsRunning alongside the doubt is Wall Street’s opposite reading: Goldman Sachs Delta-One desk head Rich Privorotsky praised Astra’s importance to the industry in a research note, arguing that a leading model genuinely breaking away is exactly what the AI bull case has been waiting for. The ones who have to re-weigh things are people using benchmarks for model selection and investment calls — when numbers in the same announcement change twice in three days, customers and investors cannot tell whether the model fits, and OpenAI’s market narrative takes a discount with them.
▪ SIGNALThe numbers moved twice and came back — what was lost is not those percentage points but the credibility of every official score that follows.