GPT-6 Astra: OpenAI's new model, and the benchmark row nobody's mentioning in the press release
OpenAI has shipped GPT-6 Astra with headline-grabbing benchmark scores, but a look at the small print - and a sceptical Hacker News crowd - suggests the comparisons doing the AGI talk aren't quite apples-to-apples.
What’s actually happened
OpenAI has released a new model, GPT-6 Astra, and published an accompanying system card as part of the rollout. That much is solid: there’s a model, there’s documentation, and it’s being discussed in real time on Hacker News, where the launch thread has racked up over 1,600 points and well over a thousand comments in less than half a day. People are testing it, arguing about it, and — inevitably — asking whether this is the moment OpenAI finally gets to say “AGI” with a straight face.
Two follow-on claims are doing the rounds: that Astra has made “major gains” on the Artificial Analysis Coding Agent Index, and that it has posted a striking result on ARC-AGI-3, the puzzle-based benchmark specifically designed to resist being gamed by pattern-matching. Both are being treated in some corners of the thread as evidence that Astra represents a genuine step change.
The benchmark claim — and the catch
Here’s where it gets murkier. One commenter on the Hacker News thread points out that the ARC-AGI-3 scorecard being circulated compares scores generated under different conditions. The chart reportedly shows GPT-5.6 Sol scoring just 7.8% — despite the benchmark’s own documentation apparently estimating that with a different test harness (OpenAI’s “responses API” setup), that same model would score somewhere around 30%. In other words, Astra’s number may have been generated with a more favourable harness than the one used to produce the comparison figure for its predecessor, making the “leap” look considerably bigger on paper than it may be in practice. The same commenter suggests a rival model, Opus 5, would likely see a similar bump if measured the same way — which rather undercuts any claim that Astra has pulled dramatically ahead.
This is a familiar problem in AI benchmarking: the scaffolding around a model (the “harness”) can matter as much as the model itself, and comparing a leader’s best-case number against a rival’s worst-case number is not a fair fight, however striking the headline chart looks.
Is this actually AGI? Don’t hold your breath
Unsurprisingly, the thread has spiralled into the now-ritual argument about whether any of this constitutes artificial general intelligence. Some commenters are ready to call it, pointing to Astra’s apparent edge over other frontier models. Others are having none of it, arguing the definition of AGI has been quietly diluted to fit whatever a leading lab has just shipped, rather than the other way round. Even the person credited with creating ARC-AGI-3 is quoted acknowledging that progress on the benchmark happened roughly twice as fast as they expected — which is a genuinely interesting data point, but is not the same thing as confirming AGI has arrived. Benchmark saturation and general intelligence are not synonyms, whatever a press release implies.
So who is actually affected
If you’re an OpenAI subscriber, you’ll likely see Astra appear as an option in the coming days, and developers building on the API will get access to whatever capabilities it brings to coding and agentic tasks. Nobody else needs to do anything. There’s no security issue, no data risk, no action required from ordinary users.
The takeaway
A new model has shipped, and it’s clearly capable. But the specific numbers being used to sell it as a historic leap rest on benchmark comparisons that don’t obviously compare like with like, and the AGI framing is opinion, not measurement. Worth watching independent, harness-consistent evaluations before taking the party line at face value.