The Hidden Setting That’s Rewriting Every AI Benchmark This Year

GPT-6 Astra, GPT-5.6 Sol, and Claude Opus 5 all hit the same wall this year, and almost nobody explained why.Same model, same weights — the only thing that moved was a setting nobody disclosed.I spent the last week untangling one number: GPT-6 Astra’s 99.9% score on ARC-AGI-3. By the end of it, I…

GPT-6 Astra, GPT-5.6 Sol, and Claude Opus 5 all hit the same wall this year, and almost nobody explained why.Same model, same weights — the only thing that moved was a setting nobody disclosed.I spent the last week untangling one number: GPT-6 Astra’s 99.9% score on ARC-AGI-3. By the end of it, I wasn’t looking at one benchmark anymore. I was looking at a pattern that’s quietly reshaping how every AI lab reports its results, and almost no coverage is naming it directly.Here’s the pattern in one sentence: the model isn’t the only thing being tested anymore. The harness around it, the settings, the scaffolding, the memory tricks, is doing as much work as the model itself. And right now, the harness is usually the thing getting left out of the headline.It started with a number that didn’t add upGPT-6 Astra launched on September 3, 2026, with a headline ARC-AGI-3 score of 99.9%, compared against GPT-5.6 Sol’s 7.8%. That gap went everywhere. It’s the kind of number that makes a launch feel like a different category of model.ARC Prize, the organization that built ARC-AGI-3, published its own breakdown the same day, and it told a more careful story. Astra’s 99.9% came from what OpenAI calls a Provider Adapter harness, a setup that lets the model preserve its internal reasoning state between requests instead of starting fresh each turn. Run through ARC Prize’s own provider-neutral Standard harness, the same one used to score every other model fairly, Astra’s result was 62.7%.Sol’s 7.8%, importantly, was measured on that same Standard harness. So the honest comparison isn’t 99.9% vs 7.8%. It’s 62.7% vs 7.8%, still roughly an eightfold jump, and still a genuinely major result. But it’s a different number than the one that spread, and ARC Prize is now insisting both scores get labeled separately going forward, specifically because the gap between them is too large to leave unexplained.This isn’t the first time. It’s not even the first time this year.Here’s what almost nobody connected to the Astra story knows: OpenAI ran into the exact same issue with Sol itself, two months earlier.In July 2026, OpenAI noticed something strange. Sol had solved long-standing open problems in mathematics and beaten entire video games, yet it scored just 7.8% on ARC-AGI-3, a benchmark built around simple 2D puzzle games. That didn’t fit. So OpenAI investigated and found the cause wasn’t the model. It was two API settings, called retained reasoning and compaction, that the official benchmark harness happened to disable.Turn those two settings on, and Sol’s score on ARC-AGI-3’s public task set went from 13.3% to 38.3%, nearly triple, while using six times fewer output tokens. OpenAI published the finding openly, which is genuinely worth crediting. But it also means the same organization now behind Astra’s 99.9% headline had already demonstrated, months earlier, exactly how sensitive this specific benchmark is to harness configuration. They knew what a favorable harness could do to this number before they published one.ARC Prize’s own leaderboard, for what it’s worth, still lists Sol’s verified score at 7.8%. The 38.3% figure came from a different, uncontrolled comparison that the official board doesn’t recognize. That distinction matters, and it’s exactly the distinction that got lost when Astra’s number made headlines two months later.Claude Opus 5 shows the same pattern from a completely different angleOpus 5 scored roughly 30.2% on ARC-AGI-3’s standard harness back in July, a genuinely strong result at the time. But wrap that same model in external scaffolding, specifically Nvidia’s AVO framework or AWS’s Strands framework, and its measured score jumps to 100% and 99.95%, respectively.Same three models, two scores each — one from the standard harness, one from a provider-specific setup. The gap is the story.Same model. Same weights. A more than threefold score difference, entirely explained by what’s wrapped around it. Nobody’s accusing Anthropic of manipulating anything here. The point is bigger than any one lab: once a benchmark becomes valuable enough to compete over, the scaffolding around the model becomes part of the competition, whether or not that’s what the benchmark was built to measure.The line ARC Prize is trying to drawARC Prize co-founder François Chollet has been explicit about where he thinks the line sits: general-purpose API settings that any developer can access, and that weren’t built specifically to game this one benchmark, are fair game. Custom harnesses engineered around a specific test are not.It’s a reasonable line. It’s also a line that’s genuinely hard to enforce after the fact, and one that most readers of a benchmark headline will never know exists. Retained reasoning and compaction are production features, available to any developer using the API.That makes them legitimate by Chollet’s standard. But a reader seeing “38.3%” or “99.9%” in a headline has no way to know whether they’re looking at a raw model capability or a raw model plus a specific configuration decision that happened to be available to whoever ran the test.The number that got buried under all of thisWhile the ARC-AGI-3 story dominated coverage, a quieter number tells you more about where these models actually stand. On the Artificial Analysis Intelligence Index, an aggregate measure built from many benchmarks rather than one, GPT-6 Astra scores roughly 61. GPT-5.6 Sol scores roughly 60.9. Claude Fable 5.1 scores roughly 66, ahead of both.That’s not a generational leap. On broad, general capability, Astra and its own predecessor are statistically tied. The models’ huge, headline-grabbing differences live almost entirely inside a small number of specific benchmarks, the ones sensitive to exactly the harness effects described above. The boring aggregate number is arguably the most honest one in this entire story, and it’s the one that got the least attention.What to actually check before you trust a benchmark headlineThree questions, every time you see a big benchmark number in a launch announcement:Was this measured on the benchmark’s own standard harness, or a provider-specific setup? If a lab is comparing its own model’s best-case score against a competitor’s standard-harness score, that’s not a fair fight, even if nobody involved is lying.Does the source publish both numbers, or just the flattering one? ARC Prize now publishes both explicitly, which is the right response to this problem. Most launch materials still lead with one.Is the gain concentrated in one narrow benchmark, or does it show up on an aggregate index too? A model that jumps 90 points on one test and stays flat on a broad index isn’t lying about its result, but it’s telling you something narrower than “smarter model” implies.None of this means these benchmark numbers are fake. Astra genuinely improved on ARC-AGI-3 by roughly eight times on the fair comparison. Opus 5 genuinely benefits from good scaffolding. Sol genuinely got more capable when its production features were switched on. The problem was never the numbers.It’s that a benchmark score without its harness attached isn’t really a number about the model. It’s a number about the model plus a decision nobody disclosed, and this year, that decision moved scores by more than most model upgrades ever have.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!The Hidden Setting That’s Rewriting Every AI Benchmark This Year was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →