The Hidden Gap in Meta’s Benchmark Denial Nobody Noticed for Nine Months
The two claims aren’t quite the same thing, and that distinction is the most interesting part of this story.Nobody noticed for nine months.In April 2025, Meta launched Llama 4 with benchmark numbers strong enough to put it near the top of LMArena, the popular head-to-head AI leaderboard. Within…
The two claims aren’t quite the same thing, and that distinction is the most interesting part of this story.Nobody noticed for nine months.In April 2025, Meta launched Llama 4 with benchmark numbers strong enough to put it near the top of LMArena, the popular head-to-head AI leaderboard. Within days, the numbers were under serious question. Nine months later, Meta’s own outgoing chief AI scientist confirmed, on the record, that the results had been manipulated.Most coverage of this story treats it as a straightforward “Meta lied, then got caught” narrative. The more accurate version is a little more interesting, and a little more uncomfortable: two separate accusations were made in April 2025, Meta specifically denied one of them, and the admission that came later confirmed something related but distinct. Getting that distinction right matters, because it’s a small case study in exactly the kind of confusion that lets benchmark disputes stay unresolved for months.What actually happened at launch, and nobody disputes this partWhen Llama 4 launched, Meta’s press materials highlighted Maverick’s LMArena score of 1417, ranking it above GPT-4o and just behind Gemini 2.5 Pro at the time. Maverick briefly took the number-two spot on the leaderboard.Four events, nine months apart, that most coverage compresses into a single ‘Meta got caught’ story.The problem: the version Meta submitted to LMArena wasn’t the model anyone could actually download. It was an unreleased, specially tuned variant, internally labeled Llama-4-Maverick-03–26-Experimental, built specifically to perform well against human preference voting.LMArena confirmed this itself, stated that Meta’s interpretation of its submission policy didn’t match what the platform expects from model providers, and updated its rules afterward specifically to prevent it happening again. This part of the story was never in dispute. Meta acknowledged using an “experimental chat version” for the benchmark.A separate rumor, and a specific denialAround the same time, a different and more serious accusation started circulating on X and Reddit: that Meta had trained Llama 4 directly on the test sets used to evaluate it, which would mean the benchmark scores weren’t measuring genuine capability at all.Meta’s VP of generative AI, Ahmad Al-Dahle, addressed this directly on X, calling the claim “simply not true.” That denial was specific. He wasn’t denying that an experimental model had been used on LMArena, since that was already public and acknowledged. He was denying the narrower, more damaging claim that the model had been trained on the actual evaluation data.Nine months later, a different admissionIn January 2026, Yann LeCun, Meta’s chief AI scientist, gave a lengthy exit interview to the Financial Times on his way out the door to start his own company. Asked about Llama 4, he said the results were “fudged a little bit,” and that the team “used different models for different benchmarks to give better results.”That’s a real admission of benchmark manipulation. It is not, however, an admission that Meta trained on test sets, the specific thing Al-Dahle denied. What LeCun described is a different practice: running several model variants, picking whichever one scored highest on each individual benchmark, and presenting the resulting composite table as if one model had achieved every number in it.Independent testers running their own evaluations on the publicly released model had consistently found lower scores than Meta’s published figures, which is consistent with LeCun’s account.So the accurate summary isn’t “Meta denied it, then admitted the exact same thing.” It’s closer to: Meta denied one specific form of manipulation and never directly addressed a different, related form that its own chief scientist confirmed nine months later. Both are legitimate reasons to distrust the original numbers.They’re just not the same reason, and conflating them is exactly the kind of imprecision that makes it easy for a company to technically deny something without the denial actually resolving the underlying doubt.Two different claims, nine months apart. Only one of them was ever directly denied.Why the fallout was real, even if the story got compressedThe consequences inside Meta suggest this wasn’t treated internally as a minor wording dispute. LeCun described Zuckerberg as “really upset,” losing confidence in the team behind the launch and sidelining the entire GenAI organization as a result.That reaction set off a chain that reshaped Meta’s AI division: a restructuring into Meta Superintelligence Labs, a roughly $14.3 to $15 billion investment for a 49% stake in Scale AI, and Scale’s 28-year-old CEO, Alexandr Wang, installed to lead the new organization, with LeCun himself reporting to him before his eventual exit.A benchmark dispute that started as a leaderboard placement ended up restructuring a research division and moving billions of dollars. That scale of consequence is unusual, and it’s a reasonable signal that whatever happened internally around Llama 4’s numbers was taken more seriously behind closed doors than Al-Dahle’s brief public denial let on.Where this fits the larger patternThis isn’t an isolated incident, and it predates every other case in this ongoing look at benchmark trust. Submitting a specially configured variant to a favorable evaluation environment, then letting a headline number spread without the configuration details attached, is the same underlying shape as the harness disputes and post-launch benchmark edits documented elsewhere in this series.The mechanism varies. Model-variant cherry-picking, harness configuration, post-publication edits. The incentive behind all of them is identical: a launch-week benchmark table is judged by how it looks in a screenshot, not by how carefully it was produced, and that pressure doesn’t appear to be unique to any one lab.The specific lesson from the Llama 4 case is narrower and more useful than “don’t trust benchmarks.” It’s this: when a company denies an accusation, check exactly what was denied. A precise, on-the-record denial of one specific claim is not the same as a denial of the broader concern, and treating it as one can leave a genuinely misleading number standing for months after the narrower rumor was already put to rest.FAQ: Meta, Llama 4, and the benchmark controversyDid Meta admit to manipulating Llama 4’s benchmark scores? Meta’s chief AI scientist, Yann LeCun, told the Financial Times in January 2026 that Llama 4’s benchmark results were “fudged a little bit,” explaining that different model variants were used for different benchmarks to produce better individual scores. This is a separate claim from the one Meta had denied nine months earlier.What exactly did Meta deny? In April 2025, Meta VP Ahmad Al-Dahle denied a specific rumor that Llama 4 had been trained directly on benchmark test sets. He did not address, and was not specifically asked about, the practice of submitting different model variants to different benchmarks.Is the LMArena controversy the same as the benchmark-fudging admission? No. The LMArena episode, where Meta submitted an unreleased, specially tuned model variant to the leaderboard, was confirmed and acknowledged by both Meta and LMArena at launch. LeCun’s later admission concerns a separate practice: selecting different model variants for different benchmarks rather than testing one consistent model across all of them.What happened at Meta after the admission became public? According to LeCun, Mark Zuckerberg lost confidence in the team behind Llama 4 and sidelined the entire GenAI organization. Meta subsequently restructured into Meta Superintelligence Labs, paid roughly $14.3 to $15 billion for a stake in Scale AI, and brought in Scale’s CEO to lead the new division.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!The Hidden Gap in Meta’s Benchmark Denial Nobody Noticed for Nine Months was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI