Gemini 4 Argon Launches: Beats GPT-6 Astra on Key Benchmarks, Takes On Google’s 800,000-Line Kernel…

Gemini 4 Argon Launches: Beats GPT-6 Astra on Key Benchmarks, Takes On Google’s 800,000-Line Kernel MigrationA million-token output ceiling and an ongoing 800,000-line Rust migration headline Google’s pitch. Wider access and everyday results are still ahead.Gemini 4 Argon’s long-output announcement…

Annons
Annons
Gemini 4 Argon Launches: Beats GPT-6 Astra on Key Benchmarks, Takes On Google’s 800,000-Line Kernel MigrationA million-token output ceiling and an ongoing 800,000-line Rust migration headline Google’s pitch. Wider access and everyday results are still ahead.Gemini 4 Argon’s long-output announcement and in-progress kernel migrationOn DeepSWE v1.1, Google reports Gemini 4 Argon at 77.9%, ahead of GPT-6 Astra’s 74.1%. But on FrontierSWE v2, Argon scores 55.0% while Astra reaches 65.5%. Meanwhile, Google says Argon agents are working on a C/C++-to-Rust migration covering 800,000-plus lines of the Fuchsia Zircon kernel. Google says audits, emulation tests, and review must precede production.Google reports Gemini 4 ArgonThat’s the story behind Google’s September 30 launch: a model aimed at longer, multi-step work, with some striking results and clear limits. The promise is relevant if you use AI to make things from large briefs, codebases, or video. But Google’s headline output figure, the benchmark table, and what you can actually use today are three different claims.Gemini 4 Argon is built for longer jobsGoogle says Argon’s maximum output is expanding from 64,000 tokens to one million. Output measures generated tokens. Input context measures what the model can read; the two are separate limits.A larger ceiling can give an agent room to continue through a long task without stopping as often. More room to work. Coherence and polish still have to be checked. A million-token limit alone can’t tell you how much of that work will survive review.A larger output cap isn’t a quality guaranteeVals’ Argon profileThere’s also a detail that deserves a careful read. Google’s announcement describes a one-million-token output limit. Vals’ Argon profile, which documents its own evaluation setup, lists a one-million-token context window and a 262,144-token maximum output setting.That may reflect the cap Vals used for its tests rather than a limit on Google’s model. It doesn’t verify that an ordinary user can request a million-token response today. The model remains in staged rollout, and Google hasn’t published a general-access date.Google is already using it for real engineering workThe 800,000-plus figure refers to the scale of code Argon agents are working on as Google migrates C and C++ codebases to Rust. Google says critical rewrites like the Fuchsia Zircon kernel are undergoing automated and manual audits, emulation tests, and review before production.Those checks are still ahead of production. In short, Google describes work in progress. It hasn’t said the kernel migration is complete or ready for deployment.The smaller decoder example is more concreteA smaller example makes the kind of work more tangible. Google says Argon agents replaced 32,000 lines of SIMD code in an existing Rust port of its libgav1 video decoder.After repeated profiling and experiments, the revised safe-Rust decoder ran 2.7 times faster than that earlier Rust port and produced identical video output, according to Google. Google’s baseline is the earlier Rust port. The 2.7x figure measures decoding speed; it says nothing about AI video generation.For anyone making videos or other media, the interesting part is the process: an agent can reportedly work through profiling data, compiler output, and repeated code changes. That’s a software-engineering example, not evidence that Argon will write a better script or find better footage for your next edit.Strong scores, with some clear exceptionsGoogle’s table shows Argon ahead of Astra on several selected tests. It reports 77.9% to 74.1% on DeepSWE v1.1, 91.9% to 89.6% on Vibe Code Bench, and 68.9% to 63.1% on the Vals Index. The independently maintained Vals profile places Argon first among 41 models on that index.The index combines task scores across work categories; its GDP weighting estimates the relative importance of those tasks. It does not measure actual GDP growth or the value a business will realize.But the same table has a direct counterexample: Astra leads on FrontierSWE v2, 65.5% to Argon’s 55.0%. Astra also edges Argon on Terminal-bench 4.0, 58.2% to 57.4%. That unevenness matters more than a “winner” label if your task looks like one of the tests Argon loses.Google’s methodology PDFThe comparison isn’t one uniform, independently run tournament. Google’s methodology PDF says some Argon scores were computed by Google with its own harnesses, while competitor results often come from provider-reported numbers or public leaderboards.DeepSWE is a particularly clear example: Google used a mini-swe harness for Argon and compared it with Astra’s public leaderboard result. These numbers are useful signals, not a guarantee about your own work.Selected benchmark results show both Argon leads and an Astra winA benchmark lead is a reason to pay attention, not a verdict on your workflow.Defenders get access first. Everyone else waits.Argon is first rolling out to trusted cyber defenders through Google’s Fairwind Program. Google says this group and its own security teams will get a version without cyber guardrails for defensive work.The company also describes Argon finding a critical vulnerability in healthcare software through a Wiz demonstration. Google reports these early findings; this article hasn’t independently reproduced them.Google’s rollout note says broader access will follow for developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers. Google describes a next phase but leaves its timing open. No date yet.If you’re a creator or developer hoping to try Argon, the announcement doesn’t tell you when your account will get it. The early access sequence also reflects the model’s cybersecurity capabilities and the safeguards Google says it is still strengthening.Price is clear; the everyday value is notGoogle’s introductory API rate is $2 per million input tokens and $10 per million output tokens. Cached input is 95% cheaper than the input rate. After the introductory period, the posted rates become $4 per million input tokens and $20 per million output tokens. Google hasn’t said when the introductory price ends. The API rates are separate from a Google AI Ultra subscription, and neither figure is a typical creator’s bill.For a long writing or research task, a million-token output ceiling sounds generous. It could let a model return a much larger draft or sustain a longer agent run. The announcement doesn’t show whether the result is useful, accurate, or cheaper after review and retries. There’s no independent production test here, and broad access hasn’t arrived.Argon’s launch points toward models that can stay with demanding tasks longer, and agents that Google says are already helping with real engineering work; the benchmark wins make that direction worth watching. But an in-progress migration, a vendor’s score table, and an announced output limit don’t yet answer the practical question: will Argon help you finish the thing you’re making with less rework?When access does arrive, give it one real task from your own pipeline. Track how much of the output survives editing and how many retries it takes. That’s the cost that turns a token ceiling into a production result.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!Gemini 4 Argon Launches: Beats GPT-6 Astra on Key Benchmarks, Takes On Google’s 800,000-Line Kernel… was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →
Annons
Annons