Benchmaxxed: Google's New Gemini 4 Aces SAT, Struggles With Actual Job, "Skeptical Employees" Admit

After a year in which Google's AI roadmap resembled a Waymo stuck in a roundabout, the search giant on Wednesday finally unveiled Gemini 4 "Argon", its long-awaited flagship model. The market cheered, briefly: according to Goldman's closing equities color, GOOG traded +2% after hours on the "Argon" announcement. 

Then Bloomberg reported that some of the people who have actually use the thing - as in Google's own engineers - aren't nearly as impressed as the leaderboard. The stock promptly faded. 

According to Bloomberg's Julia Love and Davey Alba, Gemini 4 "performed well on benchmarks widely used to gauge model efficacy" but "does less well when employees actually put it to work," particularly on coding. One insider said the model "isn't particularly adept at front-end design" - the part of software that decides how apps and websites look and feel. Which is a bit awkward for a company whose entire business is, well, things you look at on a screen.

Google, naturally, disagrees. It told Bloomberg it would be "inaccurate" to say Gemini 4 is underperforming in coding, and pointed back to last week's remarks by DeepMind boss Koray Kavukcuoglu, who said "it's a certainty that we are always gonna be at the frontier." Another Google employee said there is "large consensus" internally that the model is frontier-class. Which is the kind of thing one tends to say when there isn't.

The industry has a word for this. "Benchmaxxing" is when engineers optimize