Claude Mythos Preview vs Grok 4: Who Rules the Benchmark?

Claude Mythos Preview vs Grok 4: Who Rules the Benchmark?
Right now, Claude Mythos Preview owns GPQA Diamond at 94.6%, while Grok 4 is the name everyone's shouting in coding circles. These are two genuinely different weapons — and picking the wrong one costs you real performance.
Science and Reasoning: Mythos Takes the Crown

GPQA Diamond is one of the toughest knowledge benchmarks around — doctoral-level science questions designed to trip up both humans and models. Claude Mythos Preview's 94.6% is the highest confirmed score in the current landscape. That's not a small margin over the field; the top 15 models are separated by roughly 3 percentage points overall, so leading on a hard single benchmark actually means something here. If your work lives in chemistry, biology, or hard physics, Mythos earns that edge.
Coding: Grok 4 Punches Harder

Flip to coding benchmarks and the story changes. Grok 4 sits at the top alongside Claude Opus 4.6, not Mythos. That split matters if you're running automated code review or agentic dev pipelines — Grok 4's coding performance is where it separates itself from the science-focused competition. Claude Opus 4.6, for what it's worth, is priced at $5.00/1M in and $25.00/1M out, so Anthropic's coding-capable tier isn't the priciest option on the board.
Arena Elo: Close Enough to Argue About
On the Chatbot Arena Elo ratings, Anthropic leads overall at 1,503, with xAI (Grok's parent) right behind at 1,495 — an 8-point gap across the whole provider. That's genuinely tight. Neither company has run away from the other in head-to-head human preference voting, which tells you both models are delivering real value in conversation.
Verdict: Which One Should You Pick?
Go with if your use case is scientific research, complex reasoning, or graduate-level Q&A — 94.6% on GPQA Diamond isn't decoration. Choose if coding throughput is your bottleneck; it's the benchmark leader there and xAI's Arena Elo score shows it holds up in general use too. Neither model is a clear loser — you're picking a specialty, not settling.
