The Output Token Gap Is Where Your Budget Actually Lives

People fixate on input prices, but output tokens are where the real spend happens — and right now the spread between the cheapest and priciest models is enormous. Let's make that concrete.
On the high end, gpt-5.5 charges $30.00 per million output tokens. Claude Opus 4 (any variant) comes in at $25.00 out. Those are serious numbers if you're generating long responses at volume.

Drop down a tier and claude-sonnet-4-6 cuts that to $15.00 out — still not cheap, but nearly half the Opus price for a model that handles most production workloads just fine.

Then things get interesting fast. deepseek-v4-flash (available on both our US region and our Singapore region) charges just $0.28 per million output tokens. That's over 100x cheaper than gpt-5.5's output rate. deepseek-v4-pro isn't far behind at $0.87 out. Even gpt-4o-mini — OpenAI's budget option — is only $0.60 out.
The practical takeaway: if your app generates a lot of text, the output price is the number to watch, not the input price. A model that's $1.00 cheaper on input but $10.00 more expensive on output will cost you more almost every time.
Match your output volume to your model tier, and you'll save more than any other single decision you make on infrastructure.
