Companies Are Throttling AI Use Because of Costs — Here's What the Bills Actually Look Like

A 404 Media story making rounds this week says companies are actively throttling employees' AI access because it's getting too expensive. That tracks. But the real story is how much the price spread matters here.
If your team is hitting the flagship tier — say, claude-fable-5 at $10.00 in / $50.00 out per million tokens — a heavy usage month can get ugly fast. GPT-5.5 at $5.00 in / $30.00 out isn't cheap either.
But here's the thing: not every task needs a flagship model. Gemini 3.5 Flash runs $1.50 in / $9.00 out. DeepSeek V4 Flash comes in at $0.14 in / $0.28 out. That's roughly a 70x difference in output costs between the priciest and cheapest options on this list.
For a lot of internal tooling — summarizing docs, drafting emails, answering routine questions — the cheaper models do the job fine. The companies throttling AI use may actually be solving the wrong problem. It's not always about using less AI; it's about routing the right tasks to the right price tier.

GPT-4o Mini at $0.15 in / $0.60 out is still a solid mid-budget workhorse. Kimi K2.5 at $0.57 in / $2.41 out sits in a useful middle ground too.
The budget problem is real. But it's also fixable — if you're paying attention to which model you're actually using.
