The Cheap AI Model Just Learned to Use a Computer
Anthropic's Claude Haiku 5.5 costs about a tenth of its predecessor on most requests and scores far higher at operating software. The savings depend on keeping requests small, and a new open-source agent shows one way to do that.

Most AI news is about the top of the market: the biggest model, the hardest benchmark, the highest price. This week the bigger story is at the bottom. Anthropic released Claude Haiku 5.5, its first new small model in nearly a year. On most requests it costs about a tenth of what Haiku 4.5 cost, and on a test of operating a computer its score went from 15.7% to 72.4%.
Both changes matter a lot, because the small model is the one doing most of the routine work.
What actually changed
According to The New Stack's report on the launch, Haiku 5.5 is priced at $0.10 per million input tokens and $0.50 per million output tokens for requests under 100,000 tokens. Haiku 4.5 cost $1 and $5. Requests over 100,000 tokens cost $0.50 and $2.50, which is still half the old price.
Anthropic says about 90% of Haiku 4.5 requests fell under the 100,000-token line. It also put the average saving at about 75%, not 90%, because a new tokenizer uses slightly more tokens for the same task. That is an honest detail and I'm glad they included it. The cut is still very large.
The performance numbers, as Anthropic reported them:
- OSWorld 2.1 (using a computer): 72.4%, up from 15.7% for Haiku 4.5. OpenAI's GPT-6 Luna scored 48.9%.
- GDPval-AA v2.1 (knowledge work): 1,620, up from 735. GPT-6 Luna scored 1,437.
- Terminal-Bench 4.0: 39.2%, up from 0%. GPT-6 Luna scored 16.4%. Anthropic's larger Sonnet 5.5 scored 70.6%.
It is also the first Haiku with effort controls, which default to medium, so you can choose how hard the model thinks about each job. Anthropic is aiming it at live customer support, browser use, compaction (condensing long conversation histories) and database queries.
Why the computer-use jump is the real headline
The price cut gets attention, but the OSWorld result changes more. Computer use means the model looks at a screen, finds the right button, types into the right field and moves through software built for people. Until now that took a large, expensive model. Older small models mostly failed at it, as Haiku 4.5's 15.7% shows.
Think of an intern who could file and sort but couldn't use the office software. Now that intern can handle the screens, and their hourly cost just dropped by an order of magnitude. That makes a lot of jobs worth automating that weren't worth it before: tedious admin screens, legacy dashboards, tools that never got an API.
The price tag comes with a condition: stay under 100,000 tokens. The model is only cheap if you keep each request small.
The fine print: the 100,000-token line
The pricing works like a taxi meter with a city limit. Inside the limit you pay a fraction of the old fare. Go past it and the rate jumps to five times as much. It's still cheaper than before, but much of the saving is gone.
So the practical question for anyone building with AI agents is how much of each request is real work and how much is overhead. The answer is often uncomfortable.
The hidden cost: tool menus
A second story from the same week shows where a lot of tokens go. The creator of Pi, an open-source coding agent, measured how much space common MCP tool servers take up before the agent does any work. (MCP is the standard way AI agents connect to outside tools.) Chrome DevTools MCP took about 18,000 tokens, or 9% of a 200,000-token window. Playwright MCP took about 13,700 tokens for 21 tools.
It's like making someone read a hardware store's entire catalog every time they need a hammer. The agent hasn't started yet and it has already used a large share of its budget on a list of things it might never touch.
Pi 1.0 handles this differently. Instead of loading every tool definition into the prompt, the model sees one line per server. When it needs something, it uses "Codemode", a sandboxed JavaScript runner with no file system, network or timers, to find and call the right tool and return only the useful output. A per-tool setting can expose a tool, hide it behind Codemode or block it entirely. The example given exposes a code search tool, hides read-only "get" tools and blocks anything that deletes. The default budget for tool declarations is 3,000 tokens. In one measured request, prompt tokens fell from about 5,300 to 3,300.
Together, these two stories describe one shift: cheap models make agents affordable, and lean prompts keep them affordable.
What I'd keep in mind
- These are vendor numbers. The benchmarks are Anthropic's own as reported by The New Stack. Independent results usually follow within weeks, and I'd wait for them before treating 72.4% as settled.
- It isn't the only cheap option. Z.ai's GLM-5.3-Flash scores 1,647 on GDPval-AA according to Artificial Analysis, slightly ahead of Haiku 5.5 on that test.
- Small doesn't mean best. Sonnet 5.5 is still well ahead on hard terminal work (70.6% against 39.2%). The sensible approach is to match the model to the job: a cheap model for high-volume routine work and a bigger one when the task is hard.
- Blocking is a safety feature too. Pi's block list does more than save tokens. An agent can't misuse a delete tool it never had.
I find this exciting. For a few years, "AI agent" has often meant an expensive demo. A small, cheap model that can actually operate software, paired with a discipline for keeping prompts lean, is what turns demos into routine work. A capable model got cheap this week, and that will probably affect more people than the next frontier model does.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- The New Stack - Anthropic Claude Haiku 5.5 - https://thenewstack.io/anthropic-claude-haiku-5-5
- The New Stack - Pi agent, MCP and Codemode - https://thenewstack.io/pi-agent-mcp-codemode/
Quick answers
How much does Claude Haiku 5.5 cost?
For requests under 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens. Over 100,000 tokens it costs $0.50 and $2.50. Haiku 4.5 cost $1 and $5.
Is it really 90% cheaper?
On the per-token price for most requests, yes. Anthropic puts the average real saving at about 75% because a new tokenizer uses slightly more tokens per task.
How good is Haiku 5.5 at using a computer?
It scored 72.4% on OSWorld 2.1 offline according to Anthropic, up from 15.7% for Haiku 4.5 and ahead of GPT-6 Luna's 48.9%. These are vendor-reported numbers, and independent results have not yet been published.
Why do tool definitions matter for AI costs?
Agents often load every tool description into the prompt. Pi's creator measured Chrome DevTools MCP at about 18,000 tokens before any work. Pi 1.0 keeps tools out of the prompt until the agent needs them, and that cut one request from about 5,300 to 3,300 prompt tokens.