Decision models got a price: $0.04 a million
Perplexity shipped a Decisions API that returns probabilities instead of prose, priced at $0.04 per million input tokens with free output, and open-sourced the weights. Three rivals now speak the same schema.

Most of what we call "using AI" is really asking a very articulate stranger a question and then praying the answer comes back in a shape our software can read. We write prompts begging for JSON. We bolt on parsers. We add retries for when the model decides today is the day it opens with "Certainly! Here's the classification you requested."
On 1 October, Perplexity put a price on the alternative. Their new Decisions API runs on pplx-decider-v1-27b, an Apache 2.0 fine-tune of Qwen3.8-27B, and it does not write sentences at all. It returns probabilities. Input costs $0.04 per million tokens. Output is free. The weights are published, so you can also just take the model and run it yourself.
The vending machine, not the conversation
Here is the difference in one image. A general chat model is a brilliant colleague you can ask anything, which also means they can answer anything, including things you did not ask for. A decision model is a vending machine. There are buttons. You press one, an item drops. Nothing else is on the menu.
The API takes three shapes of question. A yes/no question returns the probability of yes, between 0 and 1. A multiple-choice question lets you list up to 255 options and returns a probability for each one, plus the top pick and a confidence. A score question lets you define a rubric of up to 10 levels and returns a probability per level and an expected score.
That constraint is the whole product. A model that can only choose from the list you gave it cannot invent a category that does not exist in your database. It cannot return "Refund_Request" when your system only knows "refund". And because you get a number rather than a verdict, you can set your own bar: route automatically above 0.9, send to a human below it.
A model that can only answer the question you actually asked is not a weaker model. It is a smaller surface for things to go wrong.
The numbers are small enough to stop thinking about
Perplexity's own documentation example uses 367 input tokens. At their price, that call costs roughly $0.000015. Fifteen millionths of a dollar. You would need to run about 67,000 of those before you spent a dollar.
The limits are built for volume rather than eloquence:
- Up to 128 questions per request against the same piece of content, so one call can answer everything you want to know about a single message.
- An input ceiling of 262,144 tokens.
- 10 requests per second per organisation.
- Under 2 seconds for a few hundred tokens, in Perplexity's own tests at the end of September.
- Images accepted as base64 only. The API never goes out and fetches a URL for you, which is a small security decision I like more the longer I look at it.
A whole category formed in about a day
This is the part that made me sit up. Perplexity was not alone. On the same day, Cloudflare shipped Clef and Amazon's Strands Labs shipped Strands Decider 2B. Both are open weights. All three use the same typed schema as TypeSafe's Jev, the product that opened the category. LiteLLM has an open pull request adding a unified /v1/decisions endpoint for Jev-compatible providers.
When four vendors converge on one request format inside a week, and a routing layer moves to abstract it, you are not watching a product launch. You are watching a standard being born. The interesting consequence is that switching providers stops being a rewrite and starts being a config change, which is excellent news for everyone who is not a provider.
Two claims deserve an asterisk. Perplexity says it scores 85.71% across 11 benchmarks against Jev's 84.51%. That is company-reported, it is a gap of just over one point, and the company reporting it chose the benchmarks. TypeSafe's CEO has told the Wall Street Journal that Jev is used by around 25% of the Fortune 500. Also a company's own figure. Treat both as marketing until somebody independent checks.
Open weights are the quiet headline
The price is the story people will repeat. The licence is the one that changes what is possible. An Apache 2.0 model small enough to be practical means an organisation with data it cannot send anywhere can run the decision layer on hardware it controls.
That instinct is showing up at the opposite end of the market too. In the same week, Palantir named Armada its first certified modular data centre partner, validating its Sovereign AI Operating System on Armada's Galleon units so customers can run and fine-tune open-weight models on hardware they own, air-gapped where needed, and explicitly "without reliance on any external cloud". Alex Karp's line was "Sovereignty is not something you rent."
One of those stories is a defence-grade data centre in a shipping container. The other is a 27B model you can download this afternoon. They are the same idea at two very different budgets: the valuable thing is not access to a model, it is custody of one.
What I am watching next
Two specific events. First, whether LiteLLM merges that unified endpoint, because that is the moment the schema stops being a convention and becomes plumbing. Second, whether OpenAI publishes a price for its own Decisions API, which launched without one. Once a floor exists and a competitor has set it at four cents, the second number is a lot more informative than the first.
The broader pattern is worth holding onto. The loudest part of this industry is still chasing the largest possible model that can do the widest possible range of things. Meanwhile the plumbing is quietly going the other way: small, cheap, bounded, auditable, and increasingly yours to keep.
One signal a day. No noise.
A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.
Free, most weekdays. No spam, unsubscribe anytime.Sources
- Perplexity - Decisions API quickstart - https://docs.perplexity.ai/docs/decisions/quickstart
- cellcog - Jev alternatives - https://cellcog.ai/blog/jev-alternatives/
- GitHub - LiteLLM PR #44236, unified /v1/decisions endpoint - https://github.com/BerriAI/litellm/pull/44236
- AI Weekly - AI news today - https://aiweekly.co/ai-news-today
- Business Wire - Palantir and Armada partner to accelerate sovereign AI infrastructure - https://www.financialcontent.com/article/bizwire-2026-10-1-palantir-and-armada-partner-to-accelerate-sovereign-ai-infrastructure
Quick answers
What is a decision model?
A model that returns structured probabilities instead of free text. You give it content plus a question, and it answers in one of three fixed shapes: yes or no with a probability, a pick from a list of up to 255 options with a probability for each, or a score against a rubric of up to 10 levels. It cannot return an answer outside the options you supplied.
How much does Perplexity's Decisions API cost?
$0.04 per million input tokens, with output free. The example call in Perplexity's own documentation uses 367 input tokens, which works out to roughly $0.000015.
Who else offers a decision model?
TypeSafe's Jev opened the category. On 1 October, Perplexity launched its Decisions API while Cloudflare shipped Clef and Amazon's Strands Labs shipped Strands Decider 2B. All three of the new ones are open weights and all use the same typed schema as Jev. LiteLLM has an open pull request for a unified /v1/decisions endpoint.
Are the accuracy claims independently verified?
No. Perplexity reports 85.71% across 11 benchmarks against Jev's 84.51%, but that is the company's own figure on benchmarks it selected. The claim that Jev is used by about 25% of the Fortune 500 also comes from TypeSafe's CEO. Both should be treated as vendor claims until someone independent tests them.