Home › Blog › Decision models got a price: $0.04 a million
AI Infrastructure

Decision models got a price: $0.04 a million

Perplexity shipped a Decisions API that returns probabilities instead of prose, priced at $0.04 per million input tokens with free output, and open-sourced the weights. Three rivals now speak the same schema.

One signal a day. No noise. A 3-minute read when something genuinely shifts.
By Tyron Dizon · October 3, 2026 · 5 min read
Perplexity shipped a Decisions API that returns probabilities instead of prose, priced at $0.04 per million input tokens with free output, and open-sourced the weights. Three rivals now speak the same schema.
Source: Perplexity Decisions API documentation (launched 1 October 2026).

Most of what we call "using AI" is really asking a very articulate stranger a question and then praying the answer comes back in a shape our software can read. We write prompts begging for JSON. We bolt on parsers. We add retries for when the model decides today is the day it opens with "Certainly! Here's the classification you requested."

On 1 October, Perplexity put a price on the alternative. Their new Decisions API runs on pplx-decider-v1-27b, an Apache 2.0 fine-tune of Qwen3.8-27B, and it does not write sentences at all. It returns probabilities. Input costs $0.04 per million tokens. Output is free. The weights are published, so you can also just take the model and run it yourself.

The vending machine, not the conversation

Here is the difference in one image. A general chat model is a brilliant colleague you can ask anything, which also means they can answer anything, including things you did not ask for. A decision model is a vending machine. There are buttons. You press one, an item drops. Nothing else is on the menu.

The API takes three shapes of question. A yes/no question returns the probability of yes, between 0 and 1. A multiple-choice question lets you list up to 255 options and returns a probability for each one, plus the top pick and a confidence. A score question lets you define a rubric of up to 10 levels and returns a probability per level and an expected score.

That constraint is the whole product. A model that can only choose from the list you gave it cannot invent a category that does not exist in your database. It cannot return "Refund_Request" when your system only knows "refund". And because you get a number rather than a verdict, you can set your own bar: route automatically above 0.9, send to a human below it.

A model that can only answer the question you actually asked is not a weaker model. It is a smaller surface for things to go wrong.

The numbers are small enough to stop thinking about

Perplexity's own documentation example uses 367 input tokens. At their price, that call costs roughly $0.000015. Fifteen millionths of a dollar. You would need to run about 67,000 of those before you spent a dollar.

The limits are built for volume rather than eloquence:

A whole category formed in about a day

This is the part that made me sit up. Perplexity was not alone. On the same day, Cloudflare shipped Clef and Amazon's Strands Labs shipped Strands Decider 2B. Both are open weights. All three use the same typed schema as TypeSafe's Jev, the product that opened the category. LiteLLM has an open pull request adding a unified /v1/decisions endpoint for Jev-compatible providers.

When four vendors converge on one request format inside a week, and a routing layer moves to abstract it, you are not watching a product launch. You are watching a standard being born. The interesting consequence is that switching providers stops being a rewrite and starts being a config change, which is excellent news for everyone who is not a provider.

Two claims deserve an asterisk. Perplexity says it scores 85.71% across 11 benchmarks against Jev's 84.51%. That is company-reported, it is a gap of just over one point, and the company reporting it chose the benchmarks. TypeSafe's CEO has told the Wall Street Journal that Jev is used by around 25% of the Fortune 500. Also a company's own figure. Treat both as marketing until somebody independent checks.

Open weights are the quiet headline

The price is the story people will repeat. The licence is the one that changes what is possible. An Apache 2.0 model small enough to be practical means an organisation with data it cannot send anywhere can run the decision layer on hardware it controls.

That instinct is showing up at the opposite end of the market too. In the same week, Palantir named Armada its first certified modular data centre partner, validating its Sovereign AI Operating System on Armada's Galleon units so customers can run and fine-tune open-weight models on hardware they own, air-gapped where needed, and explicitly "without reliance on any external cloud". Alex Karp's line was "Sovereignty is not something you rent."

One of those stories is a defence-grade data centre in a shipping container. The other is a 27B model you can download this afternoon. They are the same idea at two very different budgets: the valuable thing is not access to a model, it is custody of one.

What I am watching next

Two specific events. First, whether LiteLLM merges that unified endpoint, because that is the moment the schema stops being a convention and becomes plumbing. Second, whether OpenAI publishes a price for its own Decisions API, which launched without one. Once a floor exists and a competitor has set it at four cents, the second number is a lot more informative than the first.

The broader pattern is worth holding onto. The loudest part of this industry is still chasing the largest possible model that can do the widest possible range of things. Meanwhile the plumbing is quietly going the other way: small, cheap, bounded, auditable, and increasingly yours to keep.

The price floor for a decisionPerplexity Decisions API, pplx-decider-v1-27b, Apache 2.0, launched 1 Oct 2026INPUT$0.04per million tokensOUTPUT$0freeTHE DOCS EXAMPLE CALL$0.000015for 367 input tokens128questions per request262,144max input tokens10 / secrequests per organisationSource: Perplexity Decisions API documentation
Source: Perplexity Decisions API documentation (launched 1 October 2026).

One signal a day. No noise.

A 3-minute read when something genuinely shifts in AI, automation, or defense tech. Free, most weekdays.

Free, most weekdays. No spam, unsubscribe anytime.

Sources

  1. Perplexity - Decisions API quickstart - https://docs.perplexity.ai/docs/decisions/quickstart
  2. cellcog - Jev alternatives - https://cellcog.ai/blog/jev-alternatives/
  3. GitHub - LiteLLM PR #44236, unified /v1/decisions endpoint - https://github.com/BerriAI/litellm/pull/44236
  4. AI Weekly - AI news today - https://aiweekly.co/ai-news-today
  5. Business Wire - Palantir and Armada partner to accelerate sovereign AI infrastructure - https://www.financialcontent.com/article/bizwire-2026-10-1-palantir-and-armada-partner-to-accelerate-sovereign-ai-infrastructure

Quick answers

What is a decision model?

A model that returns structured probabilities instead of free text. You give it content plus a question, and it answers in one of three fixed shapes: yes or no with a probability, a pick from a list of up to 255 options with a probability for each, or a score against a rubric of up to 10 levels. It cannot return an answer outside the options you supplied.

How much does Perplexity's Decisions API cost?

$0.04 per million input tokens, with output free. The example call in Perplexity's own documentation uses 367 input tokens, which works out to roughly $0.000015.

Who else offers a decision model?

TypeSafe's Jev opened the category. On 1 October, Perplexity launched its Decisions API while Cloudflare shipped Clef and Amazon's Strands Labs shipped Strands Decider 2B. All three of the new ones are open weights and all use the same typed schema as Jev. LiteLLM has an open pull request for a unified /v1/decisions endpoint.

Are the accuracy claims independently verified?

No. Perplexity reports 85.71% across 11 benchmarks against Jev's 84.51%, but that is the company's own figure on benchmarks it selected. The claim that Jev is used by about 25% of the Fortune 500 also comes from TypeSafe's CEO. Both should be treated as vendor claims until someone independent tests them.

Tyron Dizon is a Chief Product Officer, AI product builder, and Techstars-backed SaaS founder based in Baguio City, Philippines. He previously co-founded and served as CPO of SanityDesk and now builds AI products, automation systems, SaaS platforms, and rapid prototypes. About · Work · Resume · LinkedIn