Ai Engineering 3 min read

Decision Models Go Institutional: OpenAI's Decisions API Enters Beta

OpenAI's Decisions API entered public beta on October 6, returning predicate, choice, and score answers from gpt-6-luna at $0.10 per million input tokens, one day after AWS's Strands shipped a 2B open-source decider, as the category Red Hat just benchmarked institutionalizes.

The decision-model category that Red Hat benchmarked into its actual niche yesterday just got its institutional layer. OpenAI’s Decisions API entered public beta on October 6: a POST /v1/decisions endpoint that evaluates text or images and returns typed answers, a probability that a condition is true (predicate), a pick from your supplied options with per-option probabilities (choice), or a score against ordered levels (score), running on gpt-6-luna at $0.10 per million input tokens with no output charges, and about 10 times faster than the Responses API. One day earlier, AWS’s Strands team shipped Strands Decider 2B, a small open-source decision model. The category has gone from one startup’s pitch to a platform feature on both major clouds in a week.

The API Design Is the Interesting Part

OpenAI’s schema decisions reveal what the company thinks the category is. Three question types only: predicate, choice, score, each returning calibrated probabilities rather than free text, with support for multiple independent questions against one input in a single request and an explicit refusal answer type. The pricing is the signal: input-only at $0.10 per million tokens with no output charges at all, which prices decisions as a classification utility rather than a generation product, roughly a fifth of GPT-6.1 Sol’s input price and directly competitive with Cloudflare’s Clef on the hosted frontier. The guardrails are enterprise-shaped: Zero Data Retention, HIPAA eligibility, and US and EU data residency, which are the requirements a compliance buyer needs before routing customer tickets through a probability API.

The Timing Answers Red Hat With Distribution

The irony of the sequencing is worth spelling out. Red Hat’s study found that Jev-style models “do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy,” and listed OpenAI’s Decisions API as untested future work. OpenAI shipped the API into public beta within days. The company’s bet is the one this month’s local-inference results imply from the other direction: if a decision is just a fast, cheap model call, then the winner is whoever puts decision endpoints inside the platform developers already use, not whoever trains a bespoke architecture. Red Hat’s own data partially supports that: Qwen3.6-35B as an LLM judge was the accuracy winner in their study, and gpt-6-luna judging is precisely that pattern, productized.

Strands Decider 2B Is the Self-Hosted Counterweight

AWS’s Strands Decider 2B, a 2-billion-parameter open-source decider in the same week, completes the market structure: a frontier-hosted option (OpenAI), an open-weights option (Cloudflare’s Clef), and now a small self-hostable option (Decider 2B) at the price point where per-call costs round to zero. Jev’s differentiation, zero-shot decisions without labeled data, survives as the remaining niche, but the market has clearly decided that decision models are a feature of model platforms rather than a standalone product category. That is the verdict Red Hat’s numbers pointed at, arrived at by distribution instead of benchmarks.

What to Watch

Three things. First, Red Hat’s promised follow-up benchmark of the Decisions API: it was explicitly on their future-work list, and a measured comparison of gpt-6-luna’s predicate accuracy against the 200M classifier will price the category’s accuracy-per-dollar. Second, whether Strands Decider 2B and Clef see real self-hosted adoption, which tests whether enterprises actually want decisions running inside their own perimeter or will accept the API for convenience. Third, Jev and TypeSafe’s next move: their zero-shot no-labeled-data niche is real but small, and the injection findings apply to every new entrant in the category too. The month’s arc is complete: decision models went from startup pitch to contested category to platform feature in four weeks.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading

Ai Engineering

How Function Calling Works in LLMs

Function calling lets LLMs interact with external systems by requesting structured tool executions. Here's how the loop works, how to define tools, and what to watch for across providers.

Ai Engineering

Red Hat Benchmarks Find Decision Models Don't Beat a 200M Classifier

A Red Hat study published October 2 benchmarked TypeSafe's Jev against nine guardrails: a 200-million-parameter pre-trained classifier matched it on accuracy at a fraction of the latency, and Jev did not reliably beat LLM-as-a-judge either.

Ai Engineering

OpenAI Publishes 722 AI-Generated Mathematics Manuscripts, From π to the Hodge Conjecture

OpenAI released 722 mathematical manuscripts in 372 families produced by an internal model, hosted on GitHub with Lean formalizations and reasoning traces, alongside an investigation clearing accusations that the model stole a researcher's results.

Ai Engineering

Simon Willison: Agents Need Default Hard Budget Caps on Everything

Simon Willison argued on October 3 that usage-based services should ship with default hard budget caps that cut off service when a monthly limit is hit, because AI agents make it trivially easy to deploy code that incurs real costs.

Ai Engineering

FTC Opens Investigation Into OpenAI, Anthropic Over AI Product Risks

The FTC opened a broad investigation on September 30 into OpenAI, Anthropic and other AI companies, using consumer-protection law to examine whether labs have misled the public about the dangers of their products.