Cloudflare Enters the Decision-Model War With Open-Weights Clef
Cloudflare announced Clef and Clef-flash on October 1, open-weight decision models that beat TypeSafe's Jev on classification benchmarks at lower latency, Jev-API-compatible, and fine-tunable through a new RL platform.
The decision-model market that TypeSafe’s Jev created three weeks ago has its first serious open-weights challenger. Cloudflare announced Clef and Clef-flash on October 1, its first Cloudflare-trained models: classifiers that return bounded, structured outputs with probabilities instead of generating text, exactly Jev’s “System One” niche. Clef is built on a frozen Qwen3.8-27B backbone, Clef-flash on Qwen3.5-9B, both open-sourced on Hugging Face under Apache 2.0 and served on Workers AI, and both are Jev-API-compatible, which means swapping them into an existing Jev integration is a base-URL change.
The Benchmarks Are a Direct Hit on Jev
Cloudflare did not publish vague comparisons. Against Jev head to head: BFCL case-exact 98.47 (Clef) and 98.76 (Clef-flash) versus Jev’s 95.75; BANKING77 macro-F1 94.20 versus 79.74; median latency 209.3 ms for Clef and 38.8 ms for Clef-flash versus Jev’s 524.1 ms. On Typesafe’s own evaluation suite, Clef won three of four workflows, including invoice processing at 64.7 versus 61.8. Two capability differences do the heavy lifting: a 64k context window against Jev’s 32k, and a vision encoder, since Jev is text-only and Clef classifies images natively. The architecture trick worth copying is the scoring pass: Clef runs a prefill-only pass and then scores the output schema in parallel, non-autoregressively, which is where the latency win comes from. Check Point’s study last week showed that this entire category breaks under prompt injection at pocket-change cost; Clef inherits that exposure, and nothing in Cloudflare’s announcement claims otherwise.
The RL Fine-Tuning Platform Is the Strategic Play
The models are the visible product; the fine-tuning service is the business. Cloudflare is launching a reinforcement-learning tuning platform that lets customers specialize Clef on their own labeled data, initially through forward-deployed engineers with self-serve later, and the stack is assembled from pieces that already existed: AI Gateway captures production traffic as training data, Workers AI handles rollouts and BYO-model redeployment, Cloudflare Containers provide RL sandboxes, and a new Trainer component runs the optimization (label-smoothed cross-entropy, Brier loss for calibration, and an RLCD method for calibrated decisions). Internal use cases, Trust & Safety triage, support ticketing, bot classification, are the exact workloads enterprises are currently paying LLM prices for. This is Cloudflare applying its standard playbook, open the weights, sell the platform, to a category that did not exist in August.
What the Open-Weights Decision Changes
The Apache 2.0 release matters more than the benchmarks, because it answers the question Jev’s research-only posture left open: can the industry audit and self-host the models making automated decisions about it? Clef’s weights can be downloaded, fine-tuned, inspected for backdoors, and run on your own GPUs, which is a materially different trust proposition for anything touching finance, legal, or safety triage. The competitive effect is also immediate: Jev’s differentiation is now speed-to-market plus the Typesafe brand, not capability (Clef wins most head-to-heads) and not openness (Clef wins outright). Typesafe’s response, a licensing change, a capability bump, or a price move, will show whether the first mover can hold a category against a platform company distribution play.
What to Watch
Three things. First, independent replication of Cloudflare’s numbers, which come from Cloudflare and were run against named competitors; the BANKING77 gap in particular is large enough that verification matters. Second, pricing for the hosted tiers and the fine-tuning platform, which was conspicuously absent from the announcement. Third, the injection-study question applied to Clef: the Jev findings were category-wide, and Clef’s docs will be read closely for whether Cloudflare’s version ships with input-screening guidance or repeats the “structured output is a boundary” mistake.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Fine-Tune Cosmos Predict 2.5 for Robotics With LoRA
Learn how to adapt NVIDIA's 2B and 14B Cosmos Predict 2.5 world foundation models using parameter-efficient fine-tuning methods like LoRA and DoRA.
Check Point Broke the 'Never Hallucinates' Decision Model for 50 Cents a Break
Check Point Research published a systematic prompt-injection study of Jev, TypeSafe AI's typed decision model, on September 24: every attack configuration broke at least once, the strongest broke 25 of 27 runs at about $0.50 per successful attack, and structured output and anti-injection instructions both failed to help.
Xiaomi's MiMo v2.6 Takes the Top Open-Weights Spot on the Intelligence Index
Xiaomi released MiMo v2.6 on September 21, a 1-trillion-parameter open-weights MoE model that debuts at number one among open models on Artificial Analysis' Intelligence Index, with an MIT license and aggressive API pricing.
Gemini 4 Argon Lands With 1M-Token Output and a Safety-Gated Rollout
Google announced Gemini 4 Argon on September 30 with a 1 million token output limit, a DeepSWE state of the art at 77.9%, and defensive-cyber focus, but most users cannot touch it yet: rollout runs through a trusted-defenders program first.
Livenerf: A Pre-Registered Benchmark Built to Catch Silent Model Downgrades
A 780-point Hacker News project called Livenerf runs a daily 78-question panel against Opus 5.5 through Claude Code to detect quiet post-launch degradation, with pre-registered decision rules, a control arm, and a first verdict expected around October 24.