Livenerf: A Pre-Registered Benchmark Built to Catch Silent Model Downgrades
A 780-point Hacker News project called Livenerf runs a daily 78-question panel against Opus 5.5 through Claude Code to detect quiet post-launch degradation, with pre-registered decision rules, a control arm, and a first verdict expected around October 24.
Every frontier-model launch is followed by the same rumor: it was quietly nerfed days later. What has been missing is evidence either way, because catching a silent downgrade requires exactly the kind of unglamorous infrastructure nobody builds. A project called Livenerf, which hit 780 points on Hacker News on September 30, is building it: an independent, long-running benchmark that tests Claude Opus 5.5 every day through headless Claude Code, designed to the standard where a “nerf” claim could actually survive scrutiny. The honest headline, per its own README: six of thirty planned days of data are collected as of September 29, no degradation has been measured, and the first statistically defensible verdict arrives around October 24.
The Methodology Is the Story
Livenerf’s design decisions read as a masterclass in measuring a moving target. The panel is 78 questions deliberately selected to be “sometimes right”: the author calibrated 2,336 questions from GPQA Diamond, MMLU-Pro, competition math, and AIME 2025-26 with four samples each, and kept only the slice where Opus 5.5 was inconsistent, because questions a model always answers carry no signal about degradation. Determinism is achieved by freezing everything that can be frozen: a hash-pinned Claude Code CLI (2.1.280), fixed prompts, hermetic single-turn calls, and pure-function graders with no LLM judges. The comparison is paired (each question against its own launch-week baseline) with clustered standard errors, and a control arm runs Claude Opus 5 daily on the same questions to distinguish model changes from harness or platform changes. Most importantly, the decision rule was pre-registered before data collection: a verdict of “nerfed” requires 99% confidence intervals excluding zero in two consecutive 10-day windows, an effect of at least 3 points, and no matching movement in the control arm.
What It Can and Cannot Catch
The calibration runs produced the study’s most practical finding: output token counts are the leading indicator of quiet degradation. Positive controls showed that lowering effort cut tokens by 62% (low) and 26% (medium) with accuracy drops of 8.3 and 4.2 points respectively, so a silent capability reduction would likely show up in token spend before accuracy. Sensitivity is quantified: the daily panel detects roughly 7.5-point accuracy changes per 10-day window at about 3.6% of a weekly Claude Max plan budget. The limitations are stated with equal care: Livenerf measures the model as served through Claude Code on a subscription, not the raw API; it cannot reliably detect same-family model swaps of the validated size; and the author retained eight apparently wrong answer keys among the 78, with a pre-registered sensitivity analysis rather than silent deletion.
Why This Matters Beyond Opus 5.5
The nerfing discourse has been unfalsifiable for years: users report degradation, labs deny changes or blame infrastructure, and past incidents (including 2025’s wave) resolved as infrastructure bugs rather than deliberate downgrades. Livenerf converts the argument from vibes to statistics, and the playbook is generalizable: the same paired-baseline, control-arm, pre-registration design works for any served model on any subscription tier. It also lands at a specific moment: Opus 5.5 launched September 22 with benchmark crowns and Sonnet 5.5 followed matching it at half price, which makes the question of whether launch-week quality persists a live commercial issue, not just a forum grievance. A verified no-nerf result would be almost as valuable to Anthropic as a verified nerf result would be damaging, which is precisely why an independent instrument with pre-registered rules is worth existing.
What to Watch
The first verdict window closes around October 24, and either outcome is news: a confirmed post-launch change would be the first statistically clean documentation of the practice everyone suspects, and a clean bill would set a precedent labs could cite. Watch also for replication, because the design is published in enough detail (Inspect framework, pinned CLI, panel construction) that other models and other subscriptions can be instrumented the same way, and a fleet of Livenerf-style monitors across frontier models would quietly become the industry’s real quality ledger.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How to Use Claude Across Excel and PowerPoint with Shared Context and Skills
Learn how to use Claude's shared Excel and PowerPoint context, Skills, and enterprise gateways for faster analyst workflows.
Claude Sonnet 5.5 Matches Opus-Class Workloads at Half the Price
Anthropic released Claude Sonnet 5.5 on September 28: Terminal-Bench 70.6%, GDPval 1844 against Opus 5.5's 1846, at $2/$10 per million tokens, with first-for-Sonnet cyber safeguards and a fallback to Sonnet 5 when tasks get risky.
Claude Opus 5.5 Matches Fable 5.1 While Cutting Running Costs 40%
Anthropic released Claude Opus 5.5 on September 22, a flagship that matches or beats Fable 5.1 on most benchmarks while costing 40% less to run than Opus 5, as the frontier price war that started with Grok 4.7 claims its biggest scalp.
GPT-6.1 Sol Arrives With Near-Astra Performance at a Fifth of the Price
Announced at DevDay on September 29, GPT-6.1 Sol claims near-GPT-6 Astra performance at $2/$10 per million tokens, with a 1.05 million token context window and a new Ultrafast tier at 300 tokens per second, one week after GPT-6 Sol launched.
Claude Agents Discovered a Novel CRISPR-like Enzyme System
Anthropic announced on September 23 that roughly 950 Claude agents, running 21 hours on 210 million tokens, found a previously uncharacterized bacteriophage enzyme system with CRISPR-like repeat arrays, now partly validated in the lab and published as a pre-print.