TL;DR
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Livenerf, an open benchmark tracking Claude Opus 5.5 after launch, had collected six of 30 planned daily runs as of Sept. 29, 2026. It has not yet published a result showing whether the model’s performance changed; the first comparison is expected after day 20, around Oct. 24.
As of Sept. 29, Livenerf had collected six of 30 planned daily benchmark runs and had not published a comparison of Claude Opus 5.5 with its launch-week performance. The GitHub project is tracking the model after its Sept. 22 release to measure changes in accuracy or output length over time.
The benchmark’s first 10 days form its launch-week baseline, followed by two 10-day measurement windows. Livenerf says the series began Sept. 24, about two and a half days after release. Its status page reported that all six runs completed without a missed day, each using the same harness hash and pinned Claude Code CLI version, 2.1.280. One run used an override to the project’s budget guard, which is recorded in its deviations log.
Livenerf tests a fixed panel of 78 questions selected from 2,336 questions spanning GPQA Diamond, MMLU-Pro, competition mathematics and AIME 2025–26. The project says Opus 5.5 answered about 93% of the original pool correctly on the first try; the panel comprises questions it answered inconsistently across repeated samples. The published design says fresh samples raised the panel’s pass rate from 54.7% to 62.0%, a selection effect the project accounts for in its power calculation.
The planned comparison measures paired score differences against baseline, with clustered standard errors. Livenerf also tracks output tokens because its authors say a change in how much a model “thinks” may show up in token counts before accuracy shifts. The project estimates that its current schedule can detect an accuracy change of about 7.5 percentage points per 10-day window. That is an estimate of the instrument’s sensitivity, not an observed change in Opus 5.5.
A Baseline for Later Comparisons
Claims that a model gets worse after release can be assessed by comparing measurements taken under similar conditions at launch and later dates. Livenerf has published a pre-registered comparison method and retained run logs. It says planned results will report both improvements and regressions.
The benchmark may provide information about performance on its selected question panel through a particular Claude Code setup. It does not establish whether every user, task or serving route sees the same behavior. A measured shift would describe results under these test conditions; the benchmark alone would not identify its cause.
Top picks for "livenerf opus nerf"
As an affiliate, we earn on qualifying purchases.
How the Test Tracks Drift
The project describes itself as a small, append-only benchmark intended to examine reports that Anthropic models may change after launch. It lists possible explanations for perceived differences, including quantization, a smaller model behind the same name, lower effort or routing changes. It also recognizes that users may perceive a change where none has been measured. Those possibilities are hypotheses in the project’s framing, not findings about Opus 5.5.
Livenerf says it uses frozen prompts, a pinned command-line interface, fixed graders and stored raw logs to reduce variation in its own procedure. It runs through a Claude Max subscription using headless Claude Code rather than an API key, and it is built on the UK AI Security Institute’s open-source Inspect evaluation framework. The project says its statistical methods follow Anthropic’s published guidance, “Adding Error Bars to Evals.”
The project also reports limitations in the question set. A report-only audit identified eight answer keys that appeared wrong and 30 ambiguous questions among the selected and later excluded items. Livenerf says it retained the affected panel questions and pre-registered a sensitivity analysis that would rerun results without them. It also says a safety classifier sometimes routes answers through Opus 5 or refuses certain biology and mathematics questions; affected samples are rejected and counted, and touched questions are excluded.
What Six Runs Cannot Show
No performance trend has been reported yet. Six days make up only part of the 10-day baseline, and the project says its first results row will appear after day 20. The first possible call on whether performance moved is expected around Oct. 24, assuming the planned schedule continues.
The benchmark’s own validation also describes limits. In one validation, swapping Opus 5 for Opus 5.5 was not distinguishable at the 99% level, with a reported difference of −3.8 ± 6.3 points and 23% fewer tokens. Livenerf says it has not shown that a 10-day window, despite having about 2.5 times as many samples as that validation, can detect a same-family model swap of that size. The source does not establish whether any change in benchmark performance would result from a model update, routing, effort or another cause.
The October Comparison Window
Livenerf plans to run once daily for 30 days: 10 baseline days, then two 10-day comparison windows. The project says its first results row is due after day 20, around Oct. 24 based on the schedule. That will provide the first planned comparison, while the remaining days can extend the series.
Readers can check the repository’s results table, calibration notes, design, validation and deviations log as they are updated. Until the baseline and later window are available, the available status information shows that data collection is underway; whether Opus 5.5 has changed remains unanswered.
Key Questions
Has Livenerf found that Opus 5.5 was nerfed?
No finding has been reported. As of Sept. 29, 2026, only six of the planned 30 daily runs had been collected, and the baseline comparison was not yet available.
When will the first comparison be available?
The project says its first results row will appear after day 20. Based on its schedule, the first possible comparison is around Oct. 24, 2026.
What does Livenerf measure?
It measures accuracy on a fixed 78-question panel and tracks output tokens, comparing later runs with a launch-week baseline. The project estimates it can detect an accuracy change of about 7.5 percentage points per 10-day window.
Would a score decline prove Anthropic changed the model?
No. A measured difference on Livenerf’s panel would show a change under its test conditions. The project’s method does not by itself identify whether model weights, routing, effort or another factor caused that difference.
Source: hn
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
