A new benchmark wants to catch AI models getting quietly worse after launch

Livenerf promises hard data on whether Anthropic's models degrade post-release - but even its own maintainer admits nobody actually knows yet.

A computer circuit board with a brain on it
Photo · Steve A Johnson / Unsplash

The claim

For months, a persistent rumour has circulated among heavy users of Anthropic’s Claude models: that after a flashy launch, the model quietly gets worse. Slower, dumber, less willing to follow instructions properly. Anthropic has never confirmed this is deliberate policy, and the people making the claim rarely have anything more rigorous than “it felt different on Tuesday” to back it up.

Enter livenerf, a new open-source project doing the rounds on Hacker News with 465 points and nearly 200 comments at time of writing. Its pitch is refreshingly blunt: it’s a “small, boring, append-only benchmark” built to answer one question - does a model actually get worse after it ships, or is everyone just pattern-matching on noise?

What livenerf actually does

According to its own README, livenerf is designed to be as deterministic as possible, repeatedly testing a model against the same tasks over time and logging the results. The idea is to build a public, ongoing record rather than relying on anecdote. Crucially, the project’s authors are careful not to claim they’ve proven anything yet - they explicitly frame the question as open, listing quantisation, swapping in a smaller model under the same name, reduced “effort” settings, or backend routing changes as plausible explanations if degradation does turn up. They’re equally clear that “nothing happened” is still on the table.

That caution is worth taking seriously, because livenerf isn’t the only outfit trying to measure this. Commenters on Hacker News pointed to at least two existing trackers doing similar work: Nerf Bench, which benchmarks models against their day-one launch performance and flags any drift over 10% as a genuine change, and a separate Claude Code tracker that’s reportedly been running for a while. Nerf Bench’s operators claim it previously caught a real degradation in an earlier Claude Opus release that Anthropic later addressed in a blog post - though that’s a claim repeated in comments, not something independently verified here.

So is Opus 5.5 actually nerfed?

Not proven either way. The headline question - “has Opus 5.5 been nerfed yet?” - is the benchmark’s stated aim, not its answer. Livenerf and its rivals are currently tracking the model, but no conclusive, published degradation has been demonstrated in the source material. What’s notable is how split even sceptical, technical commenters are: several insist they’ve never seen the paid API degrade, only the subscription-based coding tools, with one describing effectively daily fluctuations that vanish at weekends. Others think the whole phenomenon is a “honeymoon effect” - novelty wearing off, not the model itself changing. One commenter’s own logging suggested their usage quota consumption varied significantly between sessions for reasons that had nothing to do with model quality, which is a reminder that quota and cost behaviour is a separate issue from capability.

What this means for you

If you’re a casual ChatGPT or Claude user, nothing changes today - there’s no confirmed, systemic nerfing to react to, and switching providers on the strength of forum vibes isn’t advised. If you’re a developer relying heavily on Claude for coding work, projects like livenerf are worth bookmarking as an early-warning system rather than proof of anything. The honest state of play: plenty of suspicion, a genuinely interesting attempt to measure it properly, and no verified smoking gun yet.

Sources