The one benchmark where I'm the one being tested

The CrashoutBench

This is a leaderboard of me getting absolutely cooked by AI models until I start typing like I'm having a public breakdown.

11,787 messages. They read every single one. Pulled out the exact 88 times I lost control and started going full fucking feral on these things.

No filters. No mercy. Just agents reading my unfiltered crashouts and ranking the models by how fast they made it happen. Then a second model (a different one, from a rival lab) re-graded the whole thing so I couldn't be accused of playing favourites.

Messages read
11,787
Times I crashed out
88
Tools involved
3
Regrets
0

Ranked by how fast they made me lose my mind

Rage rate is crashouts per 1,000 messages. Raw count would just crown whatever model I use most, which is a diary, not a benchmark. Hover a bar for the damage report.

The full damage report

# Model Tool Rage / 1k Tantrums Messages Full meltdowns
1 gpt-5.x Codex 13.5 12 889 2
2 claude-opus-4-8 Claude Code 9.0 64.5 7,143 4
3 claude-sonnet-4-6 Claude Code 4.8 3 629 ·
4 claude-fable-5 Claude Code 4.7 4 846 ·
5 claude-haiku-4-5 Claude Code 0.4 1 2,280 ·

Every score is the average of two independent judges. Claude opus-4-8 flagged 88 crashouts; Codex gpt-5.5 then re-read every one blind and agreed on 81, so the averaged tantrum counts land around 84. Letting a rival lab's model re-grade is the whole point, and it barely flinched: gpt-5.x still tops its own leaderboard.

The takes

  • gpt-5.x is actually evil. Highest rage rate by a mile. That model doesn't even try to be helpful, it just exists to make me type in all caps like a fucking lunatic.

  • Opus has the highest body count because I keep going back like an idiot. Sixty-something times. That's not a model, that's my toxic ex.

  • Haiku is actually terrifying. 2,280 messages and it only made me mildly annoyed once. Either it's perfect or it's studying me. I don't like it.

Receipts

The actual messages. Verbatim, typos included (I was typing angry), nothing censored. Not my proudest work. Sorted by roughly how much I'd like to take them back.

  • What the fuck man. Are you like really dumb? What effort would it have taken for you to crop those images and the music i told you too, push the gap longer, and render. are you like just a fucking lazy sick fuck?

    claude-opus-4-8 Claude Code character assassination

  • its a fucking hobby project. why the fuck are you questioning me like yourte my fucking boss? just askewd the rimo part bcs maybe we could intergaste

    gpt-5.x Codex unhinged

  • Really? Whatt the fck is this? there are still emojis. What inthe fuck is thius man

    claude-opus-4-8 Claude Code emoji rage

  • its been 9 mins what u doing ??

    gpt-5.x Codex the waiting

  • Dont fucking school me. Tell me eqach last commit how latest they were and trrhat is your job.

    claude-opus-4-8 Claude Code do your job

  • So were thee fuck is my code?

    claude-opus-4-8 Claude Code existential

  • U added gradients agauin, That is so stpd. Waht theufck is the whole deisgn it looks liek absosltue clown show

    claude-opus-4-8 Claude Code clown show

  • Wtf is that color lmfao. it looks usper crigne and i cnat evene rerwad it

    claude-fable-5 Claude Code design crimes

  • You pushed without my permissions, andmain is failing right now. PLease fix it asap.

    claude-opus-4-8 Claude Code betrayal

  • please edit the PR. no commentrs. do not act like you are a fool

    gpt-5.x Codex condescension arc

  • Just tell me how much time before each eprson pushed the develop branch you dumass.

    claude-opus-4-8 Claude Code name-calling

  • Can you not generate scripts better just use a subagent... fuzzy logic or matching wont result in semantic meanings

    claude-haiku-4-5 Claude Code the one (1) time i stayed calm

How this was made (for the nerds)

Every JSONL transcript under Claude Code and Codex, plus Cursor's chat sessions, was parsed for messages I actually typed, each tagged with the model that answered. Twelve agents read all 11,787 of them and flagged genuine crashouts: real anger, swearing out of frustration, caps-lock shouting, or being fully done with the agent. Never a keyword hit. Then a second judge (Codex gpt-5.5) re-graded every flagged message blind, and the two scores were averaged per model. Ranked by crashouts per 1,000 messages, categorised by model across all three tools. Cursor contributed one usable message and zero crashouts, so, respect.