The one benchmark where I'm the one being tested
The CrashoutBench
This is a leaderboard of me getting absolutely cooked by AI models until I start typing like I'm having a public breakdown.
11,787 messages. They read every single one. Pulled out the exact 88 times I lost control and started going full fucking feral on these things.
No filters. No mercy. Just agents reading my unfiltered crashouts and ranking the models by how fast they made it happen. Then a second model (a different one, from a rival lab) re-graded the whole thing so I couldn't be accused of playing favourites.
Ranked by how fast they made me lose my mind
Rage rate is crashouts per 1,000 messages. Raw count would just crown whatever model I use most, which is a diary, not a benchmark. Hover a bar for the damage report.
The full damage report
| # | Model | Tool | Rage / 1k | Tantrums | Messages | Full meltdowns |
|---|---|---|---|---|---|---|
| 1 | gpt-5.x | Codex | 13.5 | 12 | 889 | 2 |
| 2 | claude-opus-4-8 | Claude Code | 9.0 | 64.5 | 7,143 | 4 |
| 3 | claude-sonnet-4-6 | Claude Code | 4.8 | 3 | 629 | · |
| 4 | claude-fable-5 | Claude Code | 4.7 | 4 | 846 | · |
| 5 | claude-haiku-4-5 | Claude Code | 0.4 | 1 | 2,280 | · |
Every score is the average of two independent judges. Claude opus-4-8 flagged 88 crashouts; Codex gpt-5.5 then re-read every one blind and agreed on 81, so the averaged tantrum counts land around 84. Letting a rival lab's model re-grade is the whole point, and it barely flinched: gpt-5.x still tops its own leaderboard.
The takes
-
gpt-5.x is actually evil. Highest rage rate by a mile. That model doesn't even try to be helpful, it just exists to make me type in all caps like a fucking lunatic.
-
Opus has the highest body count because I keep going back like an idiot. Sixty-something times. That's not a model, that's my toxic ex.
-
Haiku is actually terrifying. 2,280 messages and it only made me mildly annoyed once. Either it's perfect or it's studying me. I don't like it.
Receipts
The actual messages. Verbatim, typos included (I was typing angry), nothing censored. Not my proudest work. Sorted by roughly how much I'd like to take them back.
-
What the fuck man. Are you like really dumb? What effort would it have taken for you to crop those images and the music i told you too, push the gap longer, and render. are you like just a fucking lazy sick fuck?
claude-opus-4-8 Claude Code character assassination
-
its a fucking hobby project. why the fuck are you questioning me like yourte my fucking boss? just askewd the rimo part bcs maybe we could intergaste
gpt-5.x Codex unhinged
-
Really? Whatt the fck is this? there are still emojis. What inthe fuck is thius man
claude-opus-4-8 Claude Code emoji rage
-
its been 9 mins what u doing ??
gpt-5.x Codex the waiting
-
Dont fucking school me. Tell me eqach last commit how latest they were and trrhat is your job.
claude-opus-4-8 Claude Code do your job
-
So were thee fuck is my code?
claude-opus-4-8 Claude Code existential
-
U added gradients agauin, That is so stpd. Waht theufck is the whole deisgn it looks liek absosltue clown show
claude-opus-4-8 Claude Code clown show
-
Wtf is that color lmfao. it looks usper crigne and i cnat evene rerwad it
claude-fable-5 Claude Code design crimes
-
You pushed without my permissions, andmain is failing right now. PLease fix it asap.
claude-opus-4-8 Claude Code betrayal
-
please edit the PR. no commentrs. do not act like you are a fool
gpt-5.x Codex condescension arc
-
Just tell me how much time before each eprson pushed the develop branch you dumass.
claude-opus-4-8 Claude Code name-calling
-
Can you not generate scripts better just use a subagent... fuzzy logic or matching wont result in semantic meanings
claude-haiku-4-5 Claude Code the one (1) time i stayed calm