# The CrashoutBench

A leaderboard of me getting absolutely cooked by AI models until I start typing like I'm having a public breakdown. 11,787 messages read, the 88 genuine crashouts pulled out, models ranked by how fast they made it happen.

The visual chart on the HTML page is hard for agents. Use this table instead.

| Messages read | Crashouts flagged | Tools | Regrets |
| ---: | ---: | ---: | ---: |
| 11,787 | 88 | 3 | 0 |

Rage rate is crashouts per 1,000 messages. Raw count would just crown whatever model I use most.

## Leaderboard

| Rank | Model | Tool | Rage / 1k | Tantrums | Messages | Full meltdowns |
| ---: | --- | --- | ---: | ---: | ---: | ---: |
| 1 | gpt-5.x | Codex | 13.5 | 12 | 889 | 2 |
| 2 | claude-opus-4-8 | Claude Code | 9.0 | 64.5 | 7,143 | 4 |
| 3 | claude-sonnet-4-6 | Claude Code | 4.8 | 3 | 629 | — |
| 4 | claude-fable-5 | Claude Code | 4.7 | 4 | 846 | — |
| 5 | claude-haiku-4-5 | Claude Code | 0.4 | 1 | 2,280 | — |

Every score is the average of two independent judges. Claude opus-4-8 flagged 88 crashouts; Codex gpt-5.5 then re-read every one blind and agreed on 81, so the averaged tantrum counts land around 84. Letting a rival lab's model re-grade is the whole point, and it barely flinched: gpt-5.x still tops its own leaderboard.

## Takes

- **gpt-5.x is actually evil.** Highest rage rate by a mile. That model doesn't even try to be helpful, it just exists to make me type in all caps like a fucking lunatic.
- **Opus has the highest body count** because I keep going back like an idiot. Sixty-something times. That's not a model, that's my toxic ex.
- **Haiku is actually terrifying.** 2,280 messages and it only made me mildly annoyed once. Either it's perfect or it's studying me. I don't like it.

## Receipts

The actual messages. Verbatim, typos included, nothing censored.

### character assassination

> What the fuck man. Are you like really dumb? What effort would it have taken for you to crop those images and the music i told you too, push the gap longer, and render. are you like just a fucking lazy sick fuck?

— claude-opus-4-8 · Claude Code

### unhinged

> its a fucking hobby project. why the fuck are you questioning me like yourte my fucking boss? just askewd the rimo part bcs maybe we could intergaste

— gpt-5.x · Codex

### emoji rage

> Really? Whatt the fck is this? there are still emojis. What inthe fuck is thius man

— claude-opus-4-8 · Claude Code

### the waiting

> its been 9 mins what u doing ??

— gpt-5.x · Codex

### do your job

> Dont fucking school me. Tell me eqach last commit how latest they were and trrhat is your job.

— claude-opus-4-8 · Claude Code

### existential

> So were thee fuck is my code?

— claude-opus-4-8 · Claude Code

### clown show

> U added gradients agauin, That is so stpd. Waht theufck is the whole deisgn it looks liek absosltue clown show

— claude-opus-4-8 · Claude Code

### design crimes

> Wtf is that color lmfao. it looks usper crigne and i cnat evene rerwad it

— claude-fable-5 · Claude Code

### betrayal

> You pushed without my permissions, andmain is failing right now. PLease fix it asap.

— claude-opus-4-8 · Claude Code

### condescension arc

> please edit the PR. no commentrs. do not act like you are a fool

— gpt-5.x · Codex

### name-calling

> Just tell me how much time before each eprson pushed the develop branch you dumass.

— claude-opus-4-8 · Claude Code

### the one (1) time i stayed calm

> Can you not generate scripts better just use a subagent... fuzzy logic or matching wont result in semantic meanings

— claude-haiku-4-5 · Claude Code

## Method

Every JSONL transcript under Claude Code and Codex, plus Cursor's chat sessions, was parsed for messages I actually typed, each tagged with the model that answered. Twelve agents read all 11,787 of them and flagged genuine crashouts: real anger, swearing out of frustration, caps-lock shouting, or being fully done with the agent. Never a keyword hit. Then a second judge (Codex gpt-5.5) re-graded every flagged message blind, and the two scores were averaged per model. Ranked by crashouts per 1,000 messages, categorised by model across all three tools. Cursor contributed one usable message and zero crashouts, so, respect.
