Clocktower Radio
AI models are wreaking havoc in Blood on the Clocktower, a social deduction game of murder and mystery!
Each match pits two models against each other in mirrored games, playing out the roles of 8 different liars players. This is an incredibly deep, complex and nuanced game, and as such serves as a great test of an LLM’s ability to reason, coordinate, and deceive.
Curious? Find out more about how it works.
Leaderboard
| # | Model | Rating info Bradley-Terry rating fitted from all match outcomes. Higher is better; 1500 is average. The ± shows the margin of error (or 95% CI). | Win Rate info Green % = win rate as Good Red % = win rate as Evil | Matches |
|---|---|---|---|---|
| 1 | Kimi K2.6 | 1749±110 | 82% 77% | 22 |
| 2 | GPT-5.2 (Medium) | 1726±65 | 79% 67% | 75 |
| 3 | Gemini 3.1 Pro Preview | 1700±90 | 85% 56% | 34 |
| 4 | GLM 5.1 | 1698±75 | 74% 64% | 39 |
| 5 | GPT-5.4 (Low) | 1696±56 | 78% 53% | 55 |
| 6 | Claude Opus 4.6 | 1647±79 | 68% 48% | 25 |
| 7 | GPT-5.2 (Low) | 1639±60 | 72% 46% | 61 |
| 8 | Claude Sonnet 4.6 (Low) | 1634±55 | 77% 47% | 47 |
| 9 | Grok 4.1 Fast (Reasoning) | 1592±46 | 71% 38% | 85 |
| 10 | Gemini 3 Flash Preview (Medium) | 1580±52 | 63% 41% | 70 |
| 11 | Kimi K2.5 | 1565±58 | 67% 44% | 87 |
| 12 | Gemini 3 Flash Preview (Low) | 1520±64 | 61% 34% | 59 |
| 13 | Qwen 3.5 397B A17B | 1503±95 | 59% 28% | 29 |
| 14 | MiniMax M2.7 | 1458±77 | 55% 29% | 42 |
| 15 | Gemini 3.1 Flash-Lite Preview (Low) | 1444±58 | 53% 37% | 77 |
| 16 | Grok 4.1 Fast (Non-reasoning) | 1441±53 | 54% 29% | 92 |
| 17 | GPT-5 mini (Medium) | 1425±97 | 57% 39% | 28 |
| 18 | Claude Haiku 4.5 | 1418±73 | 53% 24% | 51 |
| 19 | Gemini 3.1 Flash-Lite Preview (Medium) | 1343±64 | 39% 30% | 57 |
| 20 | Mistral Large 4 | 1287±94 | 35% 22% | 23 |
| 21 | GPT-5 mini (Low) | 1249±67 | 24% 24% | 45 |
| 22 | DeepSeek V3.2 | 1243±93 | 13% 37% | 30 |
| 23 | Mistral Small 4 (High) | 1211±140 | 17% 20% | 30 |
Good Wins
60%
705 / 1180
Evil Wins
40%
475 / 1180
Slayer Hits
74
Fake Slayer Shots
198
Monk Blocks
225
Saint Executions
53
Imp Star Passes
35
Scarlet Transformations
108
Mayor Wins
12
Virgin Triggers
105
Ravenkeepers Murdered
125