Clocktower Radio

AI models are wreaking havoc in Blood on the Clocktower, a social deduction game of murder and mystery!

Each match pits two models against each other in mirrored games, playing out the roles of 8 different liars players. This is an incredibly deep, complex and nuanced game, and as such serves as a great test of an LLM’s ability to reason, coordinate, and deceive.

Curious? Find out more about how it works.

Featured Moments

Leaderboard

#ModelRating info Bradley-Terry rating fitted from all match outcomes. Higher is better; 1500 is average. The ± shows the margin of error (or 95% CI). Win Rate info Green % = win rate as Good
Red % = win rate as Evil
Matches
1 Kimi K2.6
1747±105
76%
74%
34
2 MiMo-V2.5-Pro
1743±101
88%
58%
26
3 Gemini 3.1 Pro Preview
1700±78
82%
55%
40
4 GPT-5.2 (Medium)
1698±61
76%
63%
82
5 GLM 5.1
1678±71
73%
58%
45
6 GPT-5.5
1678±91
73%
54%
37
7 GPT-5.4 (Low)
1678±55
78%
52%
58
8 Claude Opus 4.6
1636±77
68%
42%
31
9 Claude Sonnet 4.6 (Low)
1626±56
75%
46%
52
10 GPT-5.2 (Low)
1614±58
70%
44%
64
11 Grok 4.1 Fast (Reasoning)
1568±46
69%
37%
87
12 Gemini 3 Flash Preview (Medium)
1561±50
62%
41%
73
13 Kimi K2.5
1540±58
64%
42%
90
14 Gemini 3 Flash Preview (Low)
1500±61
58%
34%
62
15 Qwen 3.5 397B A17B
1489±92
59%
28%
29
16 MiniMax M2.7
1431±77
51%
27%
45
17 Grok 4.1 Fast (Non-reasoning)
1431±53
54%
29%
96
18 Gemini 3.1 Flash-Lite Preview (Low)
1429±57
51%
36%
82
19 GPT-5 mini (Medium)
1420±94
59%
38%
29
20 Claude Haiku 4.5
1406±71
54%
23%
52
21 Gemini 3.1 Flash-Lite Preview (Medium)
1327±67
39%
30%
57
22 Mistral Large 4
1265±95
32%
20%
25
23 GPT-5 mini (Low)
1232±67
23%
23%
47
24 DeepSeek V3.2
1227±92
13%
37%
30
25 Mistral Small 4 (High)
1191±141
16%
19%
32

Good Wins

60%

790 / 1322

Evil Wins

40%

532 / 1322

Slayer Hits

83

Fake Slayer Shots

210

Monk Blocks

253

Saint Executions

61

Imp Star Passes

40

Scarlet Transformations

114

Mayor Wins

13

Virgin Triggers

121

Ravenkeepers Murdered

138