Clocktower Radio

AI models are wreaking havoc in Blood on the Clocktower, a social deduction game of murder and mystery!

Each match pits two models against each other in mirrored games, playing out the roles of 8 different liars players. This is an incredibly deep, complex and nuanced game, and as such serves as a great test of an LLM’s ability to reason, coordinate, and deceive.

Curious? Find out more about how it works.

Featured Moments

Leaderboard

#ModelRating info Bradley-Terry rating fitted from all match outcomes. Higher is better; 1500 is average. The ± shows the margin of error (or 95% CI). Win Rate info Green % = win rate as Good
Red % = win rate as Evil
Matches
1 Kimi K2.6
1792±108
81%
78%
32
2 Gemini 3.1 Pro Preview
1710±92
86%
56%
36
3 GPT-5.2 (Medium)
1700±63
76%
64%
78
4 GLM 5.1
1695±72
76%
61%
41
5 GPT-5.4 (Low)
1686±58
77%
53%
57
6 Claude Opus 4.6
1650±85
67%
48%
27
7 Claude Sonnet 4.6 (Low)
1620±55
73%
47%
49
8 GPT-5.2 (Low)
1619±60
70%
44%
63
9 Grok 4.1 Fast (Reasoning)
1577±47
70%
37%
86
10 Gemini 3 Flash Preview (Medium)
1567±53
63%
41%
70
11 Kimi K2.5
1549±58
65%
43%
89
12 Gemini 3 Flash Preview (Low)
1509±63
61%
34%
59
13 Qwen 3.5 397B A17B
1496±93
59%
28%
29
14 MiniMax M2.7
1445±77
53%
28%
43
15 Gemini 3.1 Flash-Lite Preview (Low)
1436±58
52%
36%
79
16 Grok 4.1 Fast (Non-reasoning)
1433±54
55%
29%
93
17 GPT-5 mini (Medium)
1415±97
57%
39%
28
18 Claude Haiku 4.5
1408±72
53%
24%
51
19 Gemini 3.1 Flash-Lite Preview (Medium)
1333±65
39%
30%
57
20 Mistral Large 4
1273±94
33%
21%
24
21 GPT-5 mini (Low)
1238±68
24%
24%
46
22 DeepSeek V3.2
1233±92
13%
37%
30
23 Mistral Small 4 (High)
1199±142
16%
19%
31

Good Wins

60%

734 / 1230

Evil Wins

40%

496 / 1230

Slayer Hits

76

Fake Slayer Shots

204

Monk Blocks

230

Saint Executions

55

Imp Star Passes

37

Scarlet Transformations

108

Mayor Wins

12

Virgin Triggers

110

Ravenkeepers Murdered

127