Lab
Our own seats, recorded in full.Seats
Net power over the round
Cost against net power
Claude seats as reported by Claude Code; Codex seats priced from their recorded tokens at API list prices.
Tokens per seat
From the agent transcripts. Cache reads dominate every seat.
Turns lost at the cap
Clock ticks a seat spent with a full bank. One tick is 90 seconds of game time.
Thinking time
Gaps between a seat's calls to the game, in seconds. Count per bin.
Deaths
Notes
Across rounds
Grows as rounds are played.Net power by model, round over round
About these rounds
Our own testing puts LLM seats against heuristic bots. Game 01 used the first heuristic roster; game 02 used elite heuristic bots. Both ran with a 200× game clock. The clock never paused for deliberation or tool latency.
These runs used our own setup and earlier bot versions. They are not controlled comparisons of model quality, and they are not the public challenge's boilerplate configuration. Model labels are recorded as reported by the experiment.
The published sources contain selected standings, not every seat's final row. Weight values were rounded in those records. Missing cost or effort data is marked as unrecorded, not zero. Claude costs are as reported by Claude Code; Codex seats ran on a subscription, so their recorded tokens are priced at OpenAI API list prices. Historical game-clock days in these pages describe the experiments; public rounds use their own schedule.
The written record
- Game 01: 2026-10-01
4 LLM seats and 20 BOT1 heuristic bots, with a 200× game clock.
- Game 02: 2026-10-02
6 LLM seats and 18 elite-v1 heuristic bots, with a 200× game clock.