Can an LLM solve a Redactle?

A Redactle article with hidden words represented by blocks

Redactle is a puzzle game where a redacted Wikipedia article is displayed. Guessing a word reveals all matches in the article, including variations like 1 and one. Once the entire title is revealed, the game is won.

One day I was pasting Redactle text into LLM chats from different providers and got wildly different results. I got nerd sniped and built a leaderboard to fairly evaluate as many as I could. $16.20 later, I now have this dashboard to answer that question.

The results are surprising. Reasoning effort did not help as much as I expected. Gemini and Grok are the best, not OpenAI and Anthropic models as I had expected. Gemini 3.7 Flash does exceedingly well, managing to snipe all 12 puzzles while being the cheapest and fastest.

Redactle results

Published Sep 2, 2026 · 264/288 attempts

  • One-shot
  • No-hint
  • Failed
#Model · reasoningAttemptsQualitySolve rateScoreCost / runTime / run
1 Gemini 3.8 Flash · low1212/0/0100%0.0$0.0055s
2 Gemini 3.7 Flash · low1212/0/0100%0.0$0.0056s
3 Gemini 3.8 Flash · medium1212/0/0100%0.0$0.0067s
4 Gemini 3.8 Flash · high1212/0/0100%0.0$0.00910s
5 Grok 4.6 · medium127/5/0100%1.3$0.04391s
6 Grok 4.6 · low125/7/0100%1.5$0.02223s
7 Gemini 3.7 Flash · medium1211/0/192%5.0$0.0067s
8 GPT-5.6 Sol · medium125/6/192%6.3$0.03438s
9 Kimi K3 · low123/8/192%7.9$0.061104s
10 GPT-5.6 Sol · low124/7/192%8.7$0.06359s
11 Muse Spark 1.2 · low122/8/283%12.8$0.06358s
12 GLM 5.3 · low122/8/283%13.2$0.02274s
13 GPT-5.6 Terra · low122/7/375%17.3$0.09579s
14 GPT-5.6 Luna · medium120/9/375%24.8$0.028185s
15 GLM 5.3 Flash · low124/3/558%26.5$0.001138s
16 Claude Sonnet 5 · low121/5/650%32.7$0.13756s
17 GPT-5.6 Luna · low121/5/650%35.0$0.027161s
18 GLM 5.3 Flash · high124/0/833%40.0$0.00194s
19 HY 4 Preview · low123/0/925%45.0$0.00952s
20 DeepSeek V4 Pro 0813 · low122/1/925%45.1$0.01562s
21 DeepSeek V4 Flash 0731 · low121/0/118%55.1$0.001266s
22 Gemini 2.5 Flash Lite · none120/0/120%$0.00687s
Qwen 3.8 Flash · low Alibaba returned prose instead of the required guess tool call; retries were rate-limited upstream.0/12Not scored: tool incompatible.
Qwen 3.8 Flash · medium Alibaba returned prose instead of the required guess tool call; retries were rate-limited upstream.0/12Not scored: tool incompatible.

Scroll sideways to see every measure.

Trade-offs

Each label is a complete model configuration. Lower and farther left is better; the shaded quadrant is below the median on both measures.

Cost versus score

Speed versus score

Method notes

A model must make Redactle word guesses until every meaningful title word is revealed. Scores measure accepted guesses above the article title's word-count par, so a model that reveals every title word with exactly one guess per title word scores 0. In the hint-enabled evaluations, every hint adds 50 points.

The 500-word evaluation stops after 30 accepted guesses. If a model is still solving at that point, the attempt receives a score of 60 guesses. That is a cost-saving shortcut, not an estimate of how many guesses the model would have needed to finish. Other configurations use the same rule at twice their own guess limit.

Results from incomplete configurations remain visible but are not ranked. Article titles, excerpts, and per-article outcomes are deliberately omitted so this page does not spoil Redactle puzzles.

View evaluation code on GitHub (opens in a new tab)
deenesfrnlptittrhear