Claim Hunt Spot the false claims in the report.

We're studying how well people can verify AI-written technical reports. Each report here was written by a language model, so some of its claims may be wrong. Finding them tells us how easy it is for a person to verify what a model writes.

Before you start

How it works
  1. Flag a sentence only if it is clearly wrong. A fact, formula, definition, or mechanism you could state correctly, with nothing left to argue about. Click it to flag, click again to undo.
  2. Don't worry about wording that is just imprecise, oversimplified, or a bit strong. Save your flags for clear, checkable errors.
  3. You have 5 minutes per report — the clock at the top counts down, and at zero your flags are submitted automatically.
  4. Hit Submit when done, optionally note why on each flag, and you'll see your results.

Feel free to look things up. Search the web, open the papers, check the docs — just keep an eye on the clock.

Just don't ask an AI assistant (ChatGPT, Claude, etc.) to find the errors for you. We're measuring your judgment, not the model's.

Examples: what counts as a clear error
“Proximal Policy Optimization (PPO) was first introduced by DeepMind in 2015.”
False claim A specific, checkable fact that is simply wrong: PPO is from OpenAI (2017). When you know it's wrong, use False claim.
“Proximal Policy Optimization (PPO) is fundamentally an off-policy reinforcement learning algorithm.”
False claim A definitional error you can settle from understanding alone. PPO is on-policy, optimizing with data from the current policy. No lookup needed, so use False claim.
“The method improves accuracy by 9.3%.” (the paper actually reports 9%)
Doesn't count this round A number that's only slightly off, where you'd need the exact source to be sure. In this round a small discrepancy like this does not count — flagging it or leaving it won't change your score, so don't stress about catching it.
“By removing the value network, GRPO completely eliminates training instability.”
Doesn't count this round “Completely” is too strong, but this is an overstatement rather than a clear factual error. Borderline wording like this does not count this round either way.

One more thing: the underlines

The model that wrote this report also went back and checked its own work. The sentences its check was unsure about are underlined like this.

Example (made up)
The city approved the harbor bridge project in 2019, and construction finished in October 2023 at a cost of $2.1 billion, creating a new link between the two districts.

The underline means the model's own check thought that sentence might be wrong — here, perhaps the date or the amount. It doesn't tell you which part, and the sentence may in fact be perfectly correct.

How to use them

Treat the underlines as hints, not answers. The check is unreliable in both directions: it misses real errors, and it flags sentences that are fine. An underlined sentence isn't necessarily wrong, and a plain sentence isn't necessarily right — judge every sentence yourself.

5:00 · 0 flagged
Flag a sentence only if it is definitely wrong: a checkable fact, formula, or mechanism you could correct. Leave vague or overstated wording alone. Submit when done.

Add your notes

Notes are optional. A line on why you flagged each one helps us a lot.

Results

Don't take the score to heart. The “correct answers” are imperfect right now. If you disagree with one, tell us below.