We're studying how well people can verify AI-written technical reports. Each report here was written by a language model, so some of its claims may be wrong. Finding them tells us how easy it is for a person to verify what a model writes.
Feel free to look things up. Search the web, open the papers, check the docs — just keep an eye on the clock.
Just don't ask an AI assistant (ChatGPT, Claude, etc.) to find the errors for you. We're measuring your judgment, not the model's.
The model that wrote this report also went back and checked its own work. The sentences its check was unsure about are underlined like this.
The underline means the model's own check thought that sentence might be wrong — here, perhaps the date or the amount. It doesn't tell you which part, and the sentence may in fact be perfectly correct.
Treat the underlines as hints, not answers. The check is unreliable in both directions: it misses real errors, and it flags sentences that are fine. An underlined sentence isn't necessarily wrong, and a plain sentence isn't necessarily right — judge every sentence yourself.
Notes are optional. A line on why you flagged each one helps us a lot.
Don't take the score to heart. The “correct answers” are imperfect right now. If you disagree with one, tell us below.