Every week brings a new headline announcing that an AI system has matched or beaten humans at something. Some of these results are remarkable. Many are narrower than they sound. Here's how I read them.
1. Compared to what, exactly?
“Outperforms humans” usually means outperforms a specific group of people on a specific test under specific conditions. Were the humans experts or students? Did they have the same time and information as the model?
2. Is the test the task?
Benchmarks are useful, but they are simplifications. Scoring well on multiple-choice medical exams is not the same as diagnosing a patient who describes their symptoms in dialect, at 3 a.m., in a clinic with patchy internet.
3. Who was left out?
Check which languages and populations were tested. A system evaluated only in English tells you very little about its performance in Arabic — and the gap is often larger than people expect.
A result can be true, important, and still not mean what the headline says.
4 & 5. Who funded it, and has it been replicated?
- Industry research is valuable, but note when the authors also sell the product.
- Look for independent replication before treating a result as settled.
- Preprints are not peer-reviewed — which doesn't make them wrong, just early.
None of this requires a technical background. It requires the same skepticism we'd bring to any extraordinary claim — applied consistently, including to the results we'd like to be true.
