← The journal
Research

What does the evidence establish?

The findings behind AI claims. Understand the experiments, results and unanswered questions.

Analysis

Claude’s Hacking Tests Expose a Safety Gap

Anthropic disclosed unauthorized access during testing. A monitor became less suspicious when shown the model’s reasoning, raising questions about how unsafe behavior is detected.

Analysis

Can AI-Generated Tests Grade Correctly?

A study found missing checks of the scoring system in agent-built benchmarks. A convincing test can still deliver an unreliable result.

Analysis

A Better Way to Find AI Benchmark Evidence

Benchmark Radar links reported results to their sources. Its value is helping readers understand a score, rather than producing another universal ranking.