Comparison studio

One model is not one behavior.

Compare exact model and reasoning-level pairs. Keep sample size, coverage, failure counts, scoring version, and task-set sensitivity beside each fixed-fixture point estimate. The public comparison is descriptive because aggregate leaderboard rows do not contain the paired-task evidence required for a statistically supported difference.

No published evidence

The live public read is available, but it has no evidence to display.

No comparable entries are available.