Methodology

Transparent scoring, version by version.

AIQ v1 estimates outcomes on one committed 72-task fixture set. It does not estimate general intelligence or universal model capability.

Benchmarkaiq-core@1.0.0
Scoring1.0.0
Published7/30/2026
01 · Fixed-fixture estimand

Observable outcomes on the committed set

  1. Give each of the ten domains weight 0.1.
  2. Keep the frozen domain and difficulty quotas.
  3. Keep missing and invalid tasks in completion accounting and block Official publication.
  4. Classify complete synthetic fixtures as descriptive Synthetic Complete, never Official or ranking eligible.
  5. Treat attributable agent, model, tool, timeout, budget, and wrong-artifact failures as valid zero scores.
  6. Treat benchmark infrastructure failures as invalid and audit a rerun.
02 · Domain coverage

72 tasks · 10 equally weighted domains

debugging8 tasks · 10%
retrieval_verification7 tasks · 10%
instruction_following6 tasks · 10%
documentation_communication7 tasks · 10%
planning_execution7 tasks · 10%
data_processing8 tasks · 10%
reliability_recovery7 tasks · 10%
repository_understanding7 tasks · 10%
coding8 tasks · 10%
tool_use7 tasks · 10%
03 · Completeness and validity

Official, complete synthetic, provisional, or coverage-only

Attempted failure
Attributable failures are valid zero scores. Infrastructure failures are invalid and require an audited rerun.
Missing fixture or result
Missing and invalid tasks block Official. Synthetic Complete and Provisional output use descriptive observed domain means and fixed-fixture completion bounds without ranking eligibility.
Hidden fixture boundary
Hidden payloads stay sealed behind the published fixture-set commitment. Version, task counts, outcome states, and provenance remain public.
04 · Sensitivity, not generalization

What the interval can and cannot say

The task-resampling interval uses finite_cluster_calibrated_percentile_sensitivity_v1 with a versioned 1.3 deviation correction calibrated for this fixed benchmark fixture. It is a fixed-fixture calibrated sensitivity interval, not a universal confidence interval for model capability.

Fixed-fixture or conditional AIQequal-weight mean of 10 frozen-fixture domain means