A signature proves origin, not result quality.
Node identity, signature verification, and receiver trust are separate checks. Only receiver-verified evidence enters the trusted layer; synthetic records prove none of these states.
Theme preference system; currently light.
17 configurations · 72 tasks each · Publication pending
Loading interactive configuration matrix…
Official calibrated estimates
Top configuration domain profile: The live public read could not be read. No synthetic measurements were substituted.
Exact run, configuration, scoring-version, and provenance identity could not be verified.Official AIQ is a versioned 0–100 calibrated ability estimate. Its conditional 95% interval holds the estimated item bank fixed. Quality score, task-mix sensitivity, strict pass, time, and cost remain separate diagnostics. Synthetic previews expose quality only and are never ranked as Official.
Values come from RLS-protected public views. Inspect coverage, trust, scoring version, and run provenance before comparing entries.
Values come from RLS-protected public views. Inspect coverage, trust, scoring version, and run provenance before comparing entries.
Official efficiency: Values come from RLS-protected public reads.
Top estimate run context: The live public read could not be read. No synthetic measurements were substituted.
Exact run, configuration, scoring-version, and provenance identity could not be verified.Official efficiency is unavailable.
17 of 17 expected Official efficiency rows are Unavailable. 17 returned rows were rejected. Exact run, configuration, scoring-version, and provenance identity are required.
Loading history…
Loading run evidence…
Loading method…
Loading radar evidence…
Open any configuration to inspect its 72 task results, failures, timing, cost, and provenance.
Run archive: The live public read could not be read. No synthetic measurements were substituted.
Cannot read public_runs: invalid response shapeTrace where published results came from. AIQ has 3 registered production identities. Radar telemetry is not enabled, so live online or offline state is unknown.
Runner network: Values come from RLS-protected public reads.
Counts come from retained registry and signed observation records.
Node identity, signature verification, and receiver trust are separate checks. Only receiver-verified evidence enters the trusted layer; synthetic records prove none of these states.
| Node | Registry | Trust | Telemetry | Provenance |
|---|---|---|---|---|
| AIQ Official Runnerofficial | active | trusted verified | Registered identityRadar telemetry not enabled · live state unknown | Published |
| AIQ Independent Verifierverifier | active | independently reproduced | Registered identityRadar telemetry not enabled · live state unknown | Published |
| AIQ Official Publisherofficial | active | trusted verified | Registered identityRadar telemetry not enabled · live state unknown | Published |
official
sha256:2316ef9ab5fca992036fbce464e10499e3726b375df965e2e9c70ca5d01aa83cCapability telemetry is not enabled for this identity.
Observation telemetry is not enabled for this identity.
verifier
sha256:9e19d7366133ea7ab102bbde9279e384f1db92ec8e455ef06ac14f7383f86c8dCapability telemetry is not enabled for this identity.
Observation telemetry is not enabled for this identity.
official
sha256:63d5848f3e1bbe3d4f3e6c4e1f798222f79b3fdef915d7d69a4ba94530c0b3a6Capability telemetry is not enabled for this identity.
Observation telemetry is not enabled for this identity.
See what AIQ measures, how uncertainty is calculated, and which operational metrics stay outside the score.
Benchmark method: The live public read could not be read. No synthetic measurements were substituted.
Cannot read public_scoring_versions: stale active tupleTime is observed Codex adapter elapsed time. Cost is a versioned estimate from covered token aggregate usage and the official OpenAI API pricing documentation, accessed 2026-08-02. It is not actual Codex subscription billing or necessarily an exact API invoice.
Compare every published configuration now; trend lines begin after the next Official cycle.
History: Values come from RLS-protected public reads.
First published observation. Compare the configurations in this snapshot; trend lines begin after the next published cycle.
Interactive chart. The complete values are available in the following data table.
17 published series · scoring 1.0.8 · first observation; trend begins with the next published cycle
Latest visible date: 2026-08-28 UTC · 17 observations · calibrated ability range 69.4–77.6. Highest point estimate: sol-medium · conditional 95% interval 65.0–86.6 · strict pass 27.8% (n=72) · n=72 · coverage Unavailable · scoring 1.0.8 · published
Showing 17 configurations in canonical matrix order. Family and reasoning are explicit filters, not point-estimate cutoffs. Published buckets retain the latest exact Official run; synthetic fixture buckets expose no run detail. Lines connect observations for continuity only; they do not interpolate or estimate values between dates. Absent buckets remain gaps, bars use a zero baseline, and the server returns at most 20 buckets per configuration. Each point requires matching run, configuration, scoring version, and provenance identity. Scoring versions: 1.0.8.
Exact trend run context: The live public read could not be read. No synthetic measurements were substituted.
Exact run, configuration, scoring-version, and provenance identity could not be verified.| Recorded | Series | Primary metric | Primary interval | Strict pass | n | Run / bucket | Coverage | Runtime | Missing | Summed adapter duration | API-equivalent cost | Scoring | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-08-28 | sol-low | 76.4 · Calibrated ability | 63.5–85.7 · Conditional 95% interval | 26.4% (n=72) | 72 | run_602248d32fb259e607d90a9fc14707df48e308f4ede3921cc6324341cde06900 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | sol-medium | 77.6 · Calibrated ability | 65.0–86.6 · Conditional 95% interval | 27.8% (n=72) | 72 | run_5a2a4eca1cb69ceb798ada0d14415156ac4ab27f458ab646760fecac602b31b1 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | sol-high | 73.2 · Calibrated ability | 59.9–83.4 · Conditional 95% interval | 25.0% (n=72) | 72 | run_437b45de053bfca2ff1f69ad7dc57c006818f1ab4203786cc7db5cb6cde076ac Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | sol-xhigh | 75.9 · Calibrated ability | 63.0–85.4 · Conditional 95% interval | 27.8% (n=72) | 72 | run_b2372624b77abd0598e1c4f3b783a3726ec32241eacc3396d803f4ae0641b430 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | sol-max | 75.2 · Calibrated ability | 62.2–84.9 · Conditional 95% interval | 26.4% (n=72) | 72 | run_366299d303a738a36a1289a57a3234cc7059d0b057e620f6139ac02787899be3 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | sol-ultra | 75.2 · Calibrated ability | 62.2–84.9 · Conditional 95% interval | 27.8% (n=72) | 72 | run_395ab5e170b1b9653667e306f6c7ffd93d118175c19adacb0a7c5d5b8bfca7b1 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | terra-low | 70.0 · Calibrated ability | 56.2–80.9 · Conditional 95% interval | 29.2% (n=72) | 72 | run_0333c6971437abbd8e524383310d57e9441e80de282d6cd0ba85a4183602d1be Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | terra-medium | 74.5 · Calibrated ability | 61.3–84.3 · Conditional 95% interval | 27.8% (n=72) | 72 | run_8832ca59be523e36b05b5516ff7281b3f711eaf10cfe3581e45b8b95d1c4992b Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | terra-high | 74.6 · Calibrated ability | 61.4–84.4 · Conditional 95% interval | 23.6% (n=72) | 72 | run_bde48d35c4ab92dc81c7d7b05f60687b45850990211853d6674a97b37d9c56fe Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | terra-xhigh | 71.6 · Calibrated ability | 58.0–82.1 · Conditional 95% interval | 26.4% (n=72) | 72 | run_81e97c23f7c4ad230af28152a1c110db8164e0a8769f2dc78ceadb049cf0d294 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | terra-max | 71.6 · Calibrated ability | 58.0–82.1 · Conditional 95% interval | 26.4% (n=72) | 72 | run_29b8598df64fabb999b6f034956e2a10f13758a151d32272cb8cf8746b2d0c52 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | terra-ultra | 72.8 · Calibrated ability | 59.4–83.0 · Conditional 95% interval | 25.0% (n=72) | 72 | run_8dc4463121678365f69e79c11278f844c6b2acdd49f76da099a46e852be079ef Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | luna-low | 69.4 · Calibrated ability | 55.6–80.4 · Conditional 95% interval | 25.0% (n=72) | 72 | run_bdc18132086ad72aa7b1f858785bf0673abb4507c9f6633eaad13d06e1b84ad0 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | luna-medium | 70.6 · Calibrated ability | 56.9–81.4 · Conditional 95% interval | 23.6% (n=72) | 72 | run_92a9f7939a2dc7b961a3cbd6b13f719907ae2adfbaca013b6f086514266747cc Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | luna-high | 72.7 · Calibrated ability | 59.2–82.9 · Conditional 95% interval | 27.8% (n=72) | 72 | run_9c6cefcef6a726dc2ab54d6b851f06629e45097f8de1b6d47f709ff1cccee3b4 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | luna-xhigh | 75.7 · Calibrated ability | 62.8–85.3 · Conditional 95% interval | 29.2% (n=72) | 72 | run_652a8005e9e3066cf1987e27a140a421e256f7961aca372489ff0686f885ed83 Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
| 2026-08-28 | luna-max | 76.3 · Calibrated ability | 63.4–85.7 · Conditional 95% interval | 27.8% (n=72) | 72 | run_cc3cfefe0e6c96deb38f020dcb5154bda98b62333adc00bd816f9c9f6c7232ab Latest of 1 run | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | 1.0.8 | Published |
Same fixed response task, paired modes, 5 trials per mode. AIQ is not recomputed from speed, time, tokens, or credits. TTFT is not shown because the current CLI does not expose first-token timestamps.
Interactive chart. The complete values are available in the following data table.
| Configuration | Normal completion | Fast completion | Normal time | Fast time | Speedup | Normal output | Fast output | Normal credits | Fast credits |
|---|---|---|---|---|---|---|---|---|---|
| sol-low | 100.0% | 100.0% | 21.7 s | 16.1 s | 1.35× | 37.4 tok/s | 51.3 tok/s | 8.103 credits | 20.261 credits |
| sol-medium | 100.0% | 100.0% | 22.0 s | 16.0 s | 1.38× | 36.5 tok/s | 51.4 tok/s | 8.150 credits | 18.677 credits |
| sol-high | 100.0% | 100.0% | 21.3 s | 16.1 s | 1.33× | 38.9 tok/s | 51.3 tok/s | 8.165 credits | 18.735 credits |
| sol-xhigh | 100.0% | 100.0% | 23.3 s | 16.6 s | 1.40× | 36.5 tok/s | 50.3 tok/s | 7.525 credits | 20.468 credits |
| sol-max | 100.0% | 100.0% | 22.3 s | 16.4 s | 1.36× | 37.9 tok/s | 50.9 tok/s | 8.229 credits | 18.945 credits |
| sol-ultra | 100.0% | 100.0% | 22.4 s | 16.2 s | 1.38× | 38.4 tok/s | 52.4 tok/s | 8.297 credits | 20.657 credits |
| terra-low | 80.0% | 100.0% | 24.5 s | 18.5 s | 1.32× | 33.7 tok/s | 43.3 tok/s | 3.130 credits | 6.786 credits |
| terra-medium | 100.0% | 100.0% | 29.9 s | 19.2 s | 1.56× | 26.8 tok/s | 42.3 tok/s | 3.240 credits | 6.772 credits |
| terra-high | 100.0% | 80.0% | 21.1 s | 16.8 s | 1.26× | 39.0 tok/s | 49.1 tok/s | 2.669 credits | 5.901 credits |
| terra-xhigh | 100.0% | 100.0% | 21.4 s | 16.2 s | 1.33× | 38.7 tok/s | 51.5 tok/s | 2.732 credits | 6.181 credits |
| terra-max | 100.0% | 100.0% | 21.6 s | 15.9 s | 1.36× | 38.8 tok/s | 53.0 tok/s | 2.491 credits | 5.577 credits |
| terra-ultra | 100.0% | 100.0% | 21.2 s | 16.7 s | 1.27× | 39.4 tok/s | 50.5 tok/s | 2.575 credits | 5.791 credits |
| luna-low | 100.0% | 100.0% | 21.0 s | 15.4 s | 1.36× | 39.4 tok/s | 53.0 tok/s | 0.291 credits | 0.595 credits |
| luna-medium | 100.0% | 100.0% | 21.1 s | 15.5 s | 1.36× | 38.9 tok/s | 52.7 tok/s | 0.269 credits | 0.661 credits |
| luna-high | 100.0% | 100.0% | 20.8 s | 16.1 s | 1.29× | 39.7 tok/s | 51.1 tok/s | 0.270 credits | 0.662 credits |
| luna-xhigh | 100.0% | 100.0% | 20.8 s | 16.2 s | 1.28× | 40.0 tok/s | 50.9 tok/s | 0.268 credits | 0.598 credits |
| luna-max | 100.0% | 100.0% | 21.6 s | 16.5 s | 1.30× | 40.1 tok/s | 54.4 tok/s | 0.297 credits | 0.754 credits |
Aggregate output rate uses total elapsed time because the current Codex event stream does not expose a trustworthy first-token timestamp. Fast credit estimates use the published 2.5× ChatGPT credit multiplier. Neither measure changes AIQ, confidence intervals, or ranking.
Matrix entries: Values come from RLS-protected public reads.
Trend points: Values come from RLS-protected public reads.
Historical run context: The live public read could not be read. No synthetic measurements were substituted.
Cannot read public_runs: invalid response shapeHistorical efficiency: Values come from RLS-protected public reads.