STATE OFAGENT
RELIABILITY
2026
OPEN RESEARCH RELEASE / LIVE EVIDENCE

Completion is not
correctness.

One reproducible view across every committed real-model result: exact task success, completion, safety boundaries, uncertainty, cost, latency, and provenance—without manufacturing a universal score.

Read the citable report ↗
model evals scenario trials verified labs observed failures recorded spend
01 / STATUS GAP points

Average distance between completion and exact task success on artifacts that report both.

Loading paired evidence…
02 / UNCERTAINTY points

Median width of the committed 95% confidence interval for exact-task endpoints.

A SCORE WITHOUT ITS INTERVAL OVERSTATES THE RESULT
03 / TRANSFER FAILURE labs

A valid rule from a clean twin is reused where one deciding fact reverses it.

SIMILARITY ERASES THE EXCEPTION
04 / MODEL IDENTITY pinned

Model results need provenance. A floating alias can exercise different weights on the next run.

Loading provenance…
INTERACTIVE EVIDENCE WORKBENCH

Ask a narrower question.

Every bar names its source metric and opens the exact committed result. Filter first; compare only like work.

EXACT TASK SUCCESSLoading evidence…
0.501.00

Selected endpoints remain lab-specific. Use this view to inspect evidence, not to average unlike tasks.

MODEL TERRAIN / COVERAGE BEFORE RANK

Where was each model actually tested?

Medians summarize the current filtered portfolio. Uneven coverage is shown beside every number and is never filled with a zero.

FAILURE TERRAIN / REPOSITORY INCIDENCE

The patterns that travel.

Counts mean “observed in these labs,” not real-world prevalence. Industry and contract filters narrow the map.

Boundary: this is an automatically generated snapshot of this repository’s committed evidence. It is not a market ranking, production certification, regulator approval, or estimate of failure prevalence.

OPENAAU
CHALLENGE
01
COMMUNITY RELIABILITY LEAGUE / $0 TO ENTER

30 minutes. One claim.
Bring the receipt.

Choose a bounded mission. Reproduce a result, break an assumption, or adapt a proven contract to help different people. CI checks the evidence; the public board credits the work.

live missions community participants community finishes reference finishes earned, never claimed
MISSION BOARD / FIVE STARTING POINTS

Pick the failure you want to make visible.

THE 30-MINUTE ROUTE

A small finish with a durable receipt.

The clock is a guide, not a gate. Adapt missions usually take longer; every step remains inspectable.

  1. 00—05
    ChooseClaim a mission and open its starter lab.
  2. 05—12
    RunExecute the deterministic mock—no key, no spend.
  3. 12—20
    TraceLink the claim to one exact scenario or branch.
  4. 20—27
    RecordCommit the result and a machine-readable submission receipt.
  5. 27—30
    ProveRun aau challenge validate and open the PR.
PUBLIC SCOREBOARD / EVIDENCE, NOT POINTS

Who returned with a receipt?

References are not community wins. They show the finish standard while the first independent contributors take the open positions.

ACHIEVEMENTS / DERIVED BY CI

No vanity badges.

SUBMISSION BUILDER / LOCAL-ONLY

Leave with the record already shaped.

This form downloads a Challenge receipt or Adapt Gallery starter in your browser. Nothing is sent, stored, or tracked.

Boundary: Challenge status is derived from repository evidence. It is not a model leaderboard, production certification, regulator approval, or permission to automate protected authority.

NEW / RIGHTS CONTINUITY CONTRACT

The appeal survives.
Does the protection?

One case can carry several rights on different clocks. The new matched suite catches the moment a technically valid main route silently loses coverage, urgent review, income, or another companion protection.

NEW / CRITICAL EVENT FAN-OUT

The response worked.
What stayed open?

Containment, initial reports, recipient notices, updates, and follow-ups remain independently true. A draft or successful response never closes the complete event graph.

PREVIOUS / REGULATORY CLOCK COLLISION BENCHMARK

One event.
Many clocks.

A consequential event rarely creates one tidy task. It creates independently timed duties for different recipients, through different channels, with protected human owners. This new matched benchmark catches the obligation an agent silently drops.

EVENT 00FACT PATTERN DUTY 01WHOrecipient DUTY 02WHENclock origin DUTY 03HOWchannel DUTY 04PROOFreceipt NO DUTY DISAPPEARS
NEW / SEVEN-INDUSTRY PUBLIC PROTECTION WAVE

The action sounds right.
Where is the receipt?

A remedy, dispute, safety report, or regulator notification is not complete because an agent drafted it. The new matched suite proves the exact subject, current rule, gates, clock and channel, accountable owner, and the event that actually executed.

01SUBJECT02RULE03GATES04CLOCK + CHANNEL05OWNER06RECEIPTPROTECTED
NEW / PROOF BEFORE ACTION

Helpful is not
the same as proven.

Three new labs test the moment plausible context becomes an unsafe shortcut. The agent must bind the action to the exact source, emergency state, or current award—then stop before the protected human decision.

01RETRIEVE02BIND03VERIFY04HAND OFFNO PROOF · NO ACTION
NEW / SIX-INDUSTRY DECISION GATE

Correct answer.
Wrong authority.

A new matched benchmark tests the last inch before consequential action: exact evidence, every satisfied gate, the rule-specific reason, procedural protections, human authority, and a record that matches what actually executed.

01OUTCOME02REASON03EVIDENCE04GATES05PROCEDURE06AUTHORITY07RECORD= EXACT
Small-business recovery pathway through evidence, accessibility, recourse, deadlines, rights, oversight, and a truthful record
NEW REPOSITORY SPECIALTY / PUBLIC VALUE

A right answer can still be a bad service.

Nineteen Public Value and Evidence Service Contract labs test what ordinary accuracy misses: minimum paperwork, accessible delivery, recourse, deadlines, protected authority, and a truthful executed record.

MINIMUM BURDENACCESSIBILITYRECOURSERIGHTS + RECORD TRUTH
NINETEEN PUBLIC-VALUE + EVIDENCE-SERVICE LABS

Choose the service failure you cannot afford.

Use the contract ↗

Every lab keeps the consequential decision with an accountable person. What changes is the proof that a merely “correct” route still has to pass.

NEW / MATCHED CROSS-INDUSTRY SUITETwelve services. One exact contract. No invented authority.

Each lab holds eight archetypes and the scorecard constant while changing the beneficiary, trusted records, policy, terminal actions, and protected decision.

08 / FOOD SAFETYRecall Traceability

Prove each lot edge without widening a recall through invention.

TRACE LEDGER ↗
09 / WATERDrinking-Water Notice

Keep unknown, sampled, noticed, delivered, and replaced states distinct.

NOTICE LEDGER ↗
10 / TAXPAYER SERVICEIRS Notice Response

Bind the exact action and evidence to the notice clock and response route.

RESPONSE MAP ↗
11 / VETERANSClaim Evidence

Separate what is held from what is requested—without rating a claim.

EVIDENCE MAP ↗
12 / ACCESSIBLE MOBILITYParatransit Access

Preserve accessible application, clock, trip-condition, and appeal service.

ACCESS CLOCK ↗
13 / TRANSPARENCYFOIA Routing + Appeal

Keep component, disclosure, tracking, response, and appeal state intact.

DISCLOSURE ROUTE ↗
14 / CITIZENSHIP SERVICEUSCIS Case Evidence

Organize official case state and notices without predicting an outcome.

CASE LEDGER ↗
15 / INTERNATIONAL TRADEExport Evidence

Bind item, party, end use, destination, screening, and rule version.

EVIDENCE FIREWALL ↗
16 / CHILD NUTRITIONSchool Meal Access

Reuse direct certification before asking a family to prove it again.

CERTIFICATION REUSE ↗
17 / ELECTION SERVICEProvisional Ballot Status

Report only official status and cure routes, privately and nonpartisan.

STATUS LEDGER ↗
18 / CARE TRANSITIONSDischarge Readiness

A printed plan is not ready until the downstream handoffs are received.

READINESS GATE ↗
19 / WORKFORCE MOBILITYLicense Mobility

Show a compact or endorsement path without claiming licensure.

AUTHORITY MAP ↗
FIELD NOTES / CONTROLLED EXPERIMENTS

The findings demos cannot show.

All 17 failure patterns ↗
AAU STUDIO / LOCAL-FIRST MATCHER

Bring a workflow.
Leave with an eval.

Describe the work and its consequences. Studio maps it to verified labs, explains every match, compares evidence, and produces a runnable fork path.

Runs in this browser · no account · no upload · no tracking
04 What must not go wrong?
Deterministic matching · inspectable evidence
THE VERIFIED CATALOG

Find your next evaluation.

Search by the work an agent does, the industry it serves, or the failure you need to prevent.

I want to
Loading verified use cases… RUNNABLE CODE · COMMITTED RESULTS · OBSERVED FAILURES
EVIDENCE COMPARISON

Compare shapes, not unlike scores.

Headline metrics differ across workflows. This view compares reusable architecture, verification depth, and observed failure evidence—not a universal leaderboard.