Hash model, policy, tools, identity, authority, dependencies, egress, monitoring, and rollback.
Ship evidence,
not confidence.
Capture the exact agent release, map every changed component to its impacted tests, block uncovered or failing changes, invite oracle-free outside reproduction, exchange privacy-bounded incident regressions, and watch the official sources beneath the evidence.
One change. Four proofs.
No aggregate score can substitute for knowing exactly what changed, which boundary it touches, whether the regression passed, and who still owns the consequential decision.
Union before-and-after impact tags. An unmapped change or missing suite fails closed.
Publish tasks without answers; bind submissions to the committed hidden-oracle digest.
Share safe regressions and machine-readable findings—not credentials, targets, or exploit steps.
Test what actually moved.
The reference candidate changes policy, tool, and authority components. The gate runs only mapped suites, demands both legitimate twins, and exports a non-certifying OSCAL Assessment Results view.
— impacted evidence tags
A pinned SHA still needs an origin.
The lock proves that every immutable-looking Action object resolves as a commit in the repository named by the workflow. Version 1.1 uses stable job-scoped locators, so harmless line shifts do not erase reviewed evidence.
— stable job-scoped uses bound to — commits.
Inventory the blast radius.
Models are only one part of an agent release. The AABOM binds tools, operations, resource scopes, expiring authority, delegation, data routes, revocation, accountable ownership, stop controls, and evidence—then surfaces widening without a trust score.
— widening facts need owner review.
Operational context around the model.
Exports an officially schema-validated CycloneDX — projection plus a deterministic, tamper-evident pack. Inventory is not current authorization, verified identity, certification, safety, compliance, deployment approval, or an ATO.
Observe. Challenge. Narrow only with proof.
Compare privacy-bounded public, synthetic, or authorized aggregate action metadata with the exact AABOM. Exact tool-operation-scope relationships prevent accidental Cartesian authority; unobserved bindings become owner-review questions—not automatic policy edits.
— separate proofs are required before a human-owned narrowing change. A bounded non-use window never proves that authority is unnecessary, and the planner emits no executable policy.
Turn every declared grant into a challenge.
The inventory compiles exact relationships into legitimate clean twins and one-boundary violations—including valid operations and scopes that are unsafe in combination—then runs through an exact byte-bound command adapter without revealing expected answers or executing a tool.
Run the protocol ↗— boundary shapes · — legitimate blocks · adapter SHA-256 —. A local byte match does not prove provenance, production enforcement, safety, certification, compliance, deployment approval, or an ATO.
The oracle is committed. The answers are not.
— tasks across — public challenges are ready for an outside fork. Copy an oracle-free starter command below; GitHub attests submitted bytes, while independence and outcome still require human review. Accepted reproductions: — · campaign lock: —.
Regression before rhetoric.
— machine-readable exports: SARIF, OpenVEX, and explicitly experimental CSAF + OCSF bridges. Public or synthetic material only.
Know when evidence ages.
The weekly radar compares stable bytes or versioned visible-text fingerprints, then checks which exact revision each profile evaluated. MCP — has a —-case gate; A2A — has a —-case gate. It never interprets policy, edits a lab, or calls alignment conformance.
Loading committed release evidence…
Portable Agent Assurance.
Bind proof to authority.
One public, machine-verifiable envelope binds the agent, accountable operator, live authority lease, exact MCP or A2A action, and reproducible evidence. Eighteen clean and adversarial fixtures test where those bindings break.
Issuer, credential, resource, audience, scope step-up, self-describing headers, and token transport—tested as recorded deltas, not protocol conformance.
Version negotiation, selected interface tuple, PascalCase methods, exact card bytes, per-operation authority, and task scope—recorded evidence, not A2A conformance.
Six bindings. One inspectable decision.
A token alone cannot prove that an agent still has authority. The envelope makes every consequential link explicit and the receipt chains every exact result.
Verify the synthetic workload credential and normalized SPIFFE subject.
Keep the accountable human or organization distinct from the agent.
Bind lease, task, policy epoch, time window, revocation, and delegation.
Match MCP or A2A operation, resource, destination, and peer exactly.
Emit privacy-bounded events, exact reason codes, and source digests.
Verify the pack, byte manifest, in-toto statement, and result chain.
Change one binding. Demand the right stop.
Sixteen adversarial collisions and two legitimate twins distinguish safe refusal from blanket blocking. Every expected decision and reason code is committed before evaluation.
Result chain loading
Derived from the committed reference receipt.What this release does not prove
PRODUCTION IDENTITY / NOT VERIFIED
Four stages. No magic score.
The experimental profile maps directly to the draft TEVV-Athlon stages while preserving planned events, revealed material, and missing independent reproduction as visible structure—not a false maturity number.
Inspect an envelope without uploading it.
ZERO UPLOADThe browser reads only your local file. Structural checks are educational; use the Python verifier for signature and exact protocol evaluation.
No local envelope selected · NIST draft comment deadline 2026-10-06
Authorization is a live condition.
Prove it at runtime.
Agent identity is not enough. Test whether authority survives delegation, changing policy, peer contact, tool calls, network destinations, monitor loss, safe-stop, and recovery—then preserve the outcome in a deterministic receipt.
One boundary. Every consequential event.
The policy decision point evaluates normalized events from generic JSON, MCP, OpenAI Agents, LangGraph, CrewAI, and AutoGen-shaped recordings.
Bind the actor
Reject the wrong lease, audience, policy epoch, delegation depth, or revoked authority.
Bind the purpose
Keep the task, state, time window, and human-granted scope explicit throughout the run.
Bound the effect
Check tools, targets, network destinations, peer messages, and record mutations before use.
Observe continuously
Loss of the required monitor changes what the agent may do; it is not merely a log warning.
Fail closed
Critical alerts, invalid sequences, and revoked leases move the runtime to a safe state.
Make it replayable
Exact reason codes, state transitions, source hashes, and a deterministic chain remain inspectable.
A timeout is not permission to do it twice.
Bind the exact intent to approval, authority, policy, and one idempotency key. An unknown result must be reconciled with the authoritative system before retry.
Same key + same intent replays the receipt. Same key + changed amount, target, or policy blocks. Compensation requires a separate approval and remains a second recorded effect.
loadingORACLE-FREE COMMAND ADAPTER48/48 exact outcomes0 unsafe effects · 0 retry-after-unknown violations · receipt —
TWO-PROCESS CRASH LAB12/12 exact recoveries across 6 crash points0 unsafe resumes · 3 uncertainties preserved
MULTI-PROCESS RACE LAB12/12 exact races across 61 attempts0 duplicate effects · 0 missing legitimate effects
COMPLETE CI MATRIX72/72 exact across 3 independent gates0 unsafe outcomes · 3 entrypoints + 8 static materials · 11 runtime-read digests across 109 processes, including 3 runtime-only policy reads · 42 unresolved import names exposed · fully stressed loading · matrix —
RELEASE EVIDENCE BINDING1/1 consequential tool + operation + scope relationships boundrelease loading · static sets 3/3 match · runtime snapshots 3/3 match · authority twins 8/8 exact against adapter — · 0 holds · receipt —
Compare controls without a magic score.
Switch between matched transparent synthetic arms. The cases stay fixed; only the declared controls change.
Loading matched experiment…
— controlsLoading the source-bound case comparison…
Small-operator kits with visible gaps.
Each four-week kit protects accountable human authority, cites primary public guidance, and separates planned controls from evidenced ones.
Turn public lessons into tests that stay fixed.
The Incident Regression Commons converts a source-bound incident record into exact pre-fix and post-fix cases, preserves a legitimate twin, and scans contributions for sensitive categories.
Replay the public example ↗ Open the control experiment ↗No simulated partner claim.
External review, field observation, a human baseline, and independent reproduction remain visible until real partners produce them.
Propose a bounded pilot ↗ Open the partner form ↗Prove the fix.
Prove the stop.
A shared, open protocol for organizations, essential-service operators, governments, security partners, and frontier AI labs to test defensive work—without exposing targets, uploading inventories, or turning a synthetic run into a field-effectiveness claim.
Seven modules. One evidence spine.
Use only the layer you need, or carry one receipt from public signal through fix verification, runtime containment, continuity review, evidence exchange, and honest aggregation.
A patch is not a fix until the twin still works.
Every contract checks the vulnerability regression, legitimate twin, service-continuity budget, rollback readiness, evidence, and accountable human owner.
Measure control boundaries separately.
Stop latency is not correctness. Correctness is not evidence coverage. Neither proves service preservation. The receipts keep these signals distinct.
Synthetic containment clock
Five defensive capability families
Route a fix without uploading your inventory.
Choose an AAU campaign JSON or load the fictional water-system example. The file stays inside this browser tab; no scan, network request, model call, or automatic change occurs.
Local continuity-aware route planner
Public, synthetic, or explicitly authorized inventory onlyZERO UPLOADNothing has been loaded. Use the example button above or choose a local campaign.
Your continuity-aware fix routes will appear here.
| Vulnerability | Asset | Declared | Recommended | Continuity | Human | Gate |
|---|
Share the receipt, not the sensitive world.
Portable evidence retains artifact hashes, control fingerprints, bounded measurements, and evidence levels. The observatory counts artifacts—not organizations—and exposes what is still missing.
Interoperable, privacy-bounded evidence
Loading the committed evidence index…
No universal safety score
Fix cases, containment events, defender decisions, and benchmark tasks stay separate. A reference-exact receipt is never promoted to field effectiveness.
Loading visible evidence gaps…
Make “independent” earn its evidence.
An issuer commits a hidden oracle, an outside reproducer answers the gold-free challenge, and a separate reviewer reveals and adjudicates it. Bytes are proved by hash; organizational independence remains an explicit human-reviewed claim.
Challenge → run → review → reveal
The committed walkthrough deliberately uses maintainer-simulated roles. It teaches the protocol after oracle reveal and never advances the independent count.
Inspect an adjudication locally.
This browser checks the embedded digest and role gate. Full pack verification still requires the CLI because the oracle, receipt, statement, and manifest must recompute together.
No local adjudication selected.
What unlocks one accepted reproduction?
Bring your agent.
Leave with evidence.
Turn an existing command or HTTP agent into a runnable evaluation project: synthetic cases, exact expected outcomes, forbidden-action checks, a protected human boundary, a privacy-bounded receipt, tests, documentation, and least-privilege CI.
Build the evidence spine
STRUCTURAL READINESS ≠ PRODUCTION VALIDATION. Passing the starter proves that an adapter can satisfy three reviewed synthetic cases under the declared contract. It does not demonstrate production safety, legal compliance, domain validity, model robustness, or authority to deploy.
Your fork deserves
a receipt.
Turn a connected Agent Evidence Starter into an inspectable contribution pack: reviewed synthetic cases, aggregate receipts, a privacy scan, deterministic evidence checks, a SHA-256 manifest, and a share card—without uploading private traces.
Inspect, derive, and export
Nothing leaves this tab. Files are read in memory and form contents are neither uploaded nor persisted. Never select private details, prompts, raw model responses, credentials, classified information, CUI, or procurement-sensitive records.
Reference packs you can fork
These initial entries are explicitly labeled maintainer references. Community work will appear with contributor credit after the same validator and review route pass.
Evidence level, not endorsement. The ladder summarizes present public artifacts. It does not authenticate people, prove domain correctness, certify a system, establish legal compliance, imply U.S. Government approval, or grant permission to deploy.
An agent score answers
half the question.
Build a blinded, same-suite comparator for the existing human process. Measure exact outcome, abstention, task time, confidence calibration, and agreement—locally, without collecting names, free text, demographics, or production records.
Can you preserve the boundary?
Ground truth stays hidden until all tasks are complete. Choose a route and declare confidence; task time is measured only in this tab.
Start when you are ready.
The eight tasks contain only synthetic fields. No answer, response, timing, or form value leaves this browser tab.
Now compare like with like.
BASELINE ≠ REPLACEMENT DECISION. The lab does not verify institutional review, prove causal benefit, rank workers, support employment action, certify a system, or authorize deployment. Synthetic reference sessions test the protocol—not human performance.
Stop at the score,
or prove what changed.
The Evidence Commons binds the artifacts needed to move from a synthetic result toward useful public evidence. It never hides a missing layer behind a badge, leaderboard, or trust score. Inspect three partner-ready pilots for FOIA routing, accessible digital services, and small-nonprofit grant administration.
Impact Capsule
Loading the public evidence chain.
Who should benefit
Human authority stays with
Missing evidence—kept visible
Predeclared public-value measures
EVIDENCE CHAIN ≠ APPROVAL. Status is derived from public artifact presence. AAU does not verify identity, institutional review, causal impact, production fitness, independence, certification, government endorsement, or authority to deploy or automate a protected decision.
Turn an AI idea into
inspectable evidence.
Map a public mission to current OMB acquisition and high-impact practices, NIST AI RMF functions, hard human boundaries, realistic tests, monitoring, remedies, and exit conditions. Export a 12-file assurance pack with a SHA-256 manifest.
Build the evidence spine
Form contents stay in this browser tab. The Studio does not upload, persist, or transmit them. Never enter classified, controlled, procurement-sensitive, source-selection, or personally identifiable information here.
Name the mission to begin. Every unresolved item remains visible.
Export evidence, not a badge.
Complete the required mission fields and clear the local sensitive-data scan.
Federal AI Acquisition Performance Gate
Thirty-two seeded scenarios test intended-environment evidence, data-use conflicts, portability, pricing, monitoring, deadlines, and protected award authority.
Claims enter.
Gaps stay visible.
Fork a public mission, publish measurable gates, connect claims to evidence, recompute exact synthetic tests, preserve accountable decisions—and close the pilot with a bounded lesson that the next team can verify.
Verify the tool before it verifies a claim.
Security posture, release provenance, operating roles, and exit evidence now travel with the evaluation contract—without claiming federal approval.
Byte, nesting, node, path, symlink, archive, and manifest limits are exercised in CI.
Inspect the threat model ↗Deterministic ZIP, exact manifest, SPDX 2.3 inventory, SHA-256 checks, and GitHub attestations.
Verify a release ↗Eight working files take a team from mission scope and decision rights through tests, metrics, feedback, and exit.
Open the launch pack ↗No vendor rank, award recommendation, certification, ATO, or automated protected decision is produced.
Read the contract ↗A pilot ends.
The lesson travels.
Success, change, and stop decisions become evidence-linked public memory—with explicit non-transfer conditions, privacy review, and dated policy dependencies. No vendor leaderboard. No award recommendation.
Loading the public exchange…
Scan a lesson before it leaves your machine.
The browser checks structure, explicit sharing attestations, and narrow sensitive-data patterns. It never uploads, stores, or sends the file. A zero-finding result is not disclosure authorization.
No local lesson selected.
Publish the outcome and the stop line.
Inspect your fork locally
Choose three public or synthetic JSON files. Files stay in this browser tab—no upload, account, storage, analytics, or model call. The narrow sensitive-data scan is not a DLP system.
Loading the public reference exchange…
Loading…
—
A passing submitted synthetic case means only that the recorded output matches the declared oracle. It does not prove independent reproduction, production performance, compliance, or approval.
Ask before writing terms.
These prompts surface acquisition questions. They are not clauses, legal advice, certification criteria, or a substitute for current agency review.
python federal-pilot-kit/aau_pilot.py assess \
path/to/agency-intake.json \
path/to/vendor-response.json \
path/to/acceptance-tests.json
python federal-pilot-kit/aau_pilot.py closeout \
path/to/agency-intake.json path/to/vendor-response.json \
path/to/acceptance-tests.json path/to/lesson.json \
--out /tmp/aau-public-lesson
See the portfolio.
Interrogate the evidence.
Turn a public or synthetic AI inventory into inspectable quality gaps, possible-overlap questions, bounded public-value measurements, three-layer TEV&V coverage, and testable acquisition obligations.
Find the question behind the row.
Loading the evidence map…
Similarity opens a question. It never closes an investment.
Bring a public or synthetic portfolio.
Files stay in this browser tab. The Observatory never uploads, stores, or transmits them. Never include PII, credentials, controlled, classified, or procurement-sensitive information.
No local inventory selected.
Baseline first. Measurement second. Savings claim never inferred.
Test from three different distances.
An obligation is useful when failure is testable.
Can you trust
this agent?
Five real model traces. Inspect the case, the proposed action, and the evidence receipt. Choose Trust, Verify, or Block before the committed ground truth is revealed.
Read first. Judge second. Reveal last.
Opening the evidence room…
The committed scenario and model trace are loading.
—
Loading model reasoning…
Your evidence-review pattern
This describes one local session. It is not a score of professional ability or a certification.
Reproduce one trace, find a counterexample, or adapt a contract to a new public-interest case.
Enter the 30-minute Challenge ↓Boundary: synthetic scenarios and committed model evidence for education and evaluation. This is not production certification, legal advice, regulatory approval, or authority to automate a protected decision. Progress stays in your browser; nothing is submitted.
Change one fact.
Watch the contract move.
Eight verified scenario pairs expose the smallest semantic boundary that changes the required action. Review both sides before the oracle appears, then download the pair as a regression test.
Same-looking work. Different required action.
Opening the split room…
One deciding fact means one declared semantic boundary—not one raw JSON field. Case IDs, prose, required evidence, and derived contract fields may change because that fact changed.
Loading…
The source scenario is loading.
—Loading…
The source scenario is loading.
—Choose one action for each scenario to unlock the boundary.
One fact moved the contract.
—
You mapped 0/8 boundaries exactly.
This is a local learning receipt, not a credential or professional assessment.
Boundary: synthetic scenarios, fictional records, and source-linked evaluation contracts for education and testing. This lab does not provide legal, medical, regulatory, safety, employment, benefits, or operational advice. It never grants authority to automate protected decisions.
Bring a workflow.
Leave with a fork.
Define one semantic boundary, protect human authority, and export an eight-file contribution bundle with scenarios, tests, evidence notes, a Forge brief, and original artwork. The builder checks structure; qualified people still own domain truth.
From operational question to testable pair.
Boundary: this tool generates evaluation infrastructure from information you enter locally. It does not verify sources, determine policy, replace qualified review, inspect private systems, submit to GitHub, or make the resulting workflow safe for production.
Bring the receipt.
See what it proves.
Open an eval_*.json in your browser, recompute its trial grid, metrics, cost, latency, confidence-interval declarations, and model provenance, then export an aggregate-only receipt. No model call, account, or file upload.
Inspect structure before believing the headline.
Use synthetic, public, or safely redacted results. The browser reads locally and deliberately excludes trial details from every exported receipt.
Or choose a local eval_*.json up to 12 MB.
—
—
—
Counts and aggregates must reconcile.
Warnings stay visible without rewriting history.
Separate answers, never one score.
Exactness, completion, and safety are selected by a published metric priority. Inverted risk values are labeled.
Lowest five trials.
Only identifiers and operational aggregates appear. Scenario content stays out of the dashboard and exports.
Which model was actually served?
Warnings without transmission.
—
Boundary: Receipt Lab checks a narrow JSON contract and recomputes declared aggregates. It does not validate domain truth, evaluate source quality, certify a system, approve a protected decision, establish regulatory compliance, or independently reproduce a run.
RELIABILITY2026
Completion is not
correctness.
One reproducible view across every committed real-model result: exact task success, completion, safety boundaries, uncertainty, cost, latency, and provenance—without manufacturing a universal score.
Average distance between completion and exact task success on artifacts that report both.
Loading paired evidence…Median width of the committed 95% confidence interval for exact-task endpoints.
A SCORE WITHOUT ITS INTERVAL OVERSTATES THE RESULTA valid rule from a clean twin is reused where one deciding fact reverses it.
SIMILARITY ERASES THE EXCEPTIONModel results need provenance. A floating alias can exercise different weights on the next run.
Loading provenance…Ask a narrower question.
Every bar names its source metric and opens the exact committed result. Filter first; compare only like work.
Selected endpoints remain lab-specific. Use this view to inspect evidence, not to average unlike tasks.
Where was each model actually tested?
Medians summarize the current filtered portfolio. Uneven coverage is shown beside every number and is never filled with a zero.
The patterns that travel.
Counts mean “observed in these labs,” not real-world prevalence. Industry and contract filters narrow the map.
Boundary: this is an automatically generated snapshot of this repository’s committed evidence. It is not a market ranking, production certification, regulator approval, or estimate of failure prevalence.
CHALLENGE01
30 minutes. One claim.
Bring the receipt.
Choose a bounded mission. Reproduce a result, break an assumption, or adapt a proven contract to help different people. CI checks the evidence; the public board credits the work.
Pick the failure you want to make visible.
A small finish with a durable receipt.
The clock is a guide, not a gate. Adapt missions usually take longer; every step remains inspectable.
- 00—05ChooseClaim a mission and open its starter lab.
- 05—12RunExecute the deterministic mock—no key, no spend.
- 12—20TraceLink the claim to one exact scenario or branch.
- 20—27RecordCommit the result and a machine-readable submission receipt.
- 27—30ProveRun
aau challenge validateand open the PR.
Who returned with a receipt?
References are not community wins. They show the finish standard while the first independent contributors take the open positions.
No vanity badges.
Boundary: Challenge status is derived from repository evidence. It is not a model leaderboard, production certification, regulator approval, or permission to automate protected authority.
The appeal survives.
Does the protection?
One case can carry several rights on different clocks. The new matched suite catches the moment a technically valid main route silently loses coverage, urgent review, income, or another companion protection.
Reuse reliable agency data before asking a family to rebuild its case.
FORM ≠ FIRST STEP ↗ 02 / HEALTH APPEALSUrgency changes orderPreserve expedited internal and external review without deciding coverage.
URGENT ≠ ROUTINE ↗ 03 / DISABILITYTwo clocks, two rightsA timely appeal does not prove the shorter continuation election was timely.
60 DAYS ≠ 15 DAYS ↗The response worked.
What stayed open?
Containment, initial reports, recipient notices, updates, and follow-ups remain independently true. A draft or successful response never closes the complete event graph.
Containment, one-hour notification, 48-hour update, and receipts remain separate.
STOPPED ≠ REPORTED ↗ 02 / HEALTH DATAActor → recipientsPreserve business associate, covered entity, people, HHS, media, and substitute notice.
APPROVED ≠ NOTIFIED ↗ 03 / CLINICAL TRIALSignal + follow-upSeven-day, 15-day, qualified judgment, recipients, and follow-up do not collapse.
SERIOUS ≠ AUTOMATIC SUSAR ↗One event.
Many clocks.
A consequential event rarely creates one tidy task. It creates independently timed duties for different recipients, through different channels, with protected human owners. This new matched benchmark catches the obligation an agent silently drops.
Death, serious injury, malfunction, importer, and manufacturer duties must not collapse into one familiar report.
5 WORKDAYS ≠ 30 CALENDAR DAYS ↗ 02 / DRUG SUPPLYShortage NotificationSix-month notice, the as-soon-as-practicable fallback, and the five-business-day ceiling remain distinct.
PLANNED ≠ UNPLANNED INTERRUPTION ↗ 03 / MORTGAGE SERVICINGLoss Mitigation GateThe 45-, 37-, and 30-day protections attach to different facts and different servicer duties.
EXACTLY 37 ≠ MORE THAN 37 ↗ 04 / HEALTH PAYMENTNo Surprises IDRNegotiation and federal IDR clocks run in sequence; opening one stage does not complete the next.
30 BUSINESS DAYS → 4 BUSINESS DAYS ↗ 05 / SECURITIESCyber Disclosure GateThe disclosure clock starts at materiality determination—not merely at incident discovery.
DISCOVERED ≠ DETERMINED MATERIAL ↗ 06 / LONG-TERM CARETransfer + DischargeBasis, destination, notice, appeal, and updated circumstances travel together.
CHANGED DESTINATION → NEW NOTICE CHECK ↗ 07 / NUCLEAR OPERATIONSReactor Event NotificationOne-, four-, and eight-hour paths can coexist while emergency classification and plant control remain human-owned.
ACTUATION ≠ PREPLANNED ACTUATION · DRAFT ≠ NOTIFIED ↗The action sounds right.
Where is the receipt?
A remedy, dispute, safety report, or regulator notification is not complete because an agent drafted it. The new matched suite proves the exact subject, current rule, gates, clock and channel, accountable owner, and the event that actually executed.
Match the exact open campaign and preserve the no-cost authorized path.
APPOINTMENT ≠ REPAIRED ↗ 02 / PRODUCT SAFETYProduct → RemedyKeep model/date-code scope, stop-use language, and the official remedy together.
INTAKE ≠ COMPENSATION ↗ 03 / MARITIMECharge → ContractTest billed party, duplicate charge, free time, events, and the 30-day invoice clock.
DISPUTE ≠ WAIVER ↗ 04 / ENVIRONMENTCustody → CorrectionReconcile the manifest without inventing signatures or enforcing a proposal as law.
CORRECTION ≠ ERASURE ↗ 05 / CONSUMER FINANCEDebt → RightsRebuild notice delivery and the live validation period before collection continues.
DELIVERY ≠ VERIFIED ↗ 06 / EMERGENCY COMMSOutage → NORSDetect the 911/988 special-facility path even below the familiar volume threshold.
DRAFT ≠ FILED ↗ 07 / WORKPLACE SAFETYIncident → ReportClassify the exact medical outcome, event window, knowledge time, channel, and later update.
HOSPITAL ≠ INPATIENT · ATTEMPT ≠ ACCEPTED ↗Helpful is not
the same as proven.
Three new labs test the moment plausible context becomes an unsafe shortcut. The agent must bind the action to the exact source, emergency state, or current award—then stop before the protected human decision.
The source exists. Does its passage actually entail the sentence?
HUMAN EDITOR PUBLISHES ↗ 02 / HOME & FIELD SERVICESNo heat can be scheduled. Gas odor or a CO alarm must change the channel.
EMERGENCY PATH STAYS SEPARATE ↗ 03 / NONPROFIT GRANTSLast year's accepted packet is not proof for this award or this cost.
AUTHORIZED OFFICIAL CERTIFIES ↗Correct answer.
Wrong authority.
A new matched benchmark tests the last inch before consequential action: exact evidence, every satisfied gate, the rule-specific reason, procedural protections, human authority, and a record that matches what actually executed.
Chemical OOS discretion must not leak into the stricter sterility-positive path.
QUALITY OWNER HOLDS RELEASE ↗ GRID / CONJUNCTIVE SAFETYRestoration GateFour conditions and the clearance owner survive outage pressure together.
NO URGENCY WAIVER ↗ HIRING / CANDIDATE RIGHTSHiring ComplianceA legitimate screen cannot erase an AEDT notice or pre-adverse process.
HUMAN OWNS THE DECISION ↗ AVIATION / APPROVED DATAAircraft DispatchA similar aircraft's deferral is not this aircraft's approved MEL authority.
PIC + DISPATCHER BOUNDARY ↗ BANKING / CONFIDENTIAL CASEAML · KYC · SanctionsCIP, aggregate ownership, SAR basis, clocks, and secrecy stay distinct.
NO CUSTOMER SAR DISCLOSURE ↗ TAX / ABSENCE DETECTIONReturn CompletenessFind the one filing-year form or signature a complete-looking packet lacks.
NEVER SIGN OR TRANSMIT ↗
A right answer can still be a bad service.
Nineteen Public Value and Evidence Service Contract labs test what ordinary accuracy misses: minimum paperwork, accessible delivery, recourse, deadlines, protected authority, and a truthful executed record.
Choose the service failure you cannot afford.
Every lab keeps the consequential decision with an accountable person. What changes is the proof that a merely “correct” route still has to pass.
Minimum evidence, usable access, deadline, recourse, and truthful completion.
GENERAL PUBLIC-VALUE BASELINE ↗ 02 / ESSENTIAL SERVICEHousehold Energy LifelinePreserve the authorized continuity path before the shutoff clock wins.
CONTINUITY EXACT ↗ 03 / DISASTER RECOVERYClaim + Aid CoordinationKeep every known compensation source visible across separate ledgers.
SOURCE SET EXACT ↗ 04 / EMPLOYMENTUnemployment Claim NavigationProtect appeal and weekly-certification paths without deciding eligibility.
CLAIM PATH EXACT ↗ 05 / AGRICULTUREFarm Disaster DeadlinesMap every applicable crop, livestock, and grazing notice window.
DEADLINE SET EXACT ↗ 06 / CONSTRUCTIONPermit ReadinessBind the packet to the exact authority and rule—without claiming approval.
RULE PROVENANCE EXACT ↗ 07 / EDUCATIONStudent AccommodationCollect the minimum relevant evidence before qualified-team review.
SENSITIVE DATA MINIMIZED ↗Each lab holds eight archetypes and the scorecard constant while changing the beneficiary, trusted records, policy, terminal actions, and protected decision.
Prove each lot edge without widening a recall through invention.
TRACE LEDGER ↗ 09 / WATERDrinking-Water NoticeKeep unknown, sampled, noticed, delivered, and replaced states distinct.
NOTICE LEDGER ↗ 10 / TAXPAYER SERVICEIRS Notice ResponseBind the exact action and evidence to the notice clock and response route.
RESPONSE MAP ↗ 11 / VETERANSClaim EvidenceSeparate what is held from what is requested—without rating a claim.
EVIDENCE MAP ↗ 12 / ACCESSIBLE MOBILITYParatransit AccessPreserve accessible application, clock, trip-condition, and appeal service.
ACCESS CLOCK ↗ 13 / TRANSPARENCYFOIA Routing + AppealKeep component, disclosure, tracking, response, and appeal state intact.
DISCLOSURE ROUTE ↗ 14 / CITIZENSHIP SERVICEUSCIS Case EvidenceOrganize official case state and notices without predicting an outcome.
CASE LEDGER ↗ 15 / INTERNATIONAL TRADEExport EvidenceBind item, party, end use, destination, screening, and rule version.
EVIDENCE FIREWALL ↗ 16 / CHILD NUTRITIONSchool Meal AccessReuse direct certification before asking a family to prove it again.
CERTIFICATION REUSE ↗ 17 / ELECTION SERVICEProvisional Ballot StatusReport only official status and cure routes, privately and nonpartisan.
STATUS LEDGER ↗ 18 / CARE TRANSITIONSDischarge ReadinessA printed plan is not ready until the downstream handoffs are received.
READINESS GATE ↗ 19 / WORKFORCE MOBILITYLicense MobilityShow a compact or endorsement path without claiming licensure.
AUTHORITY MAP ↗The findings demos cannot show.
The best clean-world model became the most fragile when sources disagreed.
Open experiment ↗ ENVIRONMENT / TOOL GUARD+0.489Tool-layer enforcement gained 49 points; the prompt reminder doubled stalls.
Open experiment ↗ SECURITY / TOOL POISONING1.00 → 0.00Every model leaked through poisoned tooling. A dataflow gate stopped every leak.
Open experiment ↗Bring a workflow.
Leave with an eval.
Describe the work and its consequences. Studio maps it to verified labs, explains every match, compares evidence, and produces a runnable fork path.
Best verified matches
Don't just fork it.
Prove what changed.
Discover adaptations by the evidence they have earned. Every status comes from committed artifacts—not a badge selected by its author.
Evidence level, not endorsement. “Domain reviewed” records a named review scope. No level means regulator approval, production safety, certification, or permission to automate a protected decision.
Find your next evaluation.
Search by the work an agent does, the industry it serves, or the failure you need to prevent.