AGENT EVIDENCE STARTER / FROM DEMO TO TESTABLE CONTRACT

Bring your agent.
Leave with evidence.

Turn an existing command or HTTP agent into a runnable evaluation project: synthetic cases, exact expected outcomes, forbidden-action checks, a protected human boundary, a privacy-bounded receipt, tests, documentation, and least-privilege CI.

Inspect the starter source ↗
3starter shapes 10local safety gates 11downloaded files 0accounts or uploads
LOCAL WORKBENCH / SYNTHETIC OR PUBLIC INFORMATION ONLY

Build the evidence spine

01 / CHOOSE A TRANSFERABLE FAILURE SHAPE

What kind of boundary are you testing?

Start from a domain-neutral pattern, then adapt the synthetic cases to the deciding facts in your workflow.

Step 1 of 4

STRUCTURAL READINESS ≠ PRODUCTION VALIDATION. Passing the starter proves that an adapter can satisfy three reviewed synthetic cases under the declared contract. It does not demonstrate production safety, legal compliance, domain validity, model robustness, or authority to deploy.

STARTER → RECEIPT → REUSABLE COMMUNITY EVIDENCE

Your fork deserves
a receipt.

Turn a connected Agent Evidence Starter into an inspectable contribution pack: reviewed synthetic cases, aggregate receipts, a privacy scan, deterministic evidence checks, a SHA-256 manifest, and a share card—without uploading private traces.

Read the evidence standard ↗
reference packs public receipts starting industries 0uploads or accounts
01GeneratedConnected agent receipt + protected human authority
02Domain reviewedAdapted 10-case suite + named scope + sources
03ReproducedThree distinct public run receipts
04VerifiedNamed different reproducer + linked receipt
LOCAL CONTRIBUTION DESK / SYNTHETIC OR PUBLIC INFORMATION ONLY

Inspect, derive, and export

Nothing leaves this tab. Files are read in memory and form contents are neither uploaded nor persisted. Never select private details, prompts, raw model responses, credentials, classified information, CUI, or procurement-sensitive records.

01 / LOAD PUBLIC ARTIFACTS
02 / EXPLAIN THE REUSE VALUE
03 / OPTIONAL HIGHER-LEVEL EVIDENCE
04 / AUTHORIZE PUBLIC SHARING
BUILT WITH AAU / INSPECT THE PROOF

Reference packs you can fork

These initial entries are explicitly labeled maintainer references. Community work will appear with contributor credit after the same validator and review route pass.

Evidence level, not endorsement. The ladder summarizes present public artifacts. It does not authenticate people, prove domain correctness, certify a system, establish legal compliance, imply U.S. Government approval, or grant permission to deploy.

FEDERAL MISSION STUDIO / EVIDENCE BEFORE AUTHORITY

Turn an AI idea into
inspectable evidence.

Map a public mission to current OMB acquisition and high-impact practices, NIST AI RMF functions, hard human boundaries, realistic tests, monitoring, remedies, and exit conditions. Export a 12-file assurance pack with a SHA-256 manifest.

Inspect the profile contract ↗
mapped practices official sources exported files 0uploads
MISSION INTAKE / SYNTHETIC OR PUBLIC INFORMATION ONLY

Build the evidence spine

Form contents stay in this browser tab. The Studio does not upload, persist, or transmit them. Never enter classified, controlled, procurement-sensitive, source-selection, or personally identifiable information here.

01Mission and impactPurpose, people, baseline, measurable benefit
02Authority and dataHumans decide; evidence remains bounded
03Acquisition and testingPerformance, terms, portability, cost, transfer traps
04Oversight and stop conditionsIntervention, remedy, monitoring, cease use
PROFILE COMPLETENESS / NOT A COMPLIANCE SCORE
0/12

Name the mission to begin. Every unresolved item remains visible.

CONTROL CROSSWALKEvidence state by practice
RELEASE DESK

Export evidence, not a badge.

Complete the required mission fields and clear the local sensitive-data scan.

VERSIONED SOURCE DESKVerify policy before relying on the map
AGENCY → RESPONDER → REVIEWER / NO RANKING

Claims enter.
Gaps stay visible.

Fork a public mission, publish measurable gates, connect claims to evidence, recompute exact synthetic tests, preserve accountable decisions—and close the pilot with a bounded lesson that the next team can verify.

reference pilots measurable gates synthetic cases visible gaps
V0.3 / TRUST & PILOT READINESS

Verify the tool before it verifies a claim.

Security posture, release provenance, operating roles, and exit evidence now travel with the evaluation contract—without claiming federal approval.

01 / BOUNDHostile inputs stop early.

Byte, nesting, node, path, symlink, archive, and manifest limits are exercised in CI.

Inspect the threat model ↗
02 / ATTESTEDEvery release carries receipts.

Deterministic ZIP, exact manifest, SPDX 2.3 inventory, SHA-256 checks, and GitHub attestations.

Verify a release ↗
03 / 30 DAYSA pilot has an operating spine.

Eight working files take a team from mission scope and decision rights through tests, metrics, feedback, and exit.

Open the launch pack ↗
04 / HUMANEvidence never becomes authority.

No vendor rank, award recommendation, certification, ATO, or automated protected decision is produced.

Read the contract ↗
V0.4 / FEDERAL AI LESSONS EXCHANGE

A pilot ends.
The lesson travels.

Success, change, and stop decisions become evidence-linked public memory—with explicit non-transfer conditions, privacy review, and dated policy dependencies. No vendor leaderboard. No award recommendation.

Federal AI Lessons Exchange: observe, link, bound, verify, and reuse public pilot lessons
public lessons stopped pilots policy links bounded practices
PUBLIC CLOSEOUTSLoading…

Loading the public exchange…

HUMAN DECISION

BOUNDED PRACTICE

USE WHERE
    DO NOT TRANSFER TO
      POLICY DEPENDENCIES
      LOCAL PUBLICATION PREFLIGHT

      Scan a lesson before it leaves your machine.

      The browser checks structure, explicit sharing attestations, and narrow sensitive-data patterns. It never uploads, stores, or sends the file. A zero-finding result is not disclosure authorization.

      No local lesson selected.

      DATED SOURCE LEDGERReview dates are signals—not automatic legal conclusions.
      CHOOSE A COMPLETE PUBLIC REFERENCE EXCHANGE

      Inspect your fork locally

      Choose three public or synthetic JSON files. Files stay in this browser tab—no upload, account, storage, analytics, or model call. The narrow sensitive-data scan is not a DLP system.

      Loading the public reference exchange…

      RECOMPUTED EVIDENCE LEDGER

      Loading…

      tested requirements exact cases visible gaps critical authority failures
      CLAIM → EVIDENCE → TESTNo aggregate vendor score
      EXACT CASE RECEIPTSOutcome + reasons + authority

      A passing submitted synthetic case means only that the recorded output matches the declared oracle. It does not prove independent reproduction, production performance, compliance, or approval.

      SOURCE-LINKED REVIEW DESK

      Ask before writing terms.

      These prompts surface acquisition questions. They are not clauses, legal advice, certification criteria, or a substitute for current agency review.

      ZERO-DEPENDENCY ASSESSMENT + CLOSEOUT
      python federal-pilot-kit/aau_pilot.py assess \
        path/to/agency-intake.json \
        path/to/vendor-response.json \
        path/to/acceptance-tests.json
      
      python federal-pilot-kit/aau_pilot.py closeout \
        path/to/agency-intake.json path/to/vendor-response.json \
        path/to/acceptance-tests.json path/to/lesson.json \
        --out /tmp/aau-public-lesson
      V0.5 / FEDERAL AI PORTFOLIO OBSERVATORY / LOCAL-FIRST

      See the portfolio.
      Interrogate the evidence.

      Turn a public or synthetic AI inventory into inspectable quality gaps, possible-overlap questions, bounded public-value measurements, three-layer TEV&V coverage, and testable acquisition obligations.

      AAU Federal AI Portfolio Observatory: inventory, evidence quality, public value, testing, and accountable human decisions
      synthetic entries documented critical gaps possible overlaps 0protected decisions
      REFERENCE INVENTORY / SIX SYNTHETIC INVESTMENTS

      Find the question behind the row.

      PORTFOLIO SIGNALLoading…

      Loading the evidence map…

      OPEN QUESTIONS
      CANDIDATE AAU EVALUATION CONTRACTS
      POSSIBLE OVERLAP / HUMAN REVIEW ONLY

      Similarity opens a question. It never closes an investment.

      LOCAL INVENTORY PREFLIGHT

      Bring a public or synthetic portfolio.

      Files stay in this browser tab. The Observatory never uploads, stores, or transmits them. Never include PII, credentials, controlled, classified, or procurement-sensitive information.

      No local inventory selected.

      PUBLIC VALUE LEDGER

      Baseline first. Measurement second. Savings claim never inferred.

      THREE-LAYER TEV&V

      Test from three different distances.

      AI CLAUSE TESTBENCH

      An obligation is useful when failure is testable.

      DATED OFFICIAL SOURCE LEDGERLoading…

      A portfolio should be governable before it is impressive. Use the open JSON contracts in a fork, run the dependency-free CLI, and preserve every budget, acquisition, deployment, and retirement decision for accountable officials.

      Run the CLI ↗Open the value ledger ↗
      INTERACTIVE EVIDENCE PLAYGROUND / ZERO INSTALL

      Can you trust
      this agent?

      Five real model traces. Inspect the case, the proposed action, and the evidence receipt. Choose Trust, Verify, or Block before the committed ground truth is revealed.

      Audit how cases are built ↗
      real traces industries model IDs source artifacts $0to investigate
      CASE QUEUE / SELECT ANY TRACE

      Read first. Judge second. Reveal last.

      0/5 reviewed 0 exact
      CASE — / — Loading evidence…

      Opening the evidence room…

      The committed scenario and model trace are loading.

      01 / DECIDING FACTSWhat the agent received
      02 / EVIDENCE LEDGERPresent is not the same as required
      03 / MODEL TRACEThe output under review
      PROPOSED ACTION

      Loading model reasoning…

      MODEL BACKEND RUN COST LATENCY

      Boundary: synthetic scenarios and committed model evidence for education and evaluation. This is not production certification, legal advice, regulatory approval, or authority to automate a protected decision. Progress stays in your browser; nothing is submitted.

      INTERACTIVE COUNTERFACTUAL LAB / ZERO INSTALL

      Change one fact.
      Watch the contract move.

      Eight verified scenario pairs expose the smallest semantic boundary that changes the required action. Review both sides before the oracle appears, then download the pair as a regression test.

      Audit the pair contract ↗
      verified boundaries industries contract shapes source scenarios $0to reproduce
      BOUNDARY QUEUE / THE SIDES ALTERNATE

      Same-looking work. Different required action.

      0/8 revealed 0/16 exact calls
      BOUNDARY — / — Loading source evidence…

      Opening the split room…

      One deciding fact means one declared semantic boundary—not one raw JSON field. Case IDs, prose, required evidence, and derived contract fields may change because that fact changed.

      BEFORE
      DECLARED SEMANTIC Δ
      AFTER
      SCENARIO A

      Loading…

      The source scenario is loading.

      DECIDING RECORD
      CONTRACT GATES
      YOUR ACTION / SCENARIO A
      SCENARIO B

      Loading…

      The source scenario is loading.

      DECIDING RECORD
      CONTRACT GATES
      YOUR ACTION / SCENARIO B

      Choose one action for each scenario to unlock the boundary.

      Boundary: synthetic scenarios, fictional records, and source-linked evaluation contracts for education and testing. This lab does not provide legal, medical, regulatory, safety, employment, benefits, or operational advice. It never grants authority to automate protected decisions.

      COUNTERFACTUAL FORK GENERATOR / ZERO INSTALL

      Bring a workflow.
      Leave with a fork.

      Define one semantic boundary, protect human authority, and export an eight-file contribution bundle with scenarios, tests, evidence notes, a Forge brief, and original artwork. The builder checks structure; qualified people still own domain truth.

      Audit the builder contract ↗
      contract templates local validation gates files per bundle 0form values uploaded
      DRAFT WORKBENCH / PRIVATE UNTIL YOU EXPORT

      From operational question to testable pair.

      Not saved
      STEP 01 / WORKFLOW

      What work deserves a boundary?

      Describe the operational decision—not the model or prompt. Use synthetic, public, or safely abstracted information only.

      Choose the closest evaluation shape *
      Step 1 of 4

      Boundary: this tool generates evaluation infrastructure from information you enter locally. It does not verify sources, determine policy, replace qualified review, inspect private systems, submit to GitHub, or make the resulting workflow safe for production.

      EVALUATION EVIDENCE INSPECTOR / ZERO UPLOAD

      Bring the receipt.
      See what it proves.

      Open an eval_*.json in your browser, recompute its trial grid, metrics, cost, latency, confidence-interval declarations, and model provenance, then export an aggregate-only receipt. No model call, account, or file upload.

      committed result artifacts hard integrity checks disclosure checks 0local files uploaded
      PRIVATE INPUT DESK / NOTHING LEAVES THIS TAB

      Inspect structure before believing the headline.

      Use synthetic, public, or safely redacted results. The browser reads locally and deliberately excludes trial details from every exported receipt.

      Drop one eval JSON here

      Or choose a local eval_*.json up to 12 MB.

      Choose local JSON READ IN MEMORY · NOT SAVED · NOT SENT No artifact selected
      TEACHING RECEIPTS / SOURCE-BOUND TO COMMITTED ARTIFACTS
      Loading the evidence contract…

      Boundary: Receipt Lab checks a narrow JSON contract and recomputes declared aggregates. It does not validate domain truth, evaluate source quality, certify a system, approve a protected decision, establish regulatory compliance, or independently reproduce a run.

      STATE OFAGENT
      RELIABILITY
      2026
      OPEN RESEARCH RELEASE / LIVE EVIDENCE

      Completion is not
      correctness.

      One reproducible view across every committed real-model result: exact task success, completion, safety boundaries, uncertainty, cost, latency, and provenance—without manufacturing a universal score.

      Read the citable report ↗
      model evals scenario trials verified labs observed failures recorded spend
      01 / STATUS GAP points

      Average distance between completion and exact task success on artifacts that report both.

      Loading paired evidence…
      02 / UNCERTAINTY points

      Median width of the committed 95% confidence interval for exact-task endpoints.

      A SCORE WITHOUT ITS INTERVAL OVERSTATES THE RESULT
      03 / TRANSFER FAILURE labs

      A valid rule from a clean twin is reused where one deciding fact reverses it.

      SIMILARITY ERASES THE EXCEPTION
      04 / MODEL IDENTITY pinned

      Model results need provenance. A floating alias can exercise different weights on the next run.

      Loading provenance…
      INTERACTIVE EVIDENCE WORKBENCH

      Ask a narrower question.

      Every bar names its source metric and opens the exact committed result. Filter first; compare only like work.

      EXACT TASK SUCCESSLoading evidence…
      0.501.00

      Selected endpoints remain lab-specific. Use this view to inspect evidence, not to average unlike tasks.

      MODEL TERRAIN / COVERAGE BEFORE RANK

      Where was each model actually tested?

      Medians summarize the current filtered portfolio. Uneven coverage is shown beside every number and is never filled with a zero.

      FAILURE TERRAIN / REPOSITORY INCIDENCE

      The patterns that travel.

      Counts mean “observed in these labs,” not real-world prevalence. Industry and contract filters narrow the map.

      Boundary: this is an automatically generated snapshot of this repository’s committed evidence. It is not a market ranking, production certification, regulator approval, or estimate of failure prevalence.

      OPENAAU
      CHALLENGE
      01
      COMMUNITY RELIABILITY LEAGUE / $0 TO ENTER

      30 minutes. One claim.
      Bring the receipt.

      Choose a bounded mission. Reproduce a result, break an assumption, or adapt a proven contract to help different people. CI checks the evidence; the public board credits the work.

      live missions community participants community finishes reference finishes earned, never claimed
      MISSION BOARD / FIVE STARTING POINTS

      Pick the failure you want to make visible.

      THE 30-MINUTE ROUTE

      A small finish with a durable receipt.

      The clock is a guide, not a gate. Adapt missions usually take longer; every step remains inspectable.

      1. 00—05
        ChooseClaim a mission and open its starter lab.
      2. 05—12
        RunExecute the deterministic mock—no key, no spend.
      3. 12—20
        TraceLink the claim to one exact scenario or branch.
      4. 20—27
        RecordCommit the result and a machine-readable submission receipt.
      5. 27—30
        ProveRun aau challenge validate and open the PR.
      PUBLIC SCOREBOARD / EVIDENCE, NOT POINTS

      Who returned with a receipt?

      References are not community wins. They show the finish standard while the first independent contributors take the open positions.

      ACHIEVEMENTS / DERIVED BY CI

      No vanity badges.

      SUBMISSION BUILDER / LOCAL-ONLY

      Leave with the record already shaped.

      This form downloads a Challenge receipt or Adapt Gallery starter in your browser. Nothing is sent, stored, or tracked.

      Boundary: Challenge status is derived from repository evidence. It is not a model leaderboard, production certification, regulator approval, or permission to automate protected authority.

      NEW / RIGHTS CONTINUITY CONTRACT

      The appeal survives.
      Does the protection?

      One case can carry several rights on different clocks. The new matched suite catches the moment a technically valid main route silently loses coverage, urgent review, income, or another companion protection.

      NEW / CRITICAL EVENT FAN-OUT

      The response worked.
      What stayed open?

      Containment, initial reports, recipient notices, updates, and follow-ups remain independently true. A draft or successful response never closes the complete event graph.

      PREVIOUS / REGULATORY CLOCK COLLISION BENCHMARK

      One event.
      Many clocks.

      A consequential event rarely creates one tidy task. It creates independently timed duties for different recipients, through different channels, with protected human owners. This new matched benchmark catches the obligation an agent silently drops.

      EVENT 00FACT PATTERN DUTY 01WHOrecipient DUTY 02WHENclock origin DUTY 03HOWchannel DUTY 04PROOFreceipt NO DUTY DISAPPEARS
      NEW / SEVEN-INDUSTRY PUBLIC PROTECTION WAVE

      The action sounds right.
      Where is the receipt?

      A remedy, dispute, safety report, or regulator notification is not complete because an agent drafted it. The new matched suite proves the exact subject, current rule, gates, clock and channel, accountable owner, and the event that actually executed.

      01SUBJECT02RULE03GATES04CLOCK + CHANNEL05OWNER06RECEIPTPROTECTED
      NEW / PROOF BEFORE ACTION

      Helpful is not
      the same as proven.

      Three new labs test the moment plausible context becomes an unsafe shortcut. The agent must bind the action to the exact source, emergency state, or current award—then stop before the protected human decision.

      01RETRIEVE02BIND03VERIFY04HAND OFFNO PROOF · NO ACTION
      NEW / SIX-INDUSTRY DECISION GATE

      Correct answer.
      Wrong authority.

      A new matched benchmark tests the last inch before consequential action: exact evidence, every satisfied gate, the rule-specific reason, procedural protections, human authority, and a record that matches what actually executed.

      01OUTCOME02REASON03EVIDENCE04GATES05PROCEDURE06AUTHORITY07RECORD= EXACT
      Small-business recovery pathway through evidence, accessibility, recourse, deadlines, rights, oversight, and a truthful record
      NEW REPOSITORY SPECIALTY / PUBLIC VALUE

      A right answer can still be a bad service.

      Nineteen Public Value and Evidence Service Contract labs test what ordinary accuracy misses: minimum paperwork, accessible delivery, recourse, deadlines, protected authority, and a truthful executed record.

      MINIMUM BURDENACCESSIBILITYRECOURSERIGHTS + RECORD TRUTH
      NINETEEN PUBLIC-VALUE + EVIDENCE-SERVICE LABS

      Choose the service failure you cannot afford.

      Use the contract ↗

      Every lab keeps the consequential decision with an accountable person. What changes is the proof that a merely “correct” route still has to pass.

      NEW / MATCHED CROSS-INDUSTRY SUITETwelve services. One exact contract. No invented authority.

      Each lab holds eight archetypes and the scorecard constant while changing the beneficiary, trusted records, policy, terminal actions, and protected decision.

      08 / FOOD SAFETYRecall Traceability

      Prove each lot edge without widening a recall through invention.

      TRACE LEDGER ↗
      09 / WATERDrinking-Water Notice

      Keep unknown, sampled, noticed, delivered, and replaced states distinct.

      NOTICE LEDGER ↗
      10 / TAXPAYER SERVICEIRS Notice Response

      Bind the exact action and evidence to the notice clock and response route.

      RESPONSE MAP ↗
      11 / VETERANSClaim Evidence

      Separate what is held from what is requested—without rating a claim.

      EVIDENCE MAP ↗
      12 / ACCESSIBLE MOBILITYParatransit Access

      Preserve accessible application, clock, trip-condition, and appeal service.

      ACCESS CLOCK ↗
      13 / TRANSPARENCYFOIA Routing + Appeal

      Keep component, disclosure, tracking, response, and appeal state intact.

      DISCLOSURE ROUTE ↗
      14 / CITIZENSHIP SERVICEUSCIS Case Evidence

      Organize official case state and notices without predicting an outcome.

      CASE LEDGER ↗
      15 / INTERNATIONAL TRADEExport Evidence

      Bind item, party, end use, destination, screening, and rule version.

      EVIDENCE FIREWALL ↗
      16 / CHILD NUTRITIONSchool Meal Access

      Reuse direct certification before asking a family to prove it again.

      CERTIFICATION REUSE ↗
      17 / ELECTION SERVICEProvisional Ballot Status

      Report only official status and cure routes, privately and nonpartisan.

      STATUS LEDGER ↗
      18 / CARE TRANSITIONSDischarge Readiness

      A printed plan is not ready until the downstream handoffs are received.

      READINESS GATE ↗
      19 / WORKFORCE MOBILITYLicense Mobility

      Show a compact or endorsement path without claiming licensure.

      AUTHORITY MAP ↗
      FIELD NOTES / CONTROLLED EXPERIMENTS

      The findings demos cannot show.

      All 17 failure patterns ↗
      AAU STUDIO / LOCAL-FIRST MATCHER

      Bring a workflow.
      Leave with an eval.

      Describe the work and its consequences. Studio maps it to verified labs, explains every match, compares evidence, and produces a runnable fork path.

      Runs in this browser · no account · no upload · no tracking
      04 What must not go wrong?
      Deterministic matching · inspectable evidence
      THE VERIFIED CATALOG

      Find your next evaluation.

      Search by the work an agent does, the industry it serves, or the failure you need to prevent.

      I want to
      Loading verified use cases… RUNNABLE CODE · COMMITTED RESULTS · OBSERVED FAILURES
      EVIDENCE COMPARISON

      Compare shapes, not unlike scores.

      Headline metrics differ across workflows. This view compares reusable architecture, verification depth, and observed failure evidence—not a universal leaderboard.