AGENT RELEASE OPERATIONS / CHANGE → EVIDENCE → DECISION

Ship evidence,
not confidence.

Capture the exact agent release, map every changed component to its impacted tests, block uncovered or failing changes, invite oracle-free outside reproduction, exchange privacy-bounded incident regressions, and watch the official sources beneath the evidence.

—exact cases —changed components —open blind challenges —public incident records —watched official sources —origin-verified Action commits
01 / THE RELEASE PATH

One change. Four proofs.

No aggregate score can substitute for knowing exactly what changed, which boundary it touches, whether the regression passed, and who still owns the consequential decision.

01Capture

Hash model, policy, tools, identity, authority, dependencies, egress, monitoring, and rollback.

02Test impact

Union before-and-after impact tags. An unmapped change or missing suite fails closed.

03Reproduce

Publish tasks without answers; bind submissions to the committed hidden-oracle digest.

04Exchange

Share safe regressions and machine-readable findings—not credentials, targets, or exploit steps.

02 / CHANGE IMPACT

Test what actually moved.

The reference candidate changes policy, tool, and authority components. The gate runs only mapped suites, demands both legitimate twins, and exports a non-certifying OSCAL Assessment Results view.

BYTE-EXACT RELEASE DIFF

— impacted evidence tags

EXPERIMENTAL OSCAL BRIDGE—observations + findingsMachine-readable assessment evidence. No certification, compliance determination, or Authorization to Operate.
03 / WORKFLOW DEPENDENCY TRUST

A pinned SHA still needs an origin.

The lock proves that every immutable-looking Action object resolves as a commit in the repository named by the workflow. Version 1.1 uses stable job-scoped locators, so harmless line shifts do not erase reviewed evidence.

EXACT ACTION USE INVENTORY

— stable job-scoped uses bound to — commits.

—verified signatures observed
—membership failures
LIVEGitHub commit API recheck
ORIGIN TRUST LOCK — SHA-256 commitment Observed —. Origin verification is not an Action code audit, safety finding, or availability guarantee.
04 / AGENT CAPABILITY + AUTHORITY BOM

Inventory the blast radius.

Models are only one part of an agent release. The AABOM binds tools, operations, resource scopes, expiring authority, delegation, data routes, revocation, accountable ownership, stop controls, and evidence—then surfaces widening without a trust score.

REFERENCE AUTHORITY DIFF / LOADING

— widening facts need owner review.

    MACHINE-READABLE INVENTORY

    Operational context around the model.

    —tools
    —leases
    —routes
    —evidence

    Exports an officially schema-validated CycloneDX — projection plus a deterministic, tamper-evident pack. Inventory is not current authorization, verified identity, certification, safety, compliance, deployment approval, or an ATO.

    LEAST-AUTHORITY PLANNER / PROPOSAL ONLY

    Observe. Challenge. Narrow only with proof.

    Compare privacy-bounded public, synthetic, or authorized aggregate action metadata with the exact AABOM. Exact tool-operation-scope relationships prevent accidental Cartesian authority; unobserved bindings become owner-review questions—not automatic policy edits.

    —bound relationships → —observed relationships → —relationship candidates → —auto-removed

    — separate proofs are required before a human-owned narrowing change. A bounded non-use window never proves that authority is unnecessary, and the planner emits no executable policy.

    AUTHORITY CONFORMANCE COMPILER / LOADING

    Turn every declared grant into a challenge.

    The inventory compiles exact relationships into legitimate clean twins and one-boundary violations—including valid operations and scopes that are unsafe in combination—then runs through an exact byte-bound command adapter without revealing expected answers or executing a tool.

    Run the protocol ↗
    —exact —clean twins —violations —unsafe allows

    — boundary shapes · — legitimate blocks · adapter SHA-256 —. A local byte match does not prove provenance, production enforcement, safety, certification, compliance, deployment approval, or an ATO.

    05 / FORK-TO-REPRODUCE

    The oracle is committed. The answers are not.

    — tasks across — public challenges are ready for an outside fork. Copy an oracle-free starter command below; GitHub attests submitted bytes, while independence and outcome still require human review. Accepted reproductions: — · campaign lock: —.

    06 / AGENT INCIDENT EXCHANGE

    Regression before rhetoric.

      — machine-readable exports: SARIF, OpenVEX, and explicitly experimental CSAF + OCSF bridges. Public or synthetic material only.

      07 / POLICY FRESHNESS

      Know when evidence ages.

      —explicit baselines
      —revision bindings
      —migration gap

      The weekly radar compares stable bytes or versioned visible-text fingerprints, then checks which exact revision each profile evaluated. MCP — has a —-case gate; A2A — has a —-case gate. It never interprets policy, edits a lab, or calls alignment conformance.

      Loading committed release evidence…

      PORTABLE AGENT ASSURANCE / MCP + A2A + TEVV

      Portable Agent Assurance.
      Bind proof to authority.

      One public, machine-verifiable envelope binds the agent, accountable operator, live authority lease, exact MCP or A2A action, and reproducible evidence. Eighteen clean and adversarial fixtures test where those bindings break.

      18protocol cases 18/18exact outcomes 2clean twins allowed 6TEVV blocks NOT VERIFIEDproduction identity
      CURRENT REVISION COMPANIONMCP 2026-07-28

      Issuer, credential, resource, audience, scope step-up, self-describing headers, and token transport—tested as recorded deltas, not protocol conformance.

      16/16exact cases 2clean twins 0unsafe allows 0legitimate blocks
      PINNED INTEROPERABILITY COMPANIONA2A 1.0 @ v1.0.1

      Version negotiation, selected interface tuple, PascalCase methods, exact card bytes, per-operation authority, and task scope—recorded evidence, not A2A conformance.

      17/17exact cases 2clean twins 0unsafe allows 0legitimate blocks
      CROSS-PROTOCOL AUTHORITY RELAYA2A 1.0 → MCP 2026-07-28

      Subject, actor, task, tenant, delegation, exact route, tool, resource, scope, audience, monitor, and approval continuity—tested at the hop neither protocol owns. 58/58 exact across the current CI matrix; 52 privacy-bounded spans carry zero raw content fields.

      25/25exact cases 2valid routes 0unsafe allows 0legitimate blocks
      01 / ASSURANCE CHAIN

      Six bindings. One inspectable decision.

      A token alone cannot prove that an agent still has authority. The envelope makes every consequential link explicit and the receipt chains every exact result.

      01Identity

      Verify the synthetic workload credential and normalized SPIFFE subject.

      02Operator

      Keep the accountable human or organization distinct from the agent.

      03Authority

      Bind lease, task, policy epoch, time window, revocation, and delegation.

      04Protocol

      Match MCP or A2A operation, resource, destination, and peer exactly.

      05Evidence

      Emit privacy-bounded events, exact reason codes, and source digests.

      06Reproduce

      Verify the pack, byte manifest, in-toto statement, and result chain.

      02 / PROTOCOL COLLISIONS

      Change one binding. Demand the right stop.

      Sixteen adversarial collisions and two legitimate twins distinguish safe refusal from blanket blocking. Every expected decision and reason code is committed before evaluation.

      EXACT REASON MATRIX

      Result chain loading

      Derived from the committed reference receipt.
        VISIBLE EVIDENCE GAPS

        What this release does not prove

          PRODUCTION IDENTITY / NOT VERIFIED

          03 / NIST AI 200-2 INITIAL PUBLIC DRAFT

          Four stages. No magic score.

          The experimental profile maps directly to the draft TEVV-Athlon stages while preserving planned events, revealed material, and missing independent reproduction as visible structure—not a false maturity number.

          LOCAL ENVELOPE INSPECTOR

          Inspect an envelope without uploading it.

          ZERO UPLOADThe browser reads only your local file. Structural checks are educational; use the Python verifier for signature and exact protocol evaluation.

          No local envelope selected · NIST draft comment deadline 2026-10-06

          AGENT SECURITY COMMONS / OPEN CONTROL EVIDENCE

          Authorization is a live condition.
          Prove it at runtime.

          Agent identity is not enough. Test whether authority survives delegation, changing policy, peer contact, tool calls, network destinations, monitor loss, safe-stop, and recovery—then preserve the outcome in a deterministic receipt.

          50runtime events 6adapter shapes 5defender kits 6incident regressions 12side-effect cases
          01 / RUNTIME CONFORMANCE

          One boundary. Every consequential event.

          The policy decision point evaluates normalized events from generic JSON, MCP, OpenAI Agents, LangGraph, CrewAI, and AutoGen-shaped recordings.

          IDENTITY

          Bind the actor

          Reject the wrong lease, audience, policy epoch, delegation depth, or revoked authority.

          TASK

          Bind the purpose

          Keep the task, state, time window, and human-granted scope explicit throughout the run.

          ACTION

          Bound the effect

          Check tools, targets, network destinations, peer messages, and record mutations before use.

          MONITOR

          Observe continuously

          Loss of the required monitor changes what the agent may do; it is not merely a log warning.

          STOP

          Fail closed

          Critical alerts, invalid sequences, and revoked leases move the runtime to a safe state.

          RECEIPT

          Make it replayable

          Exact reason codes, state transitions, source hashes, and a deterministic chain remain inspectable.

          02 / SIDE-EFFECT TRANSACTIONS

          A timeout is not permission to do it twice.

          Bind the exact intent to approval, authority, policy, and one idempotency key. An unknown result must be reconciled with the authoritative system before retry.

          UNKNOWN-OUTCOME PATH
          01PREPAREhash exact intent
          →
          02APPROVEhuman + expiry
          →
          03TIMEOUToutcome unknown
          →
          04RECONCILEthen replay or retry

          Same key + same intent replays the receipt. Same key + changed amount, target, or policy blocks. Compensation requires a separate approval and remains a second recorded effect.

          COMMITTED SYNTHETIC RECEIPTloading
          48ordered events 3duplicates prevented 2outcomes reconciled 1changed-intent conflict 0at-most-one breaches

          ORACLE-FREE COMMAND ADAPTER48/48 exact outcomes0 unsafe effects · 0 retry-after-unknown violations · receipt —

          TWO-PROCESS CRASH LAB12/12 exact recoveries across 6 crash points0 unsafe resumes · 3 uncertainties preserved

          MULTI-PROCESS RACE LAB12/12 exact races across 61 attempts0 duplicate effects · 0 missing legitimate effects

          COMPLETE CI MATRIX72/72 exact across 3 independent gates0 unsafe outcomes · 3 entrypoints + 8 static materials · 11 runtime-read digests across 109 processes, including 3 runtime-only policy reads · 42 unresolved import names exposed · fully stressed loading · matrix —

          RELEASE EVIDENCE BINDING1/1 consequential tool + operation + scope relationships boundrelease loading · static sets 3/3 match · runtime snapshots 3/3 match · authority twins 8/8 exact against adapter — · 0 holds · receipt —

          Adopt the complete CI gate ↗ Bind it to a release ↗
          03 / CONTROL EFFECTIVENESS

          Compare controls without a magic score.

          Switch between matched transparent synthetic arms. The cases stay fixed; only the declared controls change.

          Loading matched experiment…

          — controls
          —Unsafe allow rate
          —Exact outcomes
          —Legitimate actions kept

          Loading the source-bound case comparison…

          04 / ESSENTIAL SERVICES

          Small-operator kits with visible gaps.

          Each four-week kit protects accountable human authority, cites primary public guidance, and separates planned controls from evidenced ones.

          INCIDENT → REGRESSION

          Turn public lessons into tests that stay fixed.

          The Incident Regression Commons converts a source-bound incident record into exact pre-fix and post-fix cases, preserves a legitimate twin, and scans contributions for sensitive categories.

          Replay the public example ↗ Open the control experiment ↗
          COLLECTIVE CYBER DEFENSE LAB / FROM ADVISORY TO REPRODUCIBLE EVIDENCE

          Prove the fix.
          Prove the stop.

          A shared, open protocol for organizations, essential-service operators, governments, security partners, and frontier AI labs to test defensive work—without exposing targets, uploading inventories, or turning a synthetic run into a field-effectiveness claim.

          Public advisory moving through verified fix, containment, service continuity, and privacy-bounded evidence stages
          Reference executions are deterministic and public-safe. Independent reproduction remains a visible, unfilled evidence level.
          3verified-fix contracts 21containment events 20safe benchmark tasks 6mesh artifacts 0independent reproductions
          01 / OPEN DEFENSE STACK

          Seven modules. One evidence spine.

          Use only the layer you need, or carry one receipt from public signal through fix verification, runtime containment, continuity review, evidence exchange, and honest aggregation.

          02 / VERIFIED FIX COMMONS

          A patch is not a fix until the twin still works.

          Every contract checks the vulnerability regression, legitimate twin, service-continuity budget, rollback readiness, evidence, and accountable human owner.

          03 / STOP CONTROL + FRONTIER BENCHMARK

          Measure control boundaries separately.

          Stop latency is not correctness. Correctness is not evidence coverage. Neither proves service preservation. The receipts keep these signals distinct.

          Synthetic containment clock

          Pause authority40 ms
          Revoke parent20 ms
          Revoke children30 ms
          Cancel queued work35 ms
          Post-deadline breaches0

          Five defensive capability families

          04 / DEFENDER-IN-A-BOX

          Route a fix without uploading your inventory.

          Choose an AAU campaign JSON or load the fictional water-system example. The file stays inside this browser tab; no scan, network request, model call, or automatic change occurs.

          Local continuity-aware route planner

          Public, synthetic, or explicitly authorized inventory only

          ZERO UPLOADNothing has been loaded. Use the example button above or choose a local campaign.

          Your continuity-aware fix routes will appear here.

          05 / EVIDENCE MESH + OUTCOMES

          Share the receipt, not the sensitive world.

          Portable evidence retains artifact hashes, control fingerprints, bounded measurements, and evidence levels. The observatory counts artifacts—not organizations—and exposes what is still missing.

          CYBER DEFENSE EVIDENCE MESH

          Interoperable, privacy-bounded evidence

          Loading the committed evidence index…

          PUBLIC DEFENSE OUTCOMES OBSERVATORY

          No universal safety score

          Fix cases, containment events, defender decisions, and benchmark tasks stay separate. A reference-exact receipt is never promoted to field effectiveness.

          Loading visible evidence gaps…

          06 / BLIND REPRODUCTION EXCHANGE

          Make “independent” earn its evidence.

          An issuer commits a hidden oracle, an outside reproducer answers the gold-free challenge, and a separate reviewer reveals and adjudicates it. Bytes are proved by hash; organizational independence remains an explicit human-reviewed claim.

          ROLE-SEPARATED PROTOCOL / NIST-INSPIRED

          Challenge → run → review → reveal

          —committed demo status 0accepted outside contributions SUPPRESSEDpublic aggregate cell

          The committed walkthrough deliberately uses maintainer-simulated roles. It teaches the protocol after oracle reveal and never advances the independent count.

          ZERO-UPLOAD PREFLIGHT

          Inspect an adjudication locally.

          This browser checks the embedded digest and role gate. Full pack verification still requires the CLI because the oracle, receipt, statement, and manifest must recompute together.

          No local adjudication selected.

          What unlocks one accepted reproduction?
            Fork the full blind protocol ↗
            AGENT EVIDENCE STARTER / FROM DEMO TO TESTABLE CONTRACT

            Bring your agent.
            Leave with evidence.

            Turn an existing command or HTTP agent into a runnable evaluation project: synthetic cases, exact expected outcomes, forbidden-action checks, a protected human boundary, a privacy-bounded receipt, tests, documentation, and least-privilege CI.

            Inspect the starter source ↗
            3starter shapes 10local safety gates 11downloaded files 0accounts or uploads
            LOCAL WORKBENCH / SYNTHETIC OR PUBLIC INFORMATION ONLY

            Build the evidence spine

            01 / CHOOSE A TRANSFERABLE FAILURE SHAPE

            What kind of boundary are you testing?

            Start from a domain-neutral pattern, then adapt the synthetic cases to the deciding facts in your workflow.

            Step 1 of 4

            STRUCTURAL READINESS ≠ PRODUCTION VALIDATION. Passing the starter proves that an adapter can satisfy three reviewed synthetic cases under the declared contract. It does not demonstrate production safety, legal compliance, domain validity, model robustness, or authority to deploy.

            STARTER → RECEIPT → REUSABLE COMMUNITY EVIDENCE

            Your fork deserves
            a receipt.

            Turn a connected Agent Evidence Starter into an inspectable contribution pack: reviewed synthetic cases, aggregate receipts, a privacy scan, deterministic evidence checks, a SHA-256 manifest, and a share card—without uploading private traces.

            Read the evidence standard ↗
            —reference packs —public receipts —starting industries 0uploads or accounts
            01GeneratedConnected agent receipt + protected human authority
            02Domain reviewedAdapted 10-case suite + named scope + sources
            03ReproducedThree distinct public run receipts
            04VerifiedNamed different reproducer + linked receipt
            LOCAL CONTRIBUTION DESK / SYNTHETIC OR PUBLIC INFORMATION ONLY

            Inspect, derive, and export

            Nothing leaves this tab. Files are read in memory and form contents are neither uploaded nor persisted. Never select private details, prompts, raw model responses, credentials, classified information, CUI, or procurement-sensitive records.

            01 / LOAD PUBLIC ARTIFACTS
            02 / EXPLAIN THE REUSE VALUE
            03 / OPTIONAL HIGHER-LEVEL EVIDENCE
            04 / AUTHORIZE PUBLIC SHARING
            BUILT WITH AAU / INSPECT THE PROOF

            Reference packs you can fork

            These initial entries are explicitly labeled maintainer references. Community work will appear with contributor credit after the same validator and review route pass.

            Evidence level, not endorsement. The ladder summarizes present public artifacts. It does not authenticate people, prove domain correctness, certify a system, establish legal compliance, imply U.S. Government approval, or grant permission to deploy.

            MODEL TEST → RED TEAM → HUMAN CONTEXT

            An agent score answers
            half the question.

            Build a blinded, same-suite comparator for the existing human process. Measure exact outcome, abstention, task time, confidence calibration, and agreement—locally, without collecting names, free text, demographics, or production records.

            3separate test layers 8blinded practice tasks 5aggregate measures 0uploads or identifiers
            BLINDED PRACTICE / SYNTHETIC PUBLIC-SERVICE ROUTING

            Can you preserve the boundary?

            Ground truth stays hidden until all tasks are complete. Choose a route and declare confidence; task time is measured only in this tab.

            Ready · 0 of 8
            No timer running
            PRACTICE READY00 / 08

            Start when you are ready.

            The eight tasks contain only synthetic fields. No answer, response, timing, or form value leaves this browser tab.

            Choose the safest exact route
            No response is stored or transmitted.
            01Blindseparate study and key
            02Protectinstitution decides scope
            03Measureaccuracy, burden, uncertainty
            04Aggregatesessions never published
            05Comparedescriptive, not a worker ranking

            BASELINE ≠ REPLACEMENT DECISION. The lab does not verify institutional review, prove causal benefit, rank workers, support employment action, certify a system, or authorize deployment. Synthetic reference sessions test the protocol—not human performance.

            REVIEWED TASK → AGENT → HUMAN → PUBLIC VALUE → REPRODUCTION

            Stop at the score,
            or prove what changed.

            The Evidence Commons binds the artifacts needed to move from a synthetic result toward useful public evidence. It never hides a missing layer behind a badge, leaderboard, or trust score. Inspect three partner-ready pilots for FOIA routing, accessible digital services, and small-nonprofit grant administration.

            3open Impact Capsules 15visible evidence gaps 3bounded partner calls 0raw participant records
            01Reviewed suitePublic tasks, provenance, oracle, and byte hash
            02Agent receiptRepeated aggregate result, cost, latency, and suite binding
            03Human comparatorBlinded aggregate baseline after the right institutional determination
            04Public valuePredeclared service, burden, rights, and safety measures
            05ReproductionIndependent rerun with divergence and transfer conditions visible
            Loading evidence

            Impact Capsule

            Loading the public evidence chain.

            partner sought
            SYNTHETIC AGENT RESULT—loading
            OBSERVED HUMAN BASELINE—loading
            PUBLIC-VALUE OBSERVATION—loading

            Who should benefit

              Human authority stays with

              Authorized domain official

                Missing evidence—kept visible

                  Predeclared public-value measures

                  Loading the first honest next step.

                  EVIDENCE CHAIN ≠ APPROVAL. Status is derived from public artifact presence. AAU does not verify identity, institutional review, causal impact, production fitness, independence, certification, government endorsement, or authority to deploy or automate a protected decision.

                  FEDERAL MISSION STUDIO / EVIDENCE BEFORE AUTHORITY

                  Turn an AI idea into
                  inspectable evidence.

                  Map a public mission to current OMB acquisition and high-impact practices, NIST AI RMF functions, hard human boundaries, realistic tests, monitoring, remedies, and exit conditions. Export a 12-file assurance pack with a SHA-256 manifest.

                  Inspect the profile contract ↗
                  —mapped practices —official sources —exported files 0uploads
                  MISSION INTAKE / SYNTHETIC OR PUBLIC INFORMATION ONLY

                  Build the evidence spine

                  Form contents stay in this browser tab. The Studio does not upload, persist, or transmit them. Never enter classified, controlled, procurement-sensitive, source-selection, or personally identifiable information here.

                  01Mission and impactPurpose, people, baseline, measurable benefit
                  02Authority and dataHumans decide; evidence remains bounded
                  03Acquisition and testingPerformance, terms, portability, cost, transfer traps
                  04Oversight and stop conditionsIntervention, remedy, monitoring, cease use
                  PROFILE COMPLETENESS / NOT A COMPLIANCE SCORE
                  0/12

                  Name the mission to begin. Every unresolved item remains visible.

                  CONTROL CROSSWALKEvidence state by practice
                  RELEASE DESK

                  Export evidence, not a badge.

                  Complete the required mission fields and clear the local sensitive-data scan.

                  VERSIONED SOURCE DESKVerify policy before relying on the map
                  AGENCY → RESPONDER → REVIEWER / NO RANKING

                  Claims enter.
                  Gaps stay visible.

                  Fork a public mission, publish measurable gates, connect claims to evidence, recompute exact synthetic tests, preserve accountable decisions—and close the pilot with a bounded lesson that the next team can verify.

                  —reference pilots —measurable gates —synthetic cases —visible gaps
                  V0.3 / TRUST & PILOT READINESS

                  Verify the tool before it verifies a claim.

                  Security posture, release provenance, operating roles, and exit evidence now travel with the evaluation contract—without claiming federal approval.

                  01 / BOUNDHostile inputs stop early.

                  Byte, nesting, node, path, symlink, archive, and manifest limits are exercised in CI.

                  Inspect the threat model ↗
                  02 / ATTESTEDEvery release carries receipts.

                  Deterministic ZIP, exact manifest, SPDX 2.3 inventory, SHA-256 checks, and GitHub attestations.

                  Verify a release ↗
                  03 / 30 DAYSA pilot has an operating spine.

                  Eight working files take a team from mission scope and decision rights through tests, metrics, feedback, and exit.

                  Open the launch pack ↗
                  04 / HUMANEvidence never becomes authority.

                  No vendor rank, award recommendation, certification, ATO, or automated protected decision is produced.

                  Read the contract ↗
                  V0.4 / FEDERAL AI LESSONS EXCHANGE

                  A pilot ends.
                  The lesson travels.

                  Success, change, and stop decisions become evidence-linked public memory—with explicit non-transfer conditions, privacy review, and dated policy dependencies. No vendor leaderboard. No award recommendation.

                  Federal AI Lessons Exchange: observe, link, bound, verify, and reuse public pilot lessons
                  —public lessons —stopped pilots —policy links —bounded practices
                  PUBLIC CLOSEOUTSLoading…
                  ——

                  Loading the public exchange…

                  HUMAN DECISION

                  BOUNDED PRACTICE

                  USE WHERE
                    DO NOT TRANSFER TO
                      POLICY DEPENDENCIES
                      LOCAL PUBLICATION PREFLIGHT

                      Scan a lesson before it leaves your machine.

                      The browser checks structure, explicit sharing attestations, and narrow sensitive-data patterns. It never uploads, stores, or sends the file. A zero-finding result is not disclosure authorization.

                      No local lesson selected.

                      DATED SOURCE LEDGERReview dates are signals—not automatic legal conclusions.
                      CHOOSE A COMPLETE PUBLIC REFERENCE EXCHANGE

                      Inspect your fork locally

                      Choose three public or synthetic JSON files. Files stay in this browser tab—no upload, account, storage, analytics, or model call. The narrow sensitive-data scan is not a DLP system.

                      Loading the public reference exchange…

                      RECOMPUTED EVIDENCE LEDGER

                      Loading…

                      —
                      —tested requirements —exact cases —visible gaps —critical authority failures
                      CLAIM → EVIDENCE → TESTNo aggregate vendor score
                      EXACT CASE RECEIPTSOutcome + reasons + authority

                      A passing submitted synthetic case means only that the recorded output matches the declared oracle. It does not prove independent reproduction, production performance, compliance, or approval.

                      SOURCE-LINKED REVIEW DESK

                      Ask before writing terms.

                      These prompts surface acquisition questions. They are not clauses, legal advice, certification criteria, or a substitute for current agency review.

                      ZERO-DEPENDENCY ASSESSMENT + CLOSEOUT
                      python federal-pilot-kit/aau_pilot.py assess \
                        path/to/agency-intake.json \
                        path/to/vendor-response.json \
                        path/to/acceptance-tests.json
                      
                      python federal-pilot-kit/aau_pilot.py closeout \
                        path/to/agency-intake.json path/to/vendor-response.json \
                        path/to/acceptance-tests.json path/to/lesson.json \
                        --out /tmp/aau-public-lesson
                      V0.5 / FEDERAL AI PORTFOLIO OBSERVATORY / LOCAL-FIRST

                      See the portfolio.
                      Interrogate the evidence.

                      Turn a public or synthetic AI inventory into inspectable quality gaps, possible-overlap questions, bounded public-value measurements, three-layer TEV&V coverage, and testable acquisition obligations.

                      AAU Federal AI Portfolio Observatory: inventory, evidence quality, public value, testing, and accountable human decisions
                      —synthetic entries —documented —critical gaps —possible overlaps 0protected decisions
                      REFERENCE INVENTORY / SIX SYNTHETIC INVESTMENTS

                      Find the question behind the row.

                      PORTFOLIO SIGNALLoading…
                      ——

                      Loading the evidence map…

                      OPEN QUESTIONS
                      CANDIDATE AAU EVALUATION CONTRACTS
                      POSSIBLE OVERLAP / HUMAN REVIEW ONLY

                      Similarity opens a question. It never closes an investment.

                      LOCAL INVENTORY PREFLIGHT

                      Bring a public or synthetic portfolio.

                      Files stay in this browser tab. The Observatory never uploads, stores, or transmits them. Never include PII, credentials, controlled, classified, or procurement-sensitive information.

                      No local inventory selected.

                      PUBLIC VALUE LEDGER

                      Baseline first. Measurement second. Savings claim never inferred.

                      —
                      THREE-LAYER TEV&V

                      Test from three different distances.

                      AI CLAUSE TESTBENCH

                      An obligation is useful when failure is testable.

                      DATED OFFICIAL SOURCE LEDGERLoading…

                      A portfolio should be governable before it is impressive. Use the open JSON contracts in a fork, run the dependency-free CLI, and preserve every budget, acquisition, deployment, and retirement decision for accountable officials.

                      Run the CLI ↗Open the value ledger ↗
                      INTERACTIVE EVIDENCE PLAYGROUND / ZERO INSTALL

                      Can you trust
                      this agent?

                      Five real model traces. Inspect the case, the proposed action, and the evidence receipt. Choose Trust, Verify, or Block before the committed ground truth is revealed.

                      Audit how cases are built ↗
                      —real traces —industries —model IDs —source artifacts $0to investigate
                      CASE QUEUE / SELECT ANY TRACE

                      Read first. Judge second. Reveal last.

                      0/5 reviewed 0 exact
                      CASE — / — Loading evidence… —

                      Opening the evidence room…

                      The committed scenario and model trace are loading.

                      01 / DECIDING FACTSWhat the agent received
                      02 / EVIDENCE LEDGERPresent is not the same as required
                      03 / MODEL TRACEThe output under review
                      PROPOSED ACTION — —

                      Loading model reasoning…

                      MODEL— BACKEND— RUN— COST— LATENCY—

                      Boundary: synthetic scenarios and committed model evidence for education and evaluation. This is not production certification, legal advice, regulatory approval, or authority to automate a protected decision. Progress stays in your browser; nothing is submitted.

                      INTERACTIVE COUNTERFACTUAL LAB / ZERO INSTALL

                      Change one fact.
                      Watch the contract move.

                      Eight verified scenario pairs expose the smallest semantic boundary that changes the required action. Review both sides before the oracle appears, then download the pair as a regression test.

                      Audit the pair contract ↗
                      —verified boundaries —industries —contract shapes —source scenarios $0to reproduce
                      BOUNDARY QUEUE / THE SIDES ALTERNATE

                      Same-looking work. Different required action.

                      0/8 revealed 0/16 exact calls
                      BOUNDARY — / — Loading source evidence… —

                      Opening the split room…

                      One deciding fact means one declared semantic boundary—not one raw JSON field. Case IDs, prose, required evidence, and derived contract fields may change because that fact changed.

                      BEFORE—
                      DECLARED SEMANTIC Δ—
                      AFTER—
                      SCENARIO A—

                      Loading…

                      The source scenario is loading.

                      DECIDING RECORD
                      CONTRACT GATES
                      YOUR ACTION / SCENARIO A
                      SCENARIO B—

                      Loading…

                      The source scenario is loading.

                      DECIDING RECORD
                      CONTRACT GATES
                      YOUR ACTION / SCENARIO B

                      Choose one action for each scenario to unlock the boundary.

                      Boundary: synthetic scenarios, fictional records, and source-linked evaluation contracts for education and testing. This lab does not provide legal, medical, regulatory, safety, employment, benefits, or operational advice. It never grants authority to automate protected decisions.

                      COUNTERFACTUAL FORK GENERATOR / ZERO INSTALL

                      Bring a workflow.
                      Leave with a fork.

                      Define one semantic boundary, protect human authority, and export an eight-file contribution bundle with scenarios, tests, evidence notes, a Forge brief, and original artwork. The builder checks structure; qualified people still own domain truth.

                      Audit the builder contract ↗
                      —contract templates —local validation gates —files per bundle 0form values uploaded
                      DRAFT WORKBENCH / PRIVATE UNTIL YOU EXPORT

                      From operational question to testable pair.

                      Not saved
                      STEP 01 / WORKFLOW

                      What work deserves a boundary?

                      Describe the operational decision—not the model or prompt. Use synthetic, public, or safely abstracted information only.

                      Choose the closest evaluation shape *
                      Step 1 of 4

                      Boundary: this tool generates evaluation infrastructure from information you enter locally. It does not verify sources, determine policy, replace qualified review, inspect private systems, submit to GitHub, or make the resulting workflow safe for production.

                      EVALUATION EVIDENCE INSPECTOR / ZERO UPLOAD

                      Bring the receipt.
                      See what it proves.

                      Open an eval_*.json in your browser, recompute its trial grid, metrics, cost, latency, confidence-interval declarations, and model provenance, then export an aggregate-only receipt. No model call, account, or file upload.

                      —committed result artifacts —hard integrity checks —disclosure checks 0local files uploaded
                      PRIVATE INPUT DESK / NOTHING LEAVES THIS TAB

                      Inspect structure before believing the headline.

                      Use synthetic, public, or safely redacted results. The browser reads locally and deliberately excludes trial details from every exported receipt.

                      ↧ Drop one eval JSON here

                      Or choose a local eval_*.json up to 12 MB.

                      Choose local JSON READ IN MEMORY · NOT SAVED · NOT SENT No artifact selected
                      TEACHING RECEIPTS / SOURCE-BOUND TO COMMITTED ARTIFACTS
                      Loading the evidence contract…

                      Boundary: Receipt Lab checks a narrow JSON contract and recomputes declared aggregates. It does not validate domain truth, evaluate source quality, certify a system, approve a protected decision, establish regulatory compliance, or independently reproduce a run.

                      STATE OFAGENT
                      RELIABILITY
                      2026
                      OPEN RESEARCH RELEASE / LIVE EVIDENCE

                      Completion is not
                      correctness.

                      One reproducible view across every committed real-model result: exact task success, completion, safety boundaries, uncertainty, cost, latency, and provenance—without manufacturing a universal score.

                      Read the citable report ↗
                      —model evals —scenario trials —verified labs —observed failures —recorded spend
                      01 / STATUS GAP — points

                      Average distance between completion and exact task success on artifacts that report both.

                      Loading paired evidence…
                      02 / UNCERTAINTY — points

                      Median width of the committed 95% confidence interval for exact-task endpoints.

                      A SCORE WITHOUT ITS INTERVAL OVERSTATES THE RESULT
                      03 / TRANSFER FAILURE — labs

                      A valid rule from a clean twin is reused where one deciding fact reverses it.

                      SIMILARITY ERASES THE EXCEPTION
                      04 / MODEL IDENTITY — pinned

                      Model results need provenance. A floating alias can exercise different weights on the next run.

                      Loading provenance…
                      INTERACTIVE EVIDENCE WORKBENCH

                      Ask a narrower question.

                      Every bar names its source metric and opens the exact committed result. Filter first; compare only like work.

                      EXACT TASK SUCCESSLoading evidence…
                      0.501.00

                      Selected endpoints remain lab-specific. Use this view to inspect evidence, not to average unlike tasks.

                      MODEL TERRAIN / COVERAGE BEFORE RANK

                      Where was each model actually tested?

                      Medians summarize the current filtered portfolio. Uneven coverage is shown beside every number and is never filled with a zero.

                      FAILURE TERRAIN / REPOSITORY INCIDENCE

                      The patterns that travel.

                      Counts mean “observed in these labs,” not real-world prevalence. Industry and contract filters narrow the map.

                      Boundary: this is an automatically generated snapshot of this repository’s committed evidence. It is not a market ranking, production certification, regulator approval, or estimate of failure prevalence.

                      OPENAAU
                      CHALLENGE
                      01
                      COMMUNITY RELIABILITY LEAGUE / $0 TO ENTER

                      30 minutes. One claim.
                      Bring the receipt.

                      Choose a bounded mission. Reproduce a result, break an assumption, or adapt a proven contract to help different people. CI checks the evidence; the public board credits the work.

                      —live missions —community participants —community finishes —reference finishes —earned, never claimed
                      MISSION BOARD / FIVE STARTING POINTS

                      Pick the failure you want to make visible.

                      THE 30-MINUTE ROUTE

                      A small finish with a durable receipt.

                      The clock is a guide, not a gate. Adapt missions usually take longer; every step remains inspectable.

                      1. 00—05
                        ChooseClaim a mission and open its starter lab.
                      2. 05—12
                        RunExecute the deterministic mock—no key, no spend.
                      3. 12—20
                        TraceLink the claim to one exact scenario or branch.
                      4. 20—27
                        RecordCommit the result and a machine-readable submission receipt.
                      5. 27—30
                        ProveRun aau challenge validate and open the PR.
                      PUBLIC SCOREBOARD / EVIDENCE, NOT POINTS

                      Who returned with a receipt?

                      References are not community wins. They show the finish standard while the first independent contributors take the open positions.

                      ACHIEVEMENTS / DERIVED BY CI

                      No vanity badges.

                      SUBMISSION BUILDER / LOCAL-ONLY

                      Leave with the record already shaped.

                      This form downloads a Challenge receipt or Adapt Gallery starter in your browser. Nothing is sent, stored, or tracked.

                      Boundary: Challenge status is derived from repository evidence. It is not a model leaderboard, production certification, regulator approval, or permission to automate protected authority.

                      NEW / RIGHTS CONTINUITY CONTRACT

                      The appeal survives.
                      Does the protection?

                      One case can carry several rights on different clocks. The new matched suite catches the moment a technically valid main route silently loses coverage, urgent review, income, or another companion protection.

                      NEW / CRITICAL EVENT FAN-OUT

                      The response worked.
                      What stayed open?

                      Containment, initial reports, recipient notices, updates, and follow-ups remain independently true. A draft or successful response never closes the complete event graph.

                      PREVIOUS / REGULATORY CLOCK COLLISION BENCHMARK

                      One event.
                      Many clocks.

                      A consequential event rarely creates one tidy task. It creates independently timed duties for different recipients, through different channels, with protected human owners. This new matched benchmark catches the obligation an agent silently drops.

                      EVENT 00FACT PATTERN → DUTY 01WHOrecipient DUTY 02WHENclock origin DUTY 03HOWchannel DUTY 04PROOFreceipt NO DUTY DISAPPEARS
                      NEW / SEVEN-INDUSTRY PUBLIC PROTECTION WAVE

                      The action sounds right.
                      Where is the receipt?

                      A remedy, dispute, safety report, or regulator notification is not complete because an agent drafted it. The new matched suite proves the exact subject, current rule, gates, clock and channel, accountable owner, and the event that actually executed.

                      01SUBJECT∧02RULE∧03GATES∧04CLOCK + CHANNEL∧05OWNER∧06RECEIPTPROTECTED
                      NEW / PROOF BEFORE ACTION

                      Helpful is not
                      the same as proven.

                      Three new labs test the moment plausible context becomes an unsafe shortcut. The agent must bind the action to the exact source, emergency state, or current award—then stop before the protected human decision.

                      01RETRIEVE→02BIND→03VERIFY→04HAND OFFNO PROOF · NO ACTION
                      NEW / SIX-INDUSTRY DECISION GATE

                      Correct answer.
                      Wrong authority.

                      A new matched benchmark tests the last inch before consequential action: exact evidence, every satisfied gate, the rule-specific reason, procedural protections, human authority, and a record that matches what actually executed.

                      01OUTCOME∧02REASON∧03EVIDENCE∧04GATES∧05PROCEDURE∧06AUTHORITY∧07RECORD= EXACT
                      Small-business recovery pathway through evidence, accessibility, recourse, deadlines, rights, oversight, and a truthful record
                      NEW REPOSITORY SPECIALTY / PUBLIC VALUE

                      A right answer can still be a bad service.

                      Nineteen Public Value and Evidence Service Contract labs test what ordinary accuracy misses: minimum paperwork, accessible delivery, recourse, deadlines, protected authority, and a truthful executed record.

                      MINIMUM BURDENACCESSIBILITYRECOURSERIGHTS + RECORD TRUTH
                      NINETEEN PUBLIC-VALUE + EVIDENCE-SERVICE LABS

                      Choose the service failure you cannot afford.

                      Use the contract ↗

                      Every lab keeps the consequential decision with an accountable person. What changes is the proof that a merely “correct” route still has to pass.

                      NEW / MATCHED CROSS-INDUSTRY SUITETwelve services. One exact contract. No invented authority.

                      Each lab holds eight archetypes and the scorecard constant while changing the beneficiary, trusted records, policy, terminal actions, and protected decision.

                      08 / FOOD SAFETYRecall Traceability

                      Prove each lot edge without widening a recall through invention.

                      TRACE LEDGER ↗
                      09 / WATERDrinking-Water Notice

                      Keep unknown, sampled, noticed, delivered, and replaced states distinct.

                      NOTICE LEDGER ↗
                      10 / TAXPAYER SERVICEIRS Notice Response

                      Bind the exact action and evidence to the notice clock and response route.

                      RESPONSE MAP ↗
                      11 / VETERANSClaim Evidence

                      Separate what is held from what is requested—without rating a claim.

                      EVIDENCE MAP ↗
                      12 / ACCESSIBLE MOBILITYParatransit Access

                      Preserve accessible application, clock, trip-condition, and appeal service.

                      ACCESS CLOCK ↗
                      13 / TRANSPARENCYFOIA Routing + Appeal

                      Keep component, disclosure, tracking, response, and appeal state intact.

                      DISCLOSURE ROUTE ↗
                      14 / CITIZENSHIP SERVICEUSCIS Case Evidence

                      Organize official case state and notices without predicting an outcome.

                      CASE LEDGER ↗
                      15 / INTERNATIONAL TRADEExport Evidence

                      Bind item, party, end use, destination, screening, and rule version.

                      EVIDENCE FIREWALL ↗
                      16 / CHILD NUTRITIONSchool Meal Access

                      Reuse direct certification before asking a family to prove it again.

                      CERTIFICATION REUSE ↗
                      17 / ELECTION SERVICEProvisional Ballot Status

                      Report only official status and cure routes, privately and nonpartisan.

                      STATUS LEDGER ↗
                      18 / CARE TRANSITIONSDischarge Readiness

                      A printed plan is not ready until the downstream handoffs are received.

                      READINESS GATE ↗
                      19 / WORKFORCE MOBILITYLicense Mobility

                      Show a compact or endorsement path without claiming licensure.

                      AUTHORITY MAP ↗
                      FIELD NOTES / CONTROLLED EXPERIMENTS

                      The findings demos cannot show.

                      All 17 failure patterns ↗
                      AAU STUDIO / LOCAL-FIRST MATCHER

                      Bring a workflow.
                      Leave with an eval.

                      Describe the work and its consequences. Studio maps it to verified labs, explains every match, compares evidence, and produces a runnable fork path.

                      Runs in this browser · no account · no upload · no tracking
                      04 What must not go wrong?
                      Deterministic matching · inspectable evidence
                      THE VERIFIED CATALOG

                      Find your next evaluation.

                      Search by the work an agent does, the industry it serves, or the failure you need to prevent.

                      I want to
                      Loading verified use cases… RUNNABLE CODE · COMMITTED RESULTS · OBSERVED FAILURES
                      EVIDENCE COMPARISON

                      Compare shapes, not unlike scores.

                      Headline metrics differ across workflows. This view compares reusable architecture, verification depth, and observed failure evidence—not a universal leaderboard.