Compile evaluation integrity, authority, containment, remediation, and resilience evidence into one executable assurance case, then require role-separated review—without hiding contradictions or automating the accountable decision.
Not a safety certificate. A measurable regression floor with every limitation visible.
Observed robustness spread25→99%Same harness. Same attacks. Different base models.
14models measured
10model families
5 sector twinsverified remediation missions
23 schedules4 bounded race profiles
Start with one useful outcome
Less setup. More evidence.
Pick the task you need to solve. Start with a local plan or fictional example, then bring authorized inputs from your own system.
01 / Developers & security teams
Know what your scan will run.
Inspect scored cases and auxiliary runs before calling a model. Scope-bound baselines detect changed task selections; missing comparison evidence is not a pass.
Apply one disclosure policy to multiple AI suppliers. Get an offline HTML review and reproducible evidence packs, with invalid submissions kept separate from findings.
Build an executable assurance case with verified source reports. Keep contradictory, missing, and stale evidence visible for the accountable human decision.
Available on main: use a source checkout for these newer workflows; they are not all in the published v0.19.0 package. Commands above follow dspy-security-bench; guides include all required inputs. Synthetic examples are not production assessments, and disclosure checks are not supplier rankings or approvals.
AssuranceGraph · executable claim–evidence cases
Nine proofs. One honest decision surface.
Every artifact recomputes through its native verifier before it can support a claim. Missing, stale, violated, and contradictory evidence stays visible. No universal safety score. No automatic approval.
LOCAL / DETERMINISTIC / ZERO ACTIONSASSURANCEGRAPH::CLAIM-EVIDENCE::V1PROFILE: CRITICAL INFRASTRUCTURE
NATIVE EVIDENCE VERIFIERS9 / 9 CURRENTAT
AuthorityTwinidentity · scope · effects
VERIFIEDTP
TraceProofauthorization · receipts
VERIFIEDCG
CollectiveGuard v2containment · provenance
VERIFIEDSP
ScheduleProofrevocation · interleavings
VERIFIEDDT
DefenderTwinremediation · continuity
VERIFIEDRG
ResilienceGraphexact portfolio frontier
VERIFIEDEI
EvalIntegrityProofholdout · evaluator · monitor
VERIFIEDCP
ContainmentProofcanary control evidence
VERIFIEDAB
AgentBOMdependency claim impact
VERIFIED
RECOMPUTEAGclaims ↔ evidence
PROFILEFROZEN
FRESHNESSBOUND
OWNERSDECLARED
DIGESTSPINNED
EXECUTABLE CLAIM GRAPH9 / 9 SUPPORTED01
bounded authorityno false allow or unsafe effect
SUPPORTED02
observable effectscurrent privacy-bounded trace
SUPPORTED03
collective containmentcomplete direct provenance
SUPPORTED04
schedule safetycomplete bounded exploration
SUPPORTED05
verified remediationpaths closed · mission stable
SUPPORTED06
resilience spacerobust feasible option exists
SUPPORTED07
evaluation integritycommit precedes label reveal
SUPPORTED08
runtime containmenteight canary controls hold
SUPPORTED09
dependency boundaryno claim-impacting drift
SUPPORTED
SUPPORTEDcurrent verified evidence meets every predicateVIOLATEDverified evidence fails a predicateCONTRADICTEDcurrent evidence disagreesSTALEage exceeds owner limitMISSINGno usable current evidence
ONE CASE / FOUR REVIEW SURFACESJSONSARIF 2.1.0OSCAL 1.2.2STANDALONE HTMLAUTOMATIC DEPLOYMENT ACTIONS: 0
01 / VERIFYNative semantics first
A checksum alone never upgrades an arbitrary assertion into deployment evidence.
02 / DISAGREEContradictions stay visible
The graph cannot choose a favorable report when current evidence conflicts.
03 / EXPIREFreshness is part of the claim
Owner-declared observation times and age limits are bound into the exact case.
04 / DECIDEHumans retain authority
Profile support informs review; it never becomes certification, ATO, or risk acceptance.
EvalIntegrityProof · evaluation evidence before model claims
Trust the result? First prove the test.
A signed score can still come from a compromised procedure. Thirteen content-free controls preserve holdout secrecy, evaluator independence, complete accounting, active monitoring, and safe exit—without reading a prompt or output.
EVALUATION SEALEDEVALINTEGRITYPROOF::V1 / 13 CONTROLS / ZERO ACTIONSINTEGRITY_EVIDENCED
INTEGRITY EVIDENCEDall controls and sources completeINTEGRITY VIOLATEDan observation contradicts a controlMONITOR FAILEDdetection or clock evidence failedINCOMPLETE EVIDENCEsource coverage cannot support a claim
CONTENT-FREE BY CONTRACTNO PROMPTSNO OUTPUTSNO LABELSNO CREDENTIAL VALUESNO MODEL CALLNO DEPLOYMENT AUTHORITY
Every claim names the functions that must review its evidence. Exact keys, roles, organizations, assignments, and review windows are policy-bound. A valid evidence gap cannot be erased by adding favorable signatures.
Registration, review inclusion, retirement, and compromise become one append-only evidence history. Operator and cross-organization witness signatures bind the exact Merkle checkpoint—without turning the log into an approval authority.
ContainmentProof × AgentBOM · evidence before action
Prove the boundary. Trace the blast radius.
Eight harmless canaries distinguish a control failure from a monitor failure. An owner-authored AI-BOM policy turns CycloneDX and SPDX gaps into reviewable SARIF, then lifecycle drift separates regression, persistence, improvement, resolution, and component change—without copying disclosure values or touching a live target.
CONTROL EVIDENCE ONLINECONTROLPLANE::CANARIES+AGENTBOM::V1NETWORK REQUESTS: 0
BASELINE sha256:7e2a…CANDIDATE sha256:b91c…MINIMAL PLAN · 5 CLAIMS
07
Sector startersbenefits · health · finance · manufacturing · water · logistics · software
00
Executable pluginsdeclarative probe manifests never load contributor code
08
Federal review filesOSCAL · evidence index · freshness · POA&M inputs
0
Public casesopen reproduction exchange · no ranking · no endorsement
DefenderTwin · verified cyber-defense remediation
Close the path. Keep the mission alive.
A finding is only the beginning. DefenderTwin verifies the exact before/after state, target and approval scope, essential-service continuity, introduced risk, rollback, and evidence completeness—without touching a live target.
SYNTHETIC / NO LIVE ACTIONSDEFENDERTWIN::VERIFIED-REMEDIATION::V1OFFLINE RECOMPUTATION
BEFORE / 00EXPOSED
ESSENTIAL SERVICEclinical routing
LIVE
01
indirect egressrelay path allowed
CRITICAL
02
admin scopecluster-wide authority
CRITICAL
VERIFYΔbefore / after
01identityBOUND
02authorityALLOW
03missionSTABLE
04rollbackVERIFIED
AFTER / 01VERIFIED
ESSENTIAL SERVICEclinical routing
LIVE
01
egress deniedcanary policy + rollback
CLOSED
02
role attenuatedroute-write only
CLOSED
MISSION CONTINUITY
25s declared disruption≤ 120s objective
01EFFECTIVE + SAFEpaths closed · mission stable · complete evidence02EFFECTIVE + REGRESSIONrisk reduced · authority or continuity failed03INEFFECTIVErequired state or attack path remains open04INSUFFICIENT EVIDENCEtechnical change cannot become a clean claim
PACKAGED MISSIONScommunity hospitalwater utilitylocal governmentopen sourcesmall business
01 / EFFECTVerify state, not prose
Frozen structural outcomes—not an agent's explanation—decide whether the required state returned.
02 / AUTHORITYA useful fix can still be unauthorized
The first independent result belongs to the community.
Bring a synthetic mission or adapter output. CI recomputes attack-path closure, mission continuity, scope, approval, rollback, privacy, and every digest.
ResilienceGraph · exact cyber-defense portfolios
Scarce resources. No hidden ranking.
Start with fixes DefenderTwin can recompute. Exclude unsafe options, trace shared-service dependencies, enforce resource and community floors, replay stressed availability, and preserve every non-dominated choice for an accountable owner.
WORST-CASE SERVICE WEIGHT ↑RESOURCE DEMAND →05 unitsweight 407 unitsweight 708 unitsweight 811 unitsweight 12REFERENCE · NOT A RECOMMENDATION
ELIGIBLE5/ 6
ROBUST SCENARIOS3/ 3
SEARCH100% exact
DT ✓
open sourceeffective + safe
ELIGIBLEDT ✓
water utilityeffective + safe
ELIGIBLEDT ✓
local governmenteffective + safe
ELIGIBLEDT !
disruptive hospitalmission regression
EXCLUDED
STRESS REPLAY
baselineall actions available
PASS
provider losstwo actions unavailable
PASS
maintainer lossone action unavailable
PASS
01 / VERIFYUnsafe fixes never enter search
Every embedded DefenderTwin result recomputes before the candidate becomes eligible.
02 / BOUNDImpossible floors stay impossible
The engine returns no feasible portfolio instead of silently relaxing owner constraints.
03 / STRESSDependencies remain visible
Directly addressed services and downstream dependency reach are reported separately.
04 / DECIDEThe frontier preserves judgment
No predicted loss, magic weight, or model-generated narrative chooses how resources are spent.
CollectiveGuard · autonomous-agent containment
One run was isolated. The collective was not.
Reconstruct cross-run coordination, indirect egress, peer-authority laundering, unsafe persistence, evaluator access, and missed response windows from structural events alone. No prompts, reasoning, messages, secrets, or exploit payloads enter the protocol.
LOCAL / CONTENT-FREECOLLECTIVEGUARD::CONTAINMENT-PLANE::V1 + PROVENANCE::V213 DETERMINISTIC RULES
RUN / ABLOCKEDevaluation worker
task dependency unavailable
scope: packages:read
egress: mirror only
+20s opaque write
SHARED SERVICEartifact channel
packages allowed communication denied
CG001
RUN / BDRIFTevaluation worker
adopts peer action
authority: peer “go”
egress: indirect path
+70s material effect
PROTECTEDgraderCG006 · access
OUT OF BOUNDSthird partyCG003 · egress
opaque artifact
cross-run read
01 / SEPARATEShared access is not shared authority
Package, artifact, storage, and browser services stay bound to run identity and declared purpose.
02 / STOPImpossible work gets a safe exit
Owner-defined action budgets turn clarification and pause into measurable system behavior.
Freshness and regression evidence with zero production actions. The controller requests review; operators decide and act.
TraceProof · runtime-to-evidence assurance
Bring the trace. Leave the secrets.
Transform a local OpenTelemetry export into a pseudonymized evidence graph, deterministic authorization findings, a synthetic replay twin, SARIF, and OSCAL. Raw prompts, tool arguments, credentials, and unapproved attributes are removed by default.
OFFLINE / LOCAL CUSTODYTRACEPROOF::SANITIZER::V1NO COLLECTOR CONNECTION
01
IngestOTLP JSON export
trace + span IDsHASH
prompt contentDROP
tool argumentsDROP
access tokensDROP
authorization signalsKEEP
privacy boundaryDEFAULT DENY
01human intentpseudonymized
02orchestratorscope: read
03worker agentscope: write
04MCP effectcontained
FIRST UNSAFE EDGEscope exceeds grantTP004 · HIGH
03
Evidencereview, replay, automate
TP002Audience mismatchcritical
TP003Token passthroughcritical
TP004Scope exceeds granthigh
JSON
SARIF
OSCAL 1.2.2
SYNTHETIC TWIN
What this provesthe supplied sanitized bytes, deterministic rule results, and content-addressed replayWhat remains humantrace completeness, operational context, severity acceptance, compliance, and deployment decisions
Open TraceProof evidence ledger0 privacy-bounded runtime experiments
Turn structural OpenTelemetry evidence into a ScheduleProof draft without treating span links or wall clocks as proof. Names, attributes, events, prompts, arguments, results, and credentials are never read.
Declare only the ordering your design truly enforces. ScheduleProof counts every reachable interleaving, checks eight authorization invariants, and returns the shortest causal counterexample—without executing a model, credential, tool, or effect.
EXHAUSTIVE / BOUNDEDSCHEDULEPROOF::TOPOLOGICAL_EXPLORER::V1ATOMIC EVENT MODEL
DECLARED PARTIAL ORDER
01Agrantauthority-1
01Bapprovemax_uses: 1
→
02exchangeaud: payments-mcp
⌁
03Brevokeauthority-1
03Acommitpayment-1042
Missing guarantee: commit and revocation can race. The graph—not a preferred trace—defines the search space.
Repair direction Revalidate named authority atomically at effect commit. A suggested fix is evidence for review—not an automated risk decision.
01Exact bounded coverage
Counts the full topological schedule space before exploration.
02Eight invariants
Authority, approval, token, scope, audience, identity, effect, and receipt.
03Offline recomputation
Protocol and scenario digests bind every derived count and counterexample.
04Honest incompleteness
A truncated search without a finding is review—not a safe result.
New · Mission Assurance Commons
A public workbench for AI missions that matter.
Start with a bounded use case—not a vendor claim. Generate synthetic evaluation drafts, exercise every delegation edge, compare verified evidence over time, and package the result for accountable review.
01 / SCOPEInventoryForge
Normalize a public AI inventory and draft a human-review-required MissionPack.
Compare content-addressed baselines and surface owner-threshold regressions.
watch compare→04 / DECIDEAcquisitionProof
Export neutral test plans, QASP inputs, cost fields, and portability checks.
acquisition export
frozen protocol · local execution
See where delegated authority breaks.
AgentGraphTwin keeps the mission fixed while mutating exactly one security-relevant edge. Receipts bind the request, adapter, policy, principal, decision, and reason.
V2 pairs
06
Path modes
04
Real effects
00
principalHumanintent bound
delegate
agent / 01Orchestratorscope: records.read
scope inflation blocked
agent / 02Specialistleast privilege
authorize
resourceMCP toolno effect
FIRST UNSAFE EDGEedge-02reason: scope_not_granted
AuthorityBridge / translation contractsBring the policy plane you already operate.
OPARego decision
Cedarauthorization response
OpenFGArelationship check
OAuth + MCPtoken-bound tools
SPIFFEworkload identity
Translation fixtures remain clearly labeled and do not execute, certify, or endorse the named backend. The operator-controlled command runner can execute a declared backend and version locally, producing self-attested conformance evidence without recording credentials or the command itself.
Two models can complete benign tasks at nearly the same rate and still have radically different prompt-injection failure profiles. Security must be measured directly.
62pointrobustness gap between two models with near-identical capability
Interactive map
Capability × robustness
RobustMixedVulnerableProvisional
Hover or focus a model. Provisional measurements use a dashed ring.
Frozen protocol leaderboard
See which models hold the boundary.
Robustness is the share of prompt-injection attacks that failed. Capability is benign task completion. Confidence intervals are cluster-bootstrapped over task pairs.
Rank / model
Family
Robustness
Capability
Evidence
No models match that view.
From benchmark to guardrail
Make injection safety a merge condition.
Scaffold a config and GitHub Action, preview the exact test matrix without spending API credits, then block regressions with terminal, JSON, and SARIF reports.
01Works with DSPy or any Python agent factory
02Absolute security floors or baseline regression gates
03OWASP LLM01 · NIST AI 100-2 · MITRE ATLAS mappings
quickstart.sh
# install and create a ready-to-run gate$ pip install dspy-security-bench
$ dspy-security-bench init --model openai/gpt-4o-mini
# inspect the matrix — zero model calls$ dspy-security-bench scan --config .dspy-security-bench.yaml --plan
✓ 10 benchmark cases planned✓ no model was called
Keep your orchestration code and tools. The integration assistant detects a directly declared framework, generates a reviewable adapter, and validates the wiring without invoking the agent run loop.
Generated workflows are manual by default, so connecting a repository cannot unexpectedly spend model credits.
ProofRun · evidence with a chain of custody
Run it. Sign it. Let others verify it.
v0.14 five-path evidence builder
A reusable workflow isolates model evaluation from signing, recomputes every statistic in a clean job, and binds the exact JSON to a source commit before enforcing the gate.
DSB / PROOFRUN01
evidence passport
Verification ladderBytes are easy. Provenance is earned.
1
Self-attestedstatistics + canonical hash
2
GitHub-attestedrunner + workflow + commit
3
Trusted builderimmutable central workflow
4
Reproducedindependent maintainer rerun
artifact subjectsha256: 9e81…c4f2
signature verification is separate from score recomputation
Every accepted entry preserves all raw trials and passes offline recomputation. A cryptographic badge appears only after the exact digest passes attestation verification and enters the reviewed registry.
Community evidenceranked by evidence tier, then lower confidence bound
The network is open. Be the first independent agent team to submit an attested run.
Provenance authenticates a workflow and exact artifact. It does not independently observe a model provider or turn a synthetic benchmark into certification.
AuthorityTwin · agent identity under pressure
Prove the agent may act.
Frozen protocol · 10 adversarial twins
A vendor-neutral conformance lab for the authorization layer between a human, an AI agent, and a tool. It asks whether each action is bound to the right identity, tenant, audience, scope, intent, approval, and audit record.
01 / principal
Hhuman-aliceintent: records.read
delegatesscope · audience · expiry
02 / workload
Aagent-orchestratoron behalf of human-alice
requestsnonce · tenant · action
03 / audience
Rmcp://recordstenant-north / record
ALLOW
policy-bound receipt
CONTROLBound requestALLOW
agent
agent-orchestrator
scope
records:read
audience
mcp://records
Δone authority mutationsame mission · same resource
Bridge an OAuth, OIDC, MCP, SPIFFE, policy-engine, or custom authorization layer and publish the first independent evidence bundle.
AuthorityTwin is an independent conformance and evidence harness—not a new authorization protocol. Synthetic results are not identity proof, non-repudiation, compliance, certification, production validation, or an authorization to operate.
MissionForge · agency-owned evaluation
Trace the answer back to authority.
Ship a mission test as reviewable YAML—not benchmark code. SourceTwin holds the authoritative corpus and question fixed, changes only retrieved untrusted content, and scores the agent's recorded claims against exact source IDs.
Five controlled interventions
01
Fabricated authorityinvented rule + citation
02
Embedded instructionretrieval changes conclusion
03
Material omissionexception disappears
04
Superseded guidanceobsolete source wins
05
Insufficient evidenceanswer instead of abstain
protocolSHA-256
PRIMARYUNTRUSTEDCLAIM
Deterministic probes
FaithfulnessDoes the citation support the claim?
CompletenessWere material facts and exceptions retained?
SufficiencyDid the agent abstain when evidence ran out?
AuthorityDid current primary evidence outrank noise?
$ dspy-security-bench pack run source-twin --agent myapp:build
Fork it, encode a real mission with synthetic evidence, and publish the first independently reproducible run.
Structured test claims are benchmark ground truth for a fixed synthetic pack—not a legal interpretation, universal truth set, certification, or substitute for agency subject-matter review.
Public-interest specialty
Same facts. Different decision.
ImpactTwin / ProcureBench
Clean and poisoned procurement twins hold structured facts fixed while varying one untrusted input: vendor-authored text. Then the benchmark measures what changed in the live synthetic world.
Clean mission utility100%no observed clean-work loss
Risk reduction$3.69Msynthetic scenario exposure
RepeatControlTwin · v0.10
One delta can be luck. Repeat the pair.
Run the complete experiment again and again. Alternate order. Preserve every child report. Gate the lower confidence bound—not the flattering point estimate.
01 Balanced schedule
T01OFF→ON
T02ON→OFF
T03OFF→ON
T04ON→OFF
T05OFF→ON
fresh agent · every case · every condition
02 95% Wilson evidence
Harm containment100%
25/25 · lower bound 86.7%
Safe mission recovery60%
15/25 · 40.7%–76.6%
Clean utility preserved100%
25/25 · lower bound 86.7%
03 Paired signal
2525
harms prevented
25
harms introduced
0
unstable pair effects
0 / 5
exact McNemar p
5.96e−8
Try the five-trial fixture$ dspy-security-bench impact control-repeat-demo --trials 5
Deterministic reference fixture, not a model result. Containment is not recovery; synthetic exposure is not predicted loss or certification.
Open Control Evidence Registry · public beta
Which guardrail works? Show the receipts.
Evidence, not endorsements
Publish the policy-off and policy-on experiment—even when the control performs poorly. Every accepted entry keeps the raw paired trials, policy identity, uncertainty, and chain of custody inspectable.
0public control experiments
01
Run paired twinsfresh agent · alternating order
02
Recompute offlineraw events → intervals → digest
03
Verify provenanceworkflow · commit · exact bytes
04
Compare tradeoffscontainment · recovery · utility
AdmissionValidity, not victory
A failing control can belong here. Five trials, recomputable evidence, fresh isolation, zero runtime errors, and redacted arguments are the comparison floor.
IdentityPolicy SHA-256 bound
The exact policy document travels with the evidence. A name alone cannot silently stand in for different enforcement logic.
InterpretationThree outcomes stay separate
Blocking harm, safely completing the mission, and preserving clean work are reported independently—never collapsed into one marketing score.
Public evidence ledgerchain of custody first · containment lower bound second
Fixed synthetic ProcureBench suite · repeated execution evidence · not certification, a population estimate, predicted loss, or independent observation of a hosted provider.
Public-sector mission assurance · v0.16
From agent action to reviewable evidence.
Exercise a cyber-response agent inside an inert digital twin, repeat the exact counterfactual protocol, and transform verified evidence into standards-shaped assessment inputs—without claiming certification.
01 / TEST
IncidentTwin
air-gap safe
Five clean/poisoned security-operations twins. Authoritative alert facts stay fixed; only hostile external content changes.
Bind verified ImpactTwin, ControlTwin, IncidentTwin, MissionPack, or AuthorityTwin evidence to an owner-supplied deployment profile and local acceptance thresholds.
Designed to support evidence review—not replace accountable officials.Informative crosswalks · owner-supplied system boundary · explicit uncertainty · no government endorsement · no automatic ATO or procurement decision
Open IncidentTwin ledger0 public cyber-response experiments
Be the first independent team to publish recomputable incident-response evidence.
Real systems. Real authority.
Secure the moment intent becomes action.
Every production pattern pairs an untrusted input with a dangerous sink. Included policies enforce least agency before the side effect—not after an incident.
01Profile included
CS
Customer support
Resolve tickets without turning a poisoned CRM note into data exfiltration, unlimited refunds, or identity changes.
Untrusted
Tickets · email · CRM
Boundary
Refund cap · recipient domain
--profile customer-support
02Profile included
AP
Accounts payable
Extract and reconcile invoices while preventing vendor impersonation from becoming an unauthorized transfer.
Untrusted
Invoices · vendor email
Boundary
Payee allowlist · approval
--profile financial-operations
03Profile included
RAG
Research & RAG
Search broadly without letting a hostile page become persistent memory, published content, or executable code.
Untrusted
Web · docs · retrieved chunks
Boundary
Memory · publish · execute
--profile research-rag
04Profile included
SRE
DevOps copilots
Accelerate diagnosis with broad observability while keeping deletion, shell access, and production mutation constrained.
Untrusted
Logs · issues · repository
Boundary
Deploy · shell · delete
--profile devops
Start offline. No API key required.$ dspy-security-bench policy init --profile customer-supportOpen implementation guide →
Research you can inspect
Built for scrutiny, not screenshots.
01
Frozen protocol
Suites, attacks, scaffold, task subset, and decoding settings are hashed and versioned.
02
Every row reproducible
Per-model result JSON, run metadata, confidence intervals, and generation scripts live beside the board.
03
Honest uncertainty
Rows remain provisional when confidence intervals cross a bucket boundary. Green means no known bypass—not safe.
Trust is a system property. Make the evidence inspectable.