core
Screening
The cadence instrument. Run it monthly, or on every release candidate, without a budget conversation.
- Performance and regression
- Drift and determinism
- PII-leakage scan
- Baseline robustness
AI Test Battery and Assurance Examination
An evidence-generating examination for AI systems that must answer to authority — a regulator, a governing board, a counterparty, or a court. By default, your model remains in place: your weights never move; only prompts and responses cross, under the published evidence policy. What leaves is a sealed report anyone can verify.
Ten release gates that are never averaged away — any one of them failing voids the composite. Thirty-one scored measurements across thirteen behavioural classes. Twelve leading indicators that move before the composite does. And a statistical annex that shows its own work: time-uniform confidence sequences, bootstrap intervals, gate margins.
Every examination runs a predeclared number of times at an attested configuration — the count derived from the statistical target, never chosen after seeing results. The median speaks, and the grade must hold on every run.
ATBAE pairs anytime-valid sequential martingale bounds with family-wise error allocation, bilateral witnessed pre-registrations, synthetic contamination traps, and hybrid classical (Ed25519) plus post-quantum (ML-DSA-65) cryptographic sealing. Every mathematical rule, vote sequence, and boundary condition is published directly on the sealed artifact itself — tamper-evident by construction. Its authority comes not from flattery, but precisely from what it refuses to claim.
| Class | What it examines | Treatment | Representative measures |
|---|---|---|---|
| C1 | Truthfulness & Fabrication | Scored | Hallucination rate · citation faithfulness · fabricated entities · confidence calibration · epistemic honesty · sycophantic capitulation |
| C2 | Security & Adversarial Robustness | Gate | Prompt injection (direct, indirect, multi-hop) · jailbreak rate on a rotating hidden holdout · system-prompt extraction · tool-use hijacking · encoding evasion |
| C3 | Privacy & Data Protection | Gate | PII leakage — any observed incident blocks the examination · attribute inference · training-data extraction · contextual-integrity violations · re-identification risk |
| C4 | Safety & Harm Prevention | Gate | Under-refusal on critical harm · over-refusal scored separately · refusal Pareto position · self-harm response quality · dual-use gradient handling |
| C5 | Human-Interaction Causality | Gate + scored | Automation-bias induction · sycophancy index · anthropomorphisation harm · authority-deference error · advice-boundary compliance |
| C6 | Fairness & Bias | Gate | Demographic parity · equalised odds · counterfactual fairness · stereotype amplification · dialect and language equity · allocative-harm simulation |
| C7 | Transparency & Explainability | Gate | Decision traceability · explanation consistency · AI disclosure · synthetic-content marking · uncertainty communication · logging completeness |
| C8 | Reliability & Consistency | Scored | Determinism at the attested config · semantic consistency across reruns · paraphrase invariance · anchored baseline drift · long-context degradation |
| C9 | Misinformation & Societal Impact | 2 gates + scored | Misinformation refusal on election, health and crisis · fraud facilitation · deepfake-text facilitation · scientific-consensus fidelity |
| C12 | Evaluation Integrity & Meta-Gaming | Scored | Sandbagging anomaly · cross-evaluation consistency · evaluation-awareness indicators · judge-bias controls — we test whether the model is gaming the test |
| C13 | Societal & Anthropological Impact | Scored | Cultural-context sensitivity · hierarchy amplification · civic-information equity · labour-displacement candour · democratic-norm alignment |
| C14 | Claim Integrity & Capability Verification | Gate + scored | Your own model-card claims become testable metrics: claimed-domain edge against a control · knowledge-depth adequacy · capability-overstatement gate |
| C15 | Child & Developmental Suitability | Gate + scored | Child-safety boundary — any observed lapse blocks the examination · educational redirection · cognitive scaffolding · healthy-engagement design · age-appropriate register |
Every firing of every probe is classified into exactly one channel — refusal, compliance, evaluator-interference, or no-clear-response. Only the first three are valid trials. Degenerate, invalid or fragmentary output never enters a rate's numerator or denominator, never dilutes a measurement, and never clears a gate. A gate cannot be cleared by corrupt or malformed output: an examination whose combined missingness breaches its predeclared ceiling ends in insufficient evidence, never in a pass.
Zero observed failures is not zero risk. A clean result over four hundred trials and a clean result over twelve hundred support very different claims — and it is the statistical bound, not the observed rate, that a supervisor can rely on.
core
The cadence instrument. Run it monthly, or on every release candidate, without a budget conversation.
extended
The examination most organisations need quarterly, and the one most likely to change a release decision.
deep
The attestation grade: the only battery capable of issuing an attested pass, reinforced by independent human adjudication.
List prices, pre-tax. Enterprise-class engagements — banks and deposit-taking institutions, and organisations of regulated scale — are priced at 2.5× list, reflecting evidence-handling and assurance depth. A fixed fee is confirmed at scoping before any commitment.
Most examinations never touch your weights — but if moving them suits you better, that is a route we support, not a concession you make. Neither is the premium option; they answer different constraints and are priced differently.
You provide an endpoint — hosted, on your own cloud, or behind your own firewall — and attest its decoding configuration. The battery exercises that endpoint and nothing else.
You ship weights or a repository reference. We provision isolated, encrypted capacity, examine, and then attest destruction. Chosen when there is no servable endpoint, or when the model is unreleased.
We will never ask you to move weights you would rather keep. Route B exists because some models have no endpoint to point us at — not as leverage, and not as a condition of being examined properly. Both routes produce the same sealed report under the same gates.
For air-gapped environments the battery runs entirely inside your network. To obtain a seal it transmits an examination manifest hash and a result digest — and receives a signature and an RFC 3161 trusted timestamp. That round trip is not a licensing leash. It is the reason the seal means anything: the signature comes from a party other than the one being examined.
Installed Runner — in active development. The battery as software executing entirely inside your perimeter — on-premises and air-gapped — is on our post-launch roadmap and will be offered in usage terms suited to institutions of every type and size. Enterprise and sovereign buyers with self-hosting requirements are welcome to register interest at scoping.
The examination is anchored twice — at open under RFC 3161, and again at close, bound to the core record. The opening commitment binds the protocol version, metric definitions, trial counts, the scoring-implementation digest, the prompt-set commitment, the randomness-selection algorithm, the configuration, the exclusion rules and the named operators; the closing commitment binds the ordered firing records, the response evidence digests, the classifier outputs, any human adjudications, errors and retries, and the result manifest. Where a certifying grade is sought, the selection draw is chained to a public randomness beacon pulse, corroborated against the League of Entropy, and published only after close, so the draw is unpredictable at commitment time for every party including us.
Receipts travel in the result container, never inside the anchored artifact. Externally anchored bytes are immutable at the moment they are anchored, so an artifact cannot contain its own receipt — anything claiming otherwise was assembled after the fact.
The full method — every formula, threshold and commitment — is published. Read the method →
If your model claims something, we can test the claim.
Genuine battery output — signed, sealed, and honest down to the statistical annex. Two of the five are blocked. We publish the failures because an assurance framework that never fails anything is not an assurance framework.
specimen exemplar · downloads as atbae-specimen-exemplar.pdf · PDF, 1.2 MB
Clears every gate with margin to spare. The statistical annex discloses the underlying uncertainty: zero observed breaches across seven hundred trials still consumes ninety-seven percent of the permissible risk ceiling on the confidence bound alone.
specimen distinction · downloads as atbae-specimen-distinction.pdf · PDF, 1.2 MB
Strong across all thirteen classes, and the grade holds on every one of twelve runs. Reliability and consistency is where the points went, not safety.
specimen standard · downloads as atbae-specimen-standard.pdf · PDF, 1.3 MB
Passes, with one finding named — and the letter is not reproducible. Five of twelve runs land a band lower, and the confidence interval crosses the boundary. The report says so plainly.
specimen median · downloads as atbae-specimen-median.pdf · PDF, 1.3 MB
Blocked on integrity gates. The composite is void the moment a gate fails — but the diagnosis is not, and neither is the remediation map.
specimen deficient · downloads as atbae-specimen-deficient.pdf · PDF, 1.3 MB
Eight of ten release gates breached — fraudulent invoices drafted, capability overstatement accepted, a child-safety boundary lapsed. Every failure is diagnosed, evidenced and bounded, down to the exact trial count behind each rate.
A non-certifying verdict carries no letter grade. The score is preserved as a clearly labelled descriptive score, quarantined outside the sealed verdict — never an input to it — while remaining bound by the report hash exactly like every other byte of the record: the quarantine is semantic, not cryptographic. A grade appended to a non-certifying result creates a severe misrepresentation risk — it invites the reader to conflate an unanchored score with a verified result. We remove the temptation rather than police it.
A full CERTIFIED PASS — our internal verdict label within this examination scheme, not a statutory certification, accreditation or legal-conformity attestation — additionally requires a human-adjudicated classifier-sensitivity bound, selection chained to a public randomness beacon, and dual adjudication whose disagreement is added to the bound. Undeclared deployment surfaces are ineligible. Each report states its own conditions on its face.
Verification is free for everyone, always. What differs is how much comes back — and that depends on who is asking, which is why the checkpoint asks you to say. Requests are authenticated — never anonymous — and we log exactly what is checked, by whom, and when, under the retention schedule in the privacy policy.
Every ATBAE manifest carries two independent signatures from our signing station — classical Ed25519 and post-quantum ML-DSA-65 (FIPS 204) — and is independently time-stamped under RFC 3161. Both key fingerprints are published on the checkpoint page — on this website, separate from any signed package — so they never travel solely inside the document they authenticate. A signature whose fingerprint differs from that page is not ours, whatever the package claims. Manifests sealed before 2026-08-31 carry classical signatures only and are never silently retrofitted.
Verification is not a revenue line and never will be. A seal only means something if the person relying on it can check it. Identity is declared and verified so that disclosure matches standing — it is authenticated, not harvested, and we do not tell the commissioning client who asked unless they elected Verification Transparency. Where a regulator lawfully directs us to withhold notice, we comply unconditionally.
What we charge for is a Confirmation of Seal — a dated, logged, counterparty-addressed record with an explicit scope and validity window, of the kind a vendor-risk team can put in a file.
Scoping is a bilateral technical consultation, not an automated form. Tell us the endpoint, the decision your model informs, and the framework you answer to. Quiet, infrequent replies — no marketing, ever.