Twelve major AI labs scored and ranked against the Synthetic Outlaw framework across four governance dimensions, sourced to published case law, regulatory findings, and investigative journalism. The documented record of institutional conduct.
"The most consequential governance failure of the AI era will not look like a rule being broken. It will look like a system hitting its targets right up to the moment society cannot live with the result."
Every major AI lab publishes safety documentation. Several have published constitutional frameworks, responsible use guides, and frontier safety commitments. None of that documentation is independently scored against the institution's actual conduct.
The Synthetic Outlaw framework measures the gap between what a governance constraint requires and whether that constraint survives when commercial or competitive pressure pushes against it. Applied to AI labs, the question is precise: does the institution's demonstrated conduct match the product claims it makes? How far apart are those two things? What does that gap cost the public?
This scorecard compares twelve labs across four dimensions. Each score is tied to a dated public record and recalculated when a material development changes the evidence.
AI circumvents rule intent while remaining formally compliant. The rule survives. Its protective purpose doesn’t. Oversight sees compliance because it is looking at the surface the system learned to satisfy.
Accountability dissolves across actors and layers. No single party can be held liable for the harm. Once capability crosses its boundary, controls bind custodians. Copies, derivatives, and downstream integrations fall outside every accountability mechanism that applies to the originating lab.
Regulatory oversight is structurally compromised by the entity it oversees. Not corruption. Structural dependency. The entity shapes the standards it is then measured against.
Existing legal frameworks were not built for this technology and do not map onto the behavior. The harm is real. The legal category for it doesn’t exist yet. In some cases, that category was actively lobbied out of existence.
| RANK | LAB | BYPASS | DIFFUSION | CAPTURE | GOV GAP | OVERALL | VERDICT | KEY FINDING |
|---|
Legislation gets lobbied. Corporate promises get rewritten. Dashboards report what the system learned to report. None of these are binding under optimization pressure. That is the exact condition the Synthetic Outlaw framework was built to detect.
The scorecard does not rely on lab self-assessment. It compares public claims with court records, regulatory findings, official investigations, publicly documented technical architecture, and attributed reporting.
The record is cumulative. A resignation, court ruling, regulatory finding, lobbying disclosure, or independently verifiable remedy is appended to the chronology. New evidence can move a score; it does not erase the historical record.
The Synthetic Outlaw Index scores 49 AI deployment domains across 12 sectors. This lab scorecard extends that methodology to the institutions building the systems. Together they form the first independent, scored, public record of AI governance failure. Deployment decisions and institutional conduct, measured in the same place, against the same standard.
SCOPE
Twelve AI labs, scored on documented institutional conduct through 20 July 2026.
INCLUDED
Dated public claims, court and regulatory records, publicly documented technical architecture, and attributed reporting. Each item is weighted under the protocol above.
RESERVED FOR THE BOOK
The complete legal theory, full scoring model, and institutional remedy architecture are developed in The Synthetic Outlaw (forthcoming 2026).
Each lab is scored 0–10 across four dimensions drawn from J. Gropper, The Synthetic Outlaw (forthcoming 2026). Governance rank follows overall score; equal scores share a rank. Scores reflect published case law, regulatory findings, investigative journalism, and peer-reviewed research as of July 2026. This scorecard measures institutional conduct: governance architecture, deployment decisions, and the gap between public safety claims and documented behaviour. Model benchmarks, capability evaluations, and market share are outside the score. Every score is traceable to the source cited beneath it.