Why we audited ourselves
Before anyone uses these rankings to inform a vote, they have to be defensible to a skeptical reader of any party. So we asked the hardest question first: are we grading with a thumb on the scale? The owner requested a structural check on his own bias, and this page is the unedited result.
To be plain: these rankings are not yet promoted as voting guidance. That unlock stays closed until the open items further down this page are resolved.
Audit record: spec 0417aa08 · decision 89938325
The adversarial review
We handed seven core scoring rules to two independent reviewers and told each to break them — one attacking from a progressive lens, the other from a conservative lens, each aiming at the opposite conclusion.
A separate synthesis judge then re-verified every factual claim against the live record before ruling. One prosecution exhibit was thrown out when its facts did not check out.
The verdict: all seven rules survived at their core — often because the two attacks pulled in opposite directions and cancelled each other out. The dominant confirmed problem was not the rules. It was coverage.
Audit record: progressive-lens review 8702b935 · conservative-lens review 29f5c8a5 · synthesis judgment fcdbb9f1
What we found — and fixed
Every correction below was made on 2026-07-19 and 2026-07-20.
Coverage gaps
Several prominent Biden-era actions that courts later struck down were missing as scoreable actions — they existed only as the court rulings about them. The Mayorkas impeachment was missing entirely. And Republican messaging bills from the 119th Congress sat in the record with no Democratic counterparts from the 117th.
The eviction-moratorium extension, the student-debt-relief instrument, and the Mayorkas impeachment are now ingested and scored by the same production pipeline as every other action. The OSHA vaccine emergency standard was already present but filed at a default weight; we corrected it to major impact, matching its official government significance determination. The sweep of remaining 117th-Congress bills is still open.
Until coverage parity completes, every cross-administration serious-violation comparison on the site carries this visible caveat — the same one shown on The Record today:
Coverage is still growing: some major actions — including prominent actions from recent administrations that courts later struck down — are still being added to the record. Serious-violation counts reflect actions reviewed so far and can rise for any administration as coverage completes.
A live scoring inconsistency
Two structurally identical resolutions carried different scores. Both now carry identical profiles, and an automated test now runs on every code change and fails if these two resolutions' profiles diverge.
A data error
One enacted law was also recorded as a separate, not-enacted duplicate. We deduplicated the pair and corrected it.
Audit record: synthesis judgment fcdbb9f1 · coverage-parity caveat decision 068cbb7d · structural-twin regression test 1dbb6ab0 (PR #201)
The symmetry harness
We built an automated detector that finds structural twins — pairs of actions alike enough that they should score the same. It found 2,076 such pairs, and on 50.6% of them at least one of the underlying evaluations differed.
Read alone, that number sounds alarming — but it is not "half our scores are inconsistent." Re-scoring (see below) showed that most twin disagreement is explained by ordinary evaluator boundary noise that is party-blind, not systematic bias. Every remaining pair is being triaged individually, and we will not consider this closed until every single one is resolved.
Audit record: symmetry-harness deploy report 8eeb5a03 · decision 47ea913d
The calibration test
We pre-registered this test to keep ourselves honest: the sample was published and hash-locked before any re-scoring began, so we could not quietly reshape it around a result.
We then re-scored 747 actions through the byte-identical production evaluator and measured whether it reproduced its own results.
In plain terms: we re-ran the same evaluator on the same 747 actions and checked how often it landed on the exact same answer twice, and how much the overall scoring pattern shifted.
- Per-evaluation reproduction: 96.0% (95% confidence interval 95.4–96.5), against a pre-declared floor of 90%. Pass.
- Corpus drift (how much the overall pattern of scores shifted between the original run and the re-run): 0.037 — under the 0.05 ceiling we set before testing. Pass.
- By principle, reproduction ranged from 93.4% (limited & divided power) to 97.6% (liberty).
Of the twin pairs re-scored to completion, 57.5% of 113 came back fully identical. That is what independent noise predicts: a roughly 4% per-evaluation wobble compounding across the twelve evaluations in a pair works out to about 61% expected-identical. That is consistent with random noise rather than a partisan lean — the same math as flipping twelve slightly-loaded coins and asking how often all twelve match on a repeat toss.
In full honesty: spending hit its pre-set ceiling with 113 of 150 planned pairs complete. We report the sample we actually finished, as-is — not a padded one.
Audit record: pre-registered sample manifest sha d2b275f6… · results artifact 759cd841
Evidence the metric is not tuned to a conclusion
Before a data-attribution repair on 2026-07-19, the very same "farthest from founding values" metric ranked Biden first; once the attribution error was corrected, it ranked Trump first. No one chose either outcome — fixing a data error moved the result. The machinery does not know which answer anyone wants.
Audit record: methodology memo 8144f083 (R12) · attribution-repair deploy report c9b709d2
What the audit could not clear — published anyway
Two rule-level defects survived the audit unresolved. We are publishing them rather than sitting on them; both await owner-level amendment through the public governance process.
- Downstream constitutional checks currently amplify in one direction only: a court can raise an action's weight, but a final vindication or a Senate dismissal sends no signal back. A two-way doctrine is proposed and pending.
- The draft pardons rubric still needs both parties' distinctive clemency-abuse patterns made visible before it can ship.
Also still open: completing the 117th-Congress bill sweep, adding historical electoral-certification records still missing from the archive, a pillar-by-pillar review of all 58 most severe findings, resolving the remaining twin-pair disagreements one by one, and aligning our internal quality-check thresholds across the board.
Audit record: two-way-doctrine proposal d25b2988 · pardons-rubric finding b8d36642 · severe-findings census 29f5c8a5 + fcdbb9f1
A standing invitation
Every number on this page traces to a dated audit record. When we fix a finding, we publish the correction. When we cannot fix something yet, we say so — right here, in public.
Principle First. Party Never.