LIVING LEGAL CORPUS · NO. 002
The iterations behind the living legal corpus
Storing the law is the easy half. The hard half is knowing whether what you stored is still true — and every mechanism here exists because something that was supposed to catch that reported OK instead.
> the one-time provenance backfill bypasses save() via QuerySet.update() # legal/models.py, LegalSourceSnapshot.save()
Software that cites the law has a defect no build server can see: the law changes, and the software does not know.
The code compiles. The tests pass. The article still renders the sentence it rendered last year, and the fee cap in the enforcement guard is still the integer someone typed into it. Nothing is broken in any sense a build server understands. The system is describing a world that stopped existing when a legislature adjourned.
This is the second post in the series. The first was about a ledger that records its refusals. This one is about the layer underneath it — the corpus that tells the ledger what the rules are, and how far we got before discovering that most of our confidence in it came from checks that were not looking at anything.
What follows is the scaffolding: the iterations, abstracted from production code, and then the failures. Every one of the failures reported success.
PART I
Making a statute into something that can change
A statute enters software in one of two disguises. Here is one number — the cap Washington puts on what an association may charge to prepare a resale certificate — in both of them.
# iteration 0 — the law as a constant
RC_PREP_FEE_CAP_USD = 275 # RCW 64.90.640? someone checked once
def article_body():
return "Washington caps the preparation fee at $275."
Two claims about the world. Neither carries the text it came from, the subsection it lives in, the date anyone verified it, or any way to find out that it moved. The number and the sentence will outlive the statute that justified them, and the day they stop being true is a day exactly like every other day. There is no alarm for this because there is nothing an alarm could watch.
# iteration 1 — five authority types, one target, enforced by the database
class ContentCitation(Model):
statute = FK(Statute, null=True)
case_law = FK(CaseLaw, null=True)
session_law = FK(SessionLaw, null=True)
regulation = FK(Regulation, null=True)
federal_guideline = FK(FederalUnderwritingGuideline, null=True)
class Meta:
constraints = [CheckConstraint( # exactly one, in the DB
name="content_citation_exactly_one_target", ...
)]
Five authority types, because four was not enough: Fannie Mae, Freddie Mac, FHA, and VA publish underwriting guidelines that decide whether a condominium is financeable, and to a board those function as law. Every citation record points at exactly one of them, and the exactly-one rule is a longhand database check constraint rather than a validator — the three that matter are named legal_source_snapshot_exactly_one_target, content_citation_exactly_one_target, and law_change_event_exactly_one_target. A convention survives until someone writes a migration on a Friday. A constraint does not care what day it is.
# iteration 2 — hash the text, never edit it
class LegalSourceSnapshot(Model):
full_text = Text()
content_hash = Char(64, db_index=True)
@staticmethod
def compute_content_hash(text):
normalized = ' '.join(text.lower().split())
return sha256(normalized.encode('utf-8')).hexdigest()
def save(self, *a, **kw):
if not self._state.adding:
raise ValueError("records are immutable — create a new snapshot")
A statute is not a row now. It is a row plus an ordered chain of snapshots, each one a dated capture of the text with a hash of that text. Normalization is deliberately blunt — lowercase, collapse every run of whitespace — so that a state reflowing its HTML does not read as an amendment. Refetch, hash, compare to the newest snapshot. Identical hash, write nothing. Different hash, the law moved.
Immutability is the part that makes the chain worth anything. The guard is a check in save(). Directly above it is a comment naming the way around it:
# provenance backfill bypasses save() via QuerySet.update(). legal/models.py — LegalSourceSnapshot.save()
Application-level immutability is a promise the ORM keeps only for code that goes through the ORM the polite way. Part III is the check that exists because of it.
# iteration 3 — detecting drift is worthless if nothing downstream hears it
def flag_affected_citations(self):
qs = ContentCitation.objects.filter(needs_reverification=False)
if self.statute_id: qs = qs.filter(statute_id=self.statute_id)
elif self.case_law_id: qs = qs.filter(case_law_id=self.case_law_id)
# ... one branch per authority type, else: return 0
return qs.update( # one statement, not a loop
needs_reverification=True,
drift_detected_at=timezone.now(),
)
A hash difference creates a LawChangeEvent whose core is frozen at creation — both snapshot references, the change type, the diff — while its review and publication fields stay mutable so a person can act on it. The event then walks its own foreign keys outward and marks everything in the database that depended on the old text — content citations, eligibility rules, and exclusion triggers — one bulk update per class. An article that quotes a statute that moved last night is flagged this morning, and the flag is a database column, not a task in someone's inbox.
There is one exception, and it is a gap rather than a design. Resale-certificate compliance profiles are Python dictionaries, not rows, so the event returns their keys and persists nothing. For that one dependent class the flag really is a list someone has to read.
Nothing about that flow publishes anything. A change event defaults to PENDING_REVIEW, and a comment in the automated processor's own apply() path forbids it from ever touching publication status. A machine can notice that the law moved. Only a person decides what it means.
# iteration 4 — the same number, carrying the sentence that establishes it
LegalThreshold(
code = "WA_RC_PREP_FEE_CAP",
statute = RCW_64_90_640, # pinpoint: "(2)"
statute_quote = "...", # the establishing sentence, verbatim
value = Decimal("275.00"),
value_unit = "USD",
comparison = "MAX",
)
# guards never read the number directly — only through the resolver
cache_key = f"legal_threshold:{code}:active" # TTL 300s, signal-invalidated
This closes the loop opened in iteration 0. The same $275 is no longer an integer in a guard; it is a row carrying the statute, the subsection, the verbatim sentence that establishes it, and whether it is a floor or a ceiling. Enforcement guards resolve it through a caching service. When a legislature moves the number, the fix is a data update and a cache invalidation, not a deploy — and the article that quotes it is flagged by the same mechanism from iteration 3.
That claim has an asterisk. The resolver fails open: no active threshold for a code returns permission, and one guard still carries module-level fallback constants for the case where the corpus has nothing seeded. The design intent is that statutory numbers live in the corpus. The current state is that most of them do.
PART II
The checks that lied
A corpus of this shape runs on checks: nightly scans that assert the snapshots are real, the hashes reconcile, the citations resolve, and an adapter exists for every state. We built those checks early and believed them for months. Then we started auditing the checks themselves.
The scan that crashed nightly, unread. The health check died at its first check on every 04:00 run for a month with a NameError on an unimported symbol. The scheduled task logged the traceback and re-raised into a Celery failure nobody was watching, so no scan record was written and nothing turned red. It surfaced only when a separate gate asked a different question: not did the scan pass, but when was a scan record last written. FAILURE MODE — A FAILURE NOBODY READS IS AN ANSWER FABRICATED
The check that matched nothing. Check 1 looks for placeholder text sitting where a real statute should be. It was written with the hand-typed literal '[Placeholder for' while the seed commands write '[Placeholder snapshot for'. The query matched zero rows and returned a clean result every night it ran. The comment recording this is still in the command, deliberately.
The alert that had never once fired. Every drift alert imported its alert manager from a module that did not define one. The ImportError was caught and logged below the level anyone reads — DEBUG in the drift detector, INFO in the health-check task — each handler treating a missing symbol as a missing app. The HIGH-priority legal drift alert was a no-op from the day it was written. FAILURE MODE — SEVERITY IN THE NAME, NOT IN THE PATH
Ten hashes that no longer matched their text. A full-corpus recompute found ten of 1,367 snapshots whose stored hash did not reproduce from the stored text — all from one cleanup window in May that wrote the two fields through different code paths. A stale hash is worse than a missing one: it will either fake drift or mask it on the authority's next fetch. The text was correct. The fingerprint of the text was not. FAILURE MODE — TWO WRITERS, ONE INVARIANT
Nav chrome stored as law. Texas intermittently served its app shell instead of the static chapter export, and the nightly fetch laid 1.3 KB of navigation furniture over a real 13.7 KB capture of Property Code chapter 207. Kentucky served its own this version is superseded notice, which was captured and stored as the statute. Applying the new gates retroactively surfaced forty more captures already sitting in the corpus as real text: tombstone redirects, raw PDF bytes, JavaScript apology pages, and tables of contents. FAILURE MODE — HTTP 200 IS NOT A CLAIM ABOUT CONTENT
Two checks that worked, at a severity nobody acts on. Stale snapshots and overdue verification were both real checks, correctly implemented, running nightly, finding real problems — and reporting them at INFO. Behind those lines, 714 stale snapshots and 438 statutes unverified for ninety days or more accumulated while the scan reported OK. Both are ERROR now, and both carry a comment in the source saying why they were promoted. FAILURE MODE — A FINDING NOBODY IS PAGED FOR IS A FINDING NOBODY HAS
The articles were right; the mirrors were behind. A queue of thirty-eight apparent citation errors mostly resolved the other direction. Nevada's public mirror was missing fee provisions operative July 1, 2026; Texas's was missing a subsection added by a 2025 session law. Three articles did need correcting. The rest were right, and the corpus had flagged them against sources that were themselves behind. FAILURE MODE — VERIFYING AGAINST A COPY
Not one was a crash a human saw. Not one turned a dashboard red. Every one was a green check, reporting success, looking at nothing. Each was caught because a different mechanism asked a question the broken one could not answer for itself.
A check reports on its subject and on itself simultaneously, and it cannot distinguish the two. When check 1 said no placeholders found, it was making one true statement and one false one, in the same word, and there is nothing in the output that separates them.
PART III
Checking the checker
The response was not more checks. It was a specific kind of check: one whose subject is the corpus's own machinery, written so that the mechanism it verifies is not the mechanism it runs on.
Does the stored hash still reproduce from the stored text?
Check 13 recomputes the SHA-256 of every snapshot in the corpus, nightly, and errors on any divergence. This is redundant by construction: the hash is written at creation and the record is immutable, so it should be impossible for them to disagree. It exists because of the comment in Part II — save() is bypassable, ten rows proved it, and the check makes that class of corruption impossible to reintroduce quietly. Its docstring says so, naming the incident and the date. The severity is fixed at ERROR with a one-line justification: hash-text divergence is corruption, never a judgment call.
Was the article wrong the day it was published?
Drift detection compares snapshot N to snapshot N-1. That means it can only ever notice the law moving. It is structurally incapable of noticing an article that was wrong on the day it went up, because there is no diff — the statute never changed, the sentence about it was never true. No amount of drift detection closes that gap.
So a separate check proves each pinned citation against the snapshot text it points at: that the cited subsection appears, that the quoted language is in the source, and that the numbers in the claim match the numbers in the law. It found a class of error drift could not have: a correct claim of '50 days' failing against a statute that reads 'nor more than fifty days', because the spelled-number map had been typed by hand and omitted fifty. The map is now derived for the whole range rather than enumerated.
Is this text law at all?
Because snapshots are immutable, a bad capture cannot be edited away. It can be quarantined — the row survives for audit, excluded from every consumer — and superseded by a fresh fetch. What it cannot be is un-happened, which is how all forty of the retro-validated captures were handled. So the cheaper move is to refuse the write, at the moment of creation, for automated fetches only. A human quoting a source has already exercised the judgment the gate is there to supply.
# as "real" snapshots. Each marker cites its incident; a marker must be
# a phrase that cannot occur inside genuine statutory text. legal/models.py — JUNK_CAPTURE_MARKERS
Three markers and two incidents, each annotated with the statute it damaged and the date. The list is short and will stay short, because the standard for entry is strict: a phrase that cannot occur inside real law.
Fifteen checks run nightly. The scan record they produce is itself immutable, hashed over a canonical serialization of its findings, so a night's verdict cannot be revised after the fact either.
PART III-A
The decision remembers which law it stood on
Enforcement decisions carry the law their rules stood on. Every immutable enforcement decision — both phases of the telemetry, the judgment and the confirmation — freezes the snapshot hashes of the statutes authorizing the rules in force when it was made, so a decision can be replayed against the law as it read that day — and because the authorities list is inside the decision's own hash, the citation cannot be rewritten afterward either. Collection is best-effort by contract: it never raises and never blocks, so an unreachable corpus costs provenance rather than a posting. The cost is bounded by keying the result to the policy-snapshot hash the dispatcher has already computed, which means the corpus is read once per policy rather than once per transaction.
The honest scope, because the record is permanent: this is the legal backing of the active eligibility rules and exclusion triggers in force for the evaluation. It is not the narrower and more flattering claim — the statutes this particular transaction touched.
PART IV
What it still does not do
Beyond the two admitted in Part I, these are the current gaps, in the order they bother me.
The nightly scan is red by design. Most of the checks emit ERROR when they find something, and any one of them alone puts the scan record at RED. Stale snapshots and overdue verification stay unsatisfied while the refetch and verification backlog exists; both were INFO until August 5, 2026, which is how the backlog got large enough to matter. Which checks are red on any given night is a fact about the corpus, not about this sentence — the grid below reads it from the last scan record rather than restating it here. The corpus does not muster green today.
Backing is the honest metric, not coverage. Citation patterns span all fifty states and a federal layer. Statutes that are backed — an authority with real captured source text behind it, not a placeholder — are a smaller number, and several hundred entries are still awaiting a first genuine capture. The counts in the provenance block below are coverage, not backing; the backed number is computed elsewhere and is not yet rendered live on this page. Pattern coverage is not evidence of anything except that we know how to parse a citation.
Silence still is not proof. Ingestion writes a snapshot only when the hash changes, which means a quiet statute and a dead scraper produce identical data: nothing. Over any window, absence of change is indistinguishable from absence of checking — the same failure mode as check 1, one layer down, and not yet closed. Per-run fetch logging is the fix and it is not built.
Drift does not cover federal underwriting guidelines. The fifth authority type participates in citations and constraints, but the drift sweep never scans it — it iterates the other four, and a direct call with a guideline reaches a TypeError instead of a diff. When Fannie moves a threshold, nothing fires.
Of 15 nightly checks, 11 pass and 4 are red: fact_quote_backing, stale_snapshots, threshold_value_drift, verification_staleness.
15 CHECKS RUN EVERY NIGHT.
4 ARE RED TONIGHT.