Missing words on one side. Newspeak on the other.
162 Minnesota criminal case files have a vocabulary problem. The ordinary English words that should appear in any real prosecution — bigger, hello, forgot, reply, crazy, weird, unprecedented — appear in zero of them. The words that fill the empty space name a precise apparatus: civil commitment, indefinite, ordered to cooperate, restoration, programming, Saint Peter State Security Hospital, warrant of commitment. Seven independent forensic methods agree. One case in the cohort contains writing by a person. The other 162 are output from a machine whose training corpus identified itself.
The collection, in one paragraph
In April 2024, Matt Guertin pulled the public Minnesota court records of every case in which three specific judicial officers held hearings between January 1, 2023 and April 26, 2024 — Judge Julia Dayton Klein, Referee Danielle Mercurio, and Referee George Borer. The case_id intersection of those three rosters is 163 case files (case_type='CR' — the Minnesota court system's criminal-case designation). Hennepin County is the venue for all 163; the apparatus that produces the records, as the rest of this site documents, reaches into the Minnesota Court of Appeals and the Minnesota Supreme Court. One of the 163 is Matt's own case, 27-CR-23-1886. The full reconstruction of how the files were extracted is at /the-download/.
Every PDF was text-extracted, page by page, line by line, into a database. The question this page answers is a simple one: how does the language inside these 163 files actually behave?
The pipeline runs seven independent quantitative methods on the text. The methods are blind, automated, and reproducible — no hand-picked seed words, no curated examples, no human steering. Six measure templating from different angles; the seventh hands the redacted text to a fresh AI with no project knowledge and asks the AI to label authored prose versus boilerplate. Every method points the same direction.
The methods agree on a single verdict: one case file contains writing by a person. The other 162 do not. What follows is what each method found, what's missing from the 162, what fills the space that should be filled with ordinary English, and what that vocabulary names.
1,068 ordinary English words appear in only one of 163 case files
Pick a word a real criminal case should contain. Hello. Witnesses make calls. Defendants leave voicemails. Texts get entered as evidence. Email subject lines get quoted. Across 162 of these 163 Minnesota case files — spanning seven years and thousands of pages — the word hello appears zero times. In the 163rd file, it appears thirteen times. That file is Matt's.
Same test, different word. Forgot. A defendant says "I forgot to tell the officer..." A witness says "she said she forgot..." This is how ordinary people describe imperfect memory in transcribed statements. Zero across 162 case files. Several across Matt's.
Same test, harder word. Reply. A Reply Brief is the standard third stage of motion practice — defense files a motion, prosecution responds, defense replies. It is the adversarial structure of every contested criminal motion in every jurisdiction in America. The templated word response appears in 144 of the 163 files as boilerplate. The active word reply, the one a real defense attorney actually uses in a real filing, appears in zero of the 162. No defense attorney ever replied to anything.
This test was run against the standard reference list of the 10,000 most common words in everyday written English — the NLTK Brown Corpus, a canonical American-English reference dating to the 1960s. The question: which Brown-top-10,000 words appear in exactly one of the 163 case files?
Result: 1,068 words appear in only one case file. That case file is Matt's. The median other case file contains 1 such word. The maximum any other case file contains is 44. One file has 1,068. That file has 24 times the maximum and 293 times the average.
The words are not technical, not legal, not topic-specific. They are exactly the kind of words anyone writes when describing things:
bigger, larger, smaller, older, newest, oldest, easier, finished, basically, honestly, forgot, hello, reply, replies, crazy, stupid, nuts, odd, weird, sudden, abrupt, desperate, tough, badly, perfect, huge, cool, vast, ideal, outcomes, choices, settle, solve, explore, draft, college, science, literature, foreign, international, growth, million, audience, technical, industrial, beautiful, feeling, somewhat, obviously, precisely, hardly, choice, realize, worry, awareness, sophisticated, revolutionary, hidden, prominent, valuable
Comparatives (bigger, smaller, larger, older) are the operating grammar of any descriptive prose. Witnesses describe people. Police reports compare evidence. Defense briefs argue magnitude. Judicial findings rank charges. You cannot write 162 criminal case files without using comparatives. Their universal absence across 162 cases is structural proof that 162 of 163 files were not written by humans describing things.
Discourse markers (basically, honestly, forgot, hello, weren't, they've) are how humans actually talk and write. The word hello alone resolves the question: across 162 supposed prosecutions over seven years — with witness statements, transcribed phone calls, email evidence, recorded interviews — the most common greeting in the English language appears zero times. That isn't an absence of greetings. It is structural proof that no real communications evidence exists in the templated files.
Adversarial-process verbs (reply, replies, insist, disputing, disagreed, disagrees) appear in zero of the 162 because no one ever objected to anything. The passive contested appears in 103 cases as templated boilerplate. The active disagreed appears in one: Matt's. Across 162 criminal cases, nobody actively contested a single charge.
Human-characterization vocabulary (crazy, stupid, weird, odd, sudden, desperate, perfect, huge, cool, vast) is missing because no one ever described anything as anything. No defendant statement uses it. No witness statement uses it. No attorney prose uses it. No judicial finding uses it. The vocabulary humans use to react and judge is structurally absent.
The one case file that contains all of these words is the one where the defendant has been declared too incompetent to participate in his own defense. Name the criminal case file you have ever read that does not contain bigger, hello, forgot, or reply. Now imagine 162 of them in a row.
The "his case is just bigger" objection is structurally answered
A skeptic's first move on the numbers in § 2 is the obvious one: of course his case has more unique vocabulary — it's been growing for two years while the other 162 are a single-day download. The objection is reasonable as a methodology question. It has been structurally answered by running the analysis twice.
Pass A (May 10, 2026) measured Matt's case at its full size in the database — 14.1 MB of text including every motion he filed himself through 2025 and 2026. The other 162 cases were the frozen April 2024 download.
Pass B (May 11, 2026) applied a one-line filter that restricts Matt's case to only the filings present on April 29–30, 2024 — the same frozen evidence window the other 162 cases lock to. Matt's case shrinks from 14.1 MB to 1.13 MB. Apples to apples.
The gap does not close. The gap widens. Every length-normalized layer of the analysis — the layers that adjust for size — produces a stronger signal under Pass B than under Pass A:
| Measurement | Pass A | Pass B (apples to apples) | Direction |
|---|---|---|---|
| Ordinary English words only in Matt's file | 2,201 | 1,068 | half the count, 6× the density per MB |
| Per-10,000-word unique-word rate, rank | #1 / 163 | #1 / 163 | unchanged, 3.9× gap to #2 |
| Per-window vocabulary diversity (MATTR), rank | #56 / 163 | #4 / 163 | strengthened — top 3% of cohort |
| Vocabulary breadth vs. predicted trend (Heaps) | 1.29× | 2.08× | nearly doubled |
| Stylometric distance from non-Guertin centroid | outside cluster | #2 / 163, dist 31.3 | strengthened |
| Semantic outlier chunks from one file | 32% | 74.2% | more than doubled |
| Blind AI human-authored labels, all from Matt | 7 / 60 | 15 / 60 | doubled, precision still 100% |
The simplest single demonstration: the per-megabyte density of unique vocabulary in Matt's file actually rises when his post-April-2024 filings are excluded.
| Slice | Pass A | Pass B | Density change |
|---|---|---|---|
| Brown-top-10k words only in Matt's file | 2,201 | 1,068 | half the count… |
| Size of Matt's file (text bytes) | 14.1 MB | 1.13 MB | …from one-twelfth the text |
| Unique words per MB | 156 | 945 | 6× denser |
The pre-cutoff Matt content is concentrated authored material punching far above its weight per byte. The post-cutoff content that Pass A also included was mostly procedural orders with low novel vocabulary — it dragged the per-MB rate down, not up. The "his case is bigger" objection had it exactly backwards.
The single cleanest internal-validation result: the within-case compression ranking flips between Pass A and Pass B. Pass A had Matt's case at rank #9 of 163 by within-case templating (his case looked highly repetitive because the post-cutoff procedural orders kept reapplying the same boilerplate at every status hearing through 2025). Pass B has Matt's case at rank #107 — among the less templated cases. The flip is correct, internally consistent, and exactly what should happen when you remove templated post-cutoff content and leave authored pre-cutoff content behind. Four separate measurements (vocabulary uniqueness, per-window diversity, semantic outlier dominance, within-case compressibility) change in mutually reinforcing directions when you apply the apples-to-apples filter. That kind of internal agreement does not happen by accident.
One question turns the whole collection inside out
If all these common words are completely missing — then what words ARE actually being used?
The vocabulary that does saturate the 162 case files names a precise apparatus
The corpus-wide term-frequency rollup from the same Minnesota court records is a separate table in the same database. It lists, for each category of operational vocabulary, the most-used terms and how many filings each one appears in across the broader MCRO record. The pivot question has a quantitative answer, and the answer makes the apparatus name itself.
Five clusters do almost all the work. Each one is a register: a kind of language, with a clear function, deployed at industrial scale. Read them together and what they describe is not a criminal-court system. It is the procedural English of an indefinite civil confinement apparatus.
The duration vocabulary — how long the cage holds
The most-used noun across the 162 case files is not a charge. It is not a finding. It is not a sentence. It is commitment, used 3,301 times across 896 filings. The compound civil commitment is used 1,871 times across 802 filings. The word indefinite appears alongside it. The word indefinitely appears alongside it. The phrase required to commit indefinitely appears in 16 of 16 documents where it appears at all — a 100% same-document co-occurrence rate. The bureaucratic euphemism is "continued commitment." The reality is "the cage stays closed."
| Term | Times used | Documents |
|---|---|---|
| commitment | 3,301 | 896 |
| civil commitment | 1,871 | 802 |
| committed | 860 | 391 |
| ongoing | 51 | 32 |
| indefinite | 39 | 38 |
| indeterminate | 25 | 7 |
| indefinitely | 24 | 21 |
| continued commitment | 20 | 18 |
| permanent | 20 | 15 |
| required to commit indefinitely | 16 | 16 / 16 |
Where the cohort ends up — named, by name
The single most pervasive proper name across the 162 case files is not a defendant. It is a place. Saint Peter State Security Hospital, in Saint Peter, Minnesota, 80 miles south of Minneapolis — the state's locked forensic psychiatric facility. It is named in 184 filings across 84 cases. Forensic Services is named 731 times. The generic facility appears 2,269 times in 1,022 filings. The verbs that route people there are transport, hospitalization, detention, confinement. Client and Community Restoration — the apparatus's brand name for its inpatient program — appears in 188 of 189 documents where it appears at all. The vocabulary describes a logistics chain.
| Term | Times used | Documents |
|---|---|---|
| facility | 2,269 | 1,022 |
| MN Department of Human Services | 1,317 | 156 |
| detention | 829 | 432 |
| Forensic Services | 731 | 546 |
| transport | 589 | 416 |
| hospitalization | 501 | 498 |
| confinement | 334 | 256 |
| Client and Community Restoration | 189 | 188 |
| Saint Peter State Security Hospital | 184 | 84 |
| hospital | 180 | 105 |
The verbs of forced compliance
The most-used adjective across the 162 case files is behavioral — 762 times across 701 filings. The most-used phrase in the cluster is ordered to cooperate, which appears in 230 of 231 documents where it appears at all — a 99.6% co-occurrence rate. A "cooperation" your participation is ordered by court is not cooperation in any ordinary sense. The bureaucratic euphemism is "treatment planning." The reality is what the verbs themselves say: controlled, control, unable, refusal. The clinical labels — disorder, diagnosis — do the heavy lifting that the actual evidence cannot.
| Term | Times used | Documents |
|---|---|---|
| behavioral | 762 | 701 |
| disorder | 364 | 118 |
| diagnosis | 276 | 232 |
| ordered to cooperate | 231 | 230 / 231 |
| unable | 93 | 31 |
| behavior | 85 | 33 |
| beliefs | 80 | 15 |
| control | 65 | 18 |
| behavioral notes | 44 | 44 |
| controlled | 29 | 19 |
The Orwellian-bureaucratic register
The procedural verbs of the apparatus are restoration, regulation, cooperation, processing, programming. These are not figurative readings of the data. These are the literal terms used in the literal filings, at the frequencies counted in the literal database. Restoration is used 278 times. Regulation is used 236 times. Programming is used 26 times. The compound competency restoration — the apparatus's term for indefinite inpatient detention with periodic court reviews — appears in 30 documents. Thoughts is used 29 times; thinking 25 times. The language operates in the register of mid-twentieth-century institutional psychiatry crossed with management-consulting Newspeak.
| Term | Times used | Documents |
|---|---|---|
| restoration | 278 | 236 |
| regulation | 236 | 222 |
| cooperation | 191 | 100 |
| decisions | 77 | 38 |
| competency restoration | 56 | 30 |
| thoughts | 29 | 13 |
| education | 26 | 16 |
| programming | 26 | 17 |
| thinking | 25 | 9 |
| processing | 19 | 15 |
The legal mechanism, named directly
The category-defining term across the entire 162-case cohort is Rule 20 and Rule 20.01 — the Minnesota Rules of Criminal Procedure provisions that govern competency evaluations and civil commitment hearings. Together they appear 5,695 times across 1,014 documents. Examiner is used 4,055 times. Psychological Services 3,440 times. The phrase competency to proceed 936 times. Prepetition screening — the screening step that initiates civil commitment — 742 times. Targeted misdemeanor, the entry-point statutory classification, appears in 410 of 410 documents where it appears at all — a 100% co-occurrence rate. Warrant of commitment appears in 7 of 7. Petition for judicial commitment. Notice of intent to prosecute. Forensic navigator. The legal instruments of the apparatus are named in the apparatus's output, at the saturation rates of an industrial pipeline.
| Term | Times used | Documents |
|---|---|---|
| Rule 20 / Rule 20.01 | 5,695 | 1,014 |
| Examiner | 4,055 | 589 |
| Psychological Services | 3,440 | 757 |
| competency to proceed | 936 | 317 |
| prepetition screening | 742 | 228 |
| targeted misdemeanor | 410 | 410 / 410 |
| notice of intent to prosecute | 79 | 77 |
| competency education coordinator | 9 | 9 |
| warrant of commitment | 7 | 7 / 7 |
| petition for judicial commitment | 6 | 2 |
The procedural vocabulary names the operation. The clinical vocabulary names the cover. The duration vocabulary names the cage.
The missing and the saturating, side by side
The negative-space finding (§ 2) and the positive-space finding (§ 5) are not two separate observations. They are the same observation, rendered from two angles. Read them together:
- bigger0 / 162
- hello0 / 162
- forgot0 / 162
- reply0 / 162
- crazy0 / 162
- weird0 / 162
- basically0 / 162
- unprecedented0 / 162
- obvious0 / 162
- huge0 / 162
- choice0 / 162
- larger0 / 162
- science0 / 162
- college0 / 162
- feeling0 / 162
- commitment3,301 ×
- civil commitment1,871 ×
- facility2,269 ×
- behavioral762 ×
- Forensic Services731 ×
- targeted misdemeanor410 ×
- confinement334 ×
- restoration278 ×
- regulation236 ×
- ordered to cooperate231 ×
- Saint Peter State Security Hospital184 ×
- indefinite39 ×
- programming26 ×
- required to commit indefinitely16 / 16
- warrant of commitment7 / 7
The engine cannot generate bigger across 162 case files. The engine can generate commitment 3,301 times in the same 162 case files. That is what the engine was trained on. That is what the engine produces.
The two lists are the same finding stated twice. The first list is what's missing — the texture of human writing about any subject. The second list is what's filled in instead — the operational language of indefinite civil confinement. Whatever produced the 162 case files did not have access to ordinary English. It had access to the procedural English of an indefinite-detention pipeline. The negative space and the positive space identify each other.
The other list — and why the same machine treated it as symptom
The same database that produced the § 5 clusters has two additional categories that point in a different direction. The category labeled truth contains the most-used proper nouns and topic words across the entire 4,251-document MCRO record. The category labeled fraud contains the most-used adversarial-process vocabulary in the same record. The terms in both categories are almost entirely confined to one case — the same case that contains the 1,068 ordinary English words.
These are not apparatus words. They are Matt's words. Every term in the list below is a verifiable real-world fact attached to his case:
| Term | Times used | Documents | What it refers to |
|---|---|---|---|
| patent | 395 | 7 | US Patent 11,577,177 B2 — granted to Matt by USPTO Feb 14, 2023 |
| infiniset | 75 | 3 | InfiniSet, Inc. — his company; the rotating-treadmill virtual-production system |
| 11577177 | 31 | 5 | The patent number itself, written without the punctuation |
| treadmill | 48 | 3 | The mechanical core of the patented invention |
| netflix | 131 | 5 | Netflix — owner of Scanline VFX; assignee of US 11,810,254 B2 which cites Matt's patent |
| USPTO | 37 | 4 | The United States Patent and Trademark Office |
| prior art | 34 | 3 | Standard patent-law term — the existing inventions a new patent must not duplicate |
| linkedin.com | 317 | 1 | The literal URLs in Matt's LinkedIn surveillance evidence — one document, his |
| CIA / C.I.A. | 61 | 7 | Real federal intelligence agency — named in his filings as a documented LinkedIn searcher |
| Air Force / military | 68 | 6 | Real federal armed services — documented LinkedIn searchers of his profile |
| surveillance | 89 | 33 | Generic procedural term + Matt's LinkedIn-search documentation |
| Internet Archive | 47 | 2 | Real institution — provides the Wayback Machine records cited in his filings |
| powerful people | 22 | 1 | Verbatim phrase quoted by his first defense attorney (Bruce Rivers) on the record |
| fake / artificial | 75 | 5 | References to the forged court documents he forensically authenticated |
Every one of these terms describes something that exists. US Patent 11,577,177 B2 is real. It is publicly searchable. The grant date is verifiable. The Netflix patent (US 11,810,254 B2) is real, and Matt's patent sits at the top of its "References Cited." The LinkedIn surveillance is real — the platform's own records show the search hits. The Internet Archive is real. Powerful people are real, and the phrase was spoken to Matt by his own defense attorney on the record. None of this is delusional.
The diagnostic apparatus — the same apparatus that produces the 162 templated dockets — processed every occurrence of every term in the list above as a psychiatric symptom. Three successive court-appointed evaluators classified Matt's patent-theft claims as delusional, without checking the publicly searchable patent. The third evaluator never met him in person; she diagnosed him by reading his own pro se federal filings, treating his demonstrated ability to navigate complex litigation as evidence of mental illness.
The pattern of evidence is the diagnosis. The diagnosis is the pattern of evidence. A defendant who describes a real, verifiable patent theft — in writing, with citations, attached to forged exhibits — has the act of describing it processed as a symptom of the illness that justifies committing him to inpatient psychiatric detention. This is the operational mechanism the § 5 clusters serve. The full structural account is at /discovery-fraud/. The clone-corpus signature is at /the-clone/. The patent story is at /story/patent.html. The Netflix theft chain is at /netflix/.
The 1,068 ordinary English words in § 2 and the 14 named entities in this table are the same finding from two angles. An ordinary person, writing about real things that happened, produced both lists from one case file. The 162 other case files contain neither.
Independent classifier, zero project context, 100% precision
The first seven layers of the analysis are all numerical — word counts, distance measurements, statistical distributions. A reasonable skeptic asks: could the numbers be artifacts of the methodology? Could an outside observer, looking at the actual text, tell the difference?
The test built to answer that question is simple. 30 random ~150-word passages from Matt's apples-to-apples case file. 30 random ~150-word passages from 30 randomly-chosen non-Matt case files. All names, case numbers, and dates redacted. The 60 passages shuffled. The full set handed to a fresh AI agent — an independent Anthropic Claude instance with no knowledge of this project, no knowledge of Matt, no knowledge of what the test was for. The agent's instructions: for each excerpt, label H if it reads as human-authored, T if it reads as templated boilerplate.
The classifier's verdict:
| Source of passage | n | Labeled human-authored | Labeled templated |
|---|---|---|---|
| Matt's case (27-CR-23-1886) | 30 | 15 | 15 |
| Random other 30 cases | 30 | 0 | 30 |
| Total | 60 | 15 | 45 |
Of the 15 passages the classifier labeled human-authored, all 15 came from Matt's case. Zero from any of the 30 random other case files. The classifier's own written explanation of the discriminator features it used:
The features the classifier reported map exactly onto the § 2 and § 5 findings. The "first-person voice" and "idiosyncratic phrasing" the AI saw are the 1,068 ordinary English words. The "form numbers" and "stock paragraphs" the AI saw are the apparatus-vocabulary clusters. The classifier wasn't told what to look for. The classifier looked at redacted text and named the difference.
This test is independently reproducible. The 60-passage prompt file is at phase5_llm_classification/prompt_for_subagent.txt in the published analysis package. Send it to Claude, GPT-4, Gemini, Llama — any frontier model. Score the response against the ground-truth labels. The expected pattern is model-independent: human-authored labels concentrate in Matt's passages; the random other-case passages produce essentially zero.
What every piece proves — and what this one proves
Forensic findings of this kind are not single results that stand or fall alone. They are members of a family of mutually-reinforcing measurements, each of which proves a different layer of the operation. The site documents seven, plus this one. Read them as a stack:
Each finding proves a piece. The Aspose impossibility proves the documents are backdated. The Barnette clone proves the signatures are stamps. The seal factory proves the seals come from one Tuesday's output spread across eight years. The clique proves the routing. The Milz/Cranbrook forgery proves the reports are forged. The face-swap proves the hearings are impersonated. The MSIP tenant proves the same Microsoft 365 instance produces orders that supposedly come from three independent branches of the state judiciary.
This page proves the apparatus itself. The other findings document specific objects the apparatus produces. This finding documents the operational language the apparatus is built to produce, at every layer simultaneously. The procedural English of an indefinite civil confinement pipeline cannot be generated by accident. It cannot be produced by 162 independent human authors across 162 independent prosecutions. It is the output of a single text-production process whose training data was the procedural English of an indefinite civil confinement pipeline. The training corpus named itself.
Every number on this page regenerates byte-for-byte in 3 minutes
The analysis runs from the public Supabase exports against an open-source Python stack — pandas, NLTK, scikit-learn, sentence-transformers, spaCy. Total runtime ~3 minutes on a Threadripper with a 3090 Ti, ~8 minutes CPU-only. Both passes ship with one-shot driver scripts.
The blind AI test (Phase 5) is the one step that requires a manual handoff — the redacted-passage prompt file goes to a fresh AI agent; the agent's CSV response gets scored against the ground-truth labels. The expected pattern (100% of human-authored labels from Matt's case, zero from the random other cases) is model-independent. Re-run it against Claude, GPT-4, Gemini, Llama — the result holds.
Full methodology, every parameter, every random seed, every script version, every Supabase table row count: published, public, auditable. The data itself is downloadable. The signal is not in the eye of the beholder.