How SovGuardAI Scores Documents
Methodology & Score Interpretation
1. What the Correspondence Score Is
The correspondence score is a comparison, of court-identified and research-determined behavior, ideology, tactics, language, and other indicators previously labeled "sovereign citizen," with the client-submitted document.
There is no DNA test, no blood test, no electrical measurement, and no self-identification or self-denial that might label one as a sovereign citizen. However, law enforcement, the judicial system, and peer-reviewed research have arrested, adjudicated, and studied "sovereign citizens."
Thus, there is no objective method that would result in the labeling of an individual as being a sovereign citizen. Nonetheless, persons are daily labeled sovereign citizens by law enforcement, by judges, and by peer-reviewed researchers. SovGuardAI looks at the ideology, tactics, behavior, language, and other indicators of those previously labeled sovereign citizens in peer-reviewed research and by the American court system. SovGuardAI provides a score from 0 to 100 representing the correspondence of that which has previously been identified and labeled sovereign citizen, with the client-submitted document. See score breakdown in Section 4 labeled Score Tiers. An overall correspondence score is provided and in addition, each occurrence of an identifiable item is located by page and paragraph number, verified directly against the source document. Where an exact location can't be independently confirmed, that's noted rather than guessed at. Also, importantly, provided is a specific reference as to how the identified tactic has been ruled on in an American court of law.
2. How the Score Works
The correspondence score answers one question: how closely does the language, structure, and content of this document resemble patterns identified in peer-reviewed research and American court decisions as associated with sovereign citizen ideology?
The score runs from 0 to 100 across four tiers (None, Low, Moderate, and High). A higher score reflects greater density and variety of known pattern matches. A lower score reflects their absence — not their impossibility. The taxonomy is actively maintained, but no tool can detect patterns that have not yet been documented.
The correspondence score measures how closely a document resembles patterns documented in the taxonomy. It is an assessment, not a legal determination, and it is not calculated from a fixed formula or a published weighting table. Section 3 describes precisely how the number is arrived at.
Scores should be interpreted as indicators, not measurements. The same document may score within a range across separate analyses — this is a known characteristic of large language models, not a flaw specific to SovGuard. Princeton University's 2026 HAL Reliability Evaluation, studying 14 AI agents across multiple benchmarks, confirmed that outcome consistency remains a key challenge across all frontier AI systems, finding a substantial gap between what models can do and what they do consistently. SovGuard is designed around this reality: scoring thresholds and tactic counts are used as floors, not precise targets, and attorney review of the underlying document remains essential.
A Judicial Document badge appears when the analyzed document is a court opinion, Report and Recommendation, or other judicial document reproducing sovereign citizen arguments in the context of rejection. Scores on judicial documents reflect the density of sovereign citizen language present in the record, not an assessment of the court's ruling.
3. How the Number Is Assigned
A correspondence score is produced in four stages. We describe them separately because they have different bases, and because the distinction matters when interpreting a result.
Stage 1 — The taxonomy: what the system looks for
SovGuardAI evaluates documents against a structured taxonomy of sovereign citizen tactics, paired with a curated library of federal and state decisions addressing each. This taxonomy derives from the research of Dr. Christine Sarteschi — including her court-records database and her published academic work on sovereign citizen ideology — supplemented by peer-reviewed literature, government and non-governmental threat assessments, and primary court filings and opinions. This layer is the substantive foundation of the system and reflects domain expertise, not automation.
Stage 2 — Detection: what is found in a given document
The submitted document is analyzed against that taxonomy by a large language model. The model identifies specific phrases in the document and associates them with named tactics. Detected tactics and the phrases supporting them are reported in full, so that every finding can be traced back to language actually present in the submission.
Stage 3 — The score: how the number is set
The numeric score is assigned by the model on the basis of the tactics and phrases it has detected, reflecting their number, density, and severity. It is not computed from a fixed formula or a published weighting table. It is an assessment, and it should be read as one.
Two mechanical rules operate on top of the model's assessment:
- If no tactics are detected and no flagged phrase is tied to a known tactic, the score is set to 0.
- In the rare case where the model returns a score our system cannot read as a number, a fallback substitutes the count of distinct detected tactics multiplied by nine, capped at 100. This is a backstop against malformed output, not the normal operating path.
Stage 4 — The tier: how None, Low, Moderate, or High is set
The tier is calculated by our software from the final numeric score using the fixed thresholds listed in Section 4. The tier is not supplied by the model; where a model-stated tier would be inconsistent with the score, the score governs and the tier is recalculated.
Those thresholds are reporting bands chosen for clarity of presentation. They are not derived from a validated statistical model, and they do not express a probability that a document is, or is not, any particular thing.
Proprietary elements
The contents of our analytical prompts, the internal structure of the taxonomy, and our detection patterns are proprietary and are not published. The mechanism described above is not.
4. Score Tiers
| Score | Color | Label |
|---|---|---|
| 0–10 | None — Below detectable limits | |
| 11–40 | Low — Present but not prevalent | |
| 41–70 | Moderate — Notable ideology present | |
| 71–100 | High — Strong sovereign citizen ideology detected |
The thresholds in this table are authoritative. Narrative summaries generated within an individual report may characterize score ranges in approximate terms; where that narrative language differs from this table, this table governs.
5. What the Score Is Not
- —Not a statement about a person. A document can score high because someone copied a template without understanding it.
- —Not a legal conclusion. A qualified reviewer must decide how to respond to the findings.
- —Not a threat assessment. A high score says nothing about whether someone will act.
- —Not exhaustive. Sovereign citizen tactics evolve. The SovGuardAI taxonomy is actively maintained as new patterns are identified in peer-reviewed research and American court decisions, but no tool can detect what has not yet been documented.
What SovGuard Does Not Do
- Provide legal advice or legal opinions
- Make determinations about any individual's intent, competence, or legal status
- Substitute for attorney review of the underlying document
- Guarantee detection of all sovereign citizen tactics in all documents
- Verify the authenticity of submitted documents
- Recommend specific filings, motions, or legal strategy in response
- Predict future behavior, such as whether a filer will escalate, reject counsel, or continue filing
6. How to Use This Tool
SovGuardAI is a triage instrument. It surfaces documents that resemble known sovereign citizen filings so that qualified reviewers can make informed decisions about how to respond. It is not a substitute for professional review.
A None or Low score is not a clean bill of health. The taxonomy is actively maintained, but no tool can detect patterns that have not yet been documented.
The Taxonomy
SovGuard's detection is built on a taxonomy derived from peer-reviewed academic research, published court decisions, and analysis of primary source filings — see Section 3, Stage 1, for the sources it draws on. The taxonomy covers the major sovereign citizen ideological clusters. The taxonomy is updated as new patterns emerge in case law and court filings.
Document Quality and Limitations
Detection quality is affected by document quality. Native text PDFs produce the most reliable results. Heavily scanned or handwritten documents may limit phrase extraction, particularly where key language appears in degraded form. Where OCR degradation is detected, SovGuard notes it in the analysis and recommends manual review of the original filing.
SovGuard does not detect tactics that are paraphrased or summarized rather than asserted. A court order that describes and rejects sovereign citizen arguments will score differently than a pro se filing asserting those same arguments — this is by design.
7. The Model
SovGuard uses Anthropic's Claude Sonnet, a frontier large language model, via direct API integration. The model is guided by a structured system prompt embodying the tactic taxonomy and scoring architecture. A threshold-based stress test suite is run against a corpus of known documents before any prompt changes are deployed, so that updates do not silently degrade detection on document types already in the suite.
That suite is a regression guard rather than an accuracy measurement. It compares each run against an expected result set in advance; it does not establish whether that expected result is correct. Precision and recall figures require human-adjudicated ground truth — a labeled corpus in which qualified reviewers, rather than the model, establish the correct answer for each document — which the suite does not provide. All testing described here is internal. The score is intended as triage and is not offered as a measurement or as evidence.
A Note on Consistency
No AI system produces identical outputs on every run. SovGuard's scoring architecture is designed to be robust to this: threshold-based pass/fail criteria, required tactic detection floors, and multi-run validation are used in development. In production, scores should be treated as indicators within a range, not precise measurements. The tactic list and flagged phrases are more stable across runs than the numeric score and should be weighted accordingly in attorney review.
A Note on Case Citations
SovGuard does not generate case citations. The case name, court, reporter citation and date of every authority in a report are reproduced from a fixed, manually curated library rather than produced by the model, so the fabricated-citation failure mode documented in generative legal research tools does not arise in the same way. That is a statement about architecture, not a measurement of accuracy.
Citations
The HAL Reliability Evaluation referenced on this page is: Rabanser, S., Kapoor, S., Kirgis, P., Liu, K., Utpala, S., & Narayanan, A. (2026). Towards a Science of AI Agent Reliability. arXiv:2602.16666. Princeton University.
The fabricated-citation failure mode in generative legal research tools is documented in: Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., & Ho, D. E. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362.
© 2026SovGuard AI. All rights reserved. · Not a law firm. Not legal advice.