RecordStrata.ai

Method / Note 3

How RecordStrata Uses AI - and Where Human Review Begins

A transparent division of labor between machine assistance, source authority, and accountable human judgment.

At a glance

  • The official source remains authoritative; the model is never the evidence.
  • AI operates inside a defined task and, for Ask the Record, inside a defined source collection.
  • Human reviewers make the decisions that require evidentiary classification, context, or professional judgment.
  • Uncertainty, contradiction, and source silence are preserved rather than completed by prediction.
  • The published product states the role of AI and the point at which human review occurred.

AI is an analytical tool, not an authority

RecordStrata uses artificial intelligence because the underlying record is large, repetitive, and fragmented. Investigation reports contain recurring sections, structured datasets contain millions of rows, and similar hazards may be described in inconsistent language across years. AI can accelerate the work of locating, extracting, comparing, and organizing that material.

The value of the product, however, does not come from asking a model for an answer and accepting the response. It comes from the combination of source control, page-level provenance, evidence labels, documented joins, and human review. The official record remains the authority. AI helps the analyst work through the record; it does not replace the record.

The model may propose. The source must support. The human reviewer must approve.

Tasks AI may assist with

Task AI-assisted function Control
Document segmentation Identify recurring sections such as overview, discussion, root cause, conclusion, enforcement actions, and appendices. Page boundaries and headings remain tied to the original PDF.
Text and field extraction Propose dates, mine IDs, entities, equipment, standards, findings, corrective actions, and quoted passages. Fields are not final until checked against the source.
Passage retrieval Locate potentially relevant passages in response to a research question. A search match is treated as a candidate, not a finding.
Quotation matching Compare proposed quotations with page-level extracted text and flag mismatches. Human review confirms the visible source and context.
Entity and record comparison Suggest possible matches across accident, violation, penalty, mine, contractor, and controller records. Ambiguous joins require documented human confirmation.
Cross-record comparison Group potentially similar incidents, equipment conditions, standards, or corrective actions. Similarity is reviewed for factual and legal relevance.
Draft support Organize supported facts into a chronology, table, or draft narrative. The reviewer controls evidentiary language and final publication.
Ask the Record answers Generate a response from the selected source collection and attach supporting pages. The system must remain within scope, surface uncertainty, and permit source verification.

Tasks reserved to human review

Some decisions cannot be reduced to extraction or similarity. They require context, restraint, and accountability. A human reviewer must control the following stages before a commissioned Deep-Read is released:

  • Defining the source collection and deciding whether it is sufficient for the stated assignment.

  • Confirming the incident, mine, operator, contractor, and temporal identity of the matter.

  • Deciding which structured field is canonical when sources differ, while preserving the conflicting value.

  • Classifying a statement as an agency finding, interview account, documented fact, synthesis, inference, silence, contradiction, or open question.

  • Confirming causation language and distinguishing contributing, non-contributing, compliant, and indeterminate findings.

  • Confirming equipment and entity joins, especially where names, ownership, contractor IDs, or dates create more than one candidate match.

  • Determining whether an analytical label fairly describes the supported pattern without being attributed to the agency.

  • Reviewing every published quotation and page anchor.

  • Approving the limitations section and ensuring the report does not imply facts outside the source collection.

  • Approving corrections, revision notices, and client communications about material changes.

What AI must not do

Prohibited behavior Why it is prohibited
Determine civil or criminal liability The source may inform legal analysis, but liability depends on law, evidence, procedure, and professional judgment outside the model’s role.
Convert a citation into proof of causation An enforcement action and a causal finding may overlap, but they are not automatically the same proposition.
Choose between conflicting source statements without disclosure A cleaner answer is not more accurate when the underlying record is inconsistent.
Treat silence as exoneration or proof of absence A report may omit a fact because it was outside scope, unavailable, or not material to the agency’s investigation.
Supply missing names, dates, manufacturers, or motives from probability A likely completion is still unsupported if it is not in the source collection.
Present an analyst label as agency language Themes such as "management-knew record" or "dual citation" must remain visibly attributed to RecordStrata.
Use the open internet inside a source-bounded answer without disclosure The user must know exactly which collection supports the response.
Release an unreviewed Deep-Read as final Commissioned reports require accountable human approval of evidence, joins, quotations, and limitations.

The review boundary in practice

  1. Machine-assisted intake

    The system registers the document, extracts page-level text, identifies sections, and proposes fields and candidate quotations.

  2. Automated checks

    The system tests quotation matches, required fields, identifier formats, page anchors, and potential inconsistencies.

  3. Analyst review

    A human compares the proposed extraction with the source, confirms labels, reconstructs chronology, and reviews joins.

  4. Editorial review

    The report is checked for attribution, overstatement, omissions, internal consistency, and a clear limitations section.

  5. Release record

    The publication records the source edition, review status, first publication date, last reviewed date, and any later corrections.

Why the boundary matters: three examples

Ask the Record: source-bounded generation

For Ask the Record, the collection boundary is part of the answer contract. The user selects the Nevada Gold Mines or Bear Run demonstration collection. The system may retrieve and synthesize only from the documents listed in that collection’s manifest. It must identify the document and page supporting the response and state when the collection does not contain the requested fact.

This design limits one of the principal risks of general-purpose generative systems: the tendency to produce a plausible completion when the evidence is incomplete. In a source-bounded interface, "the record does not state" is a successful answer when that is what the collection supports.

Public demo and private matter rooms

The public demonstration should contain only preselected public records. It should not accept confidential client uploads. A private matter room requires a separate access, storage, retention, and engagement framework. The existence of source-bounded answer generation does not by itself create confidentiality, privilege, or an attorney-client relationship.

For private work, the source manifest should identify client-provided materials separately from official records, preserve access controls, and record which collection supported each answer. Any expansion of the source collection should be visible to the user and versioned.

Human review must be evidenced, not merely asserted

A credible claim of human review should correspond to an actual review record. For each commissioned Deep-Read, RecordStrata should be able to identify the reviewer, review date, source edition, review status, material corrections, and whether all critical quotations and joins were confirmed.

The sample Deep-Reads already state a concrete standard: character-level machine verification of quotations, automatic page anchoring, followed by line-by-line human verification of the extraction record. Future publications should preserve that specificity and should not broaden the claim beyond what the production workflow can demonstrate.

The objective

The purpose of this division of labor is not to minimize AI. It is to use AI where it is strongest – scale, consistency, retrieval, and comparison – while reserving evidentiary judgment and publication responsibility for a human reviewer. That combination produces a research product that is faster than purely manual review without asking the user to trust an untraceable machine conclusion.

Source-bounded AI is valuable precisely because the boundary is visible, testable, and enforced.

Source notes

The examples in this Method Note are drawn from the following records. Page references are to the cited as-published or Deep-Read edition.

1. "A Berm That Management Chose Not to Build," MSHA Fatality Deep-Read Series No. 1, Nevada Gold Mines, LLC, dated September 7, 2026, especially pp. 2 and 4.

2. "Eleven Minutes: A Coordination Failure, Cited Twice," MSHA Fatality Deep-Read Series No. 2, Peabody Bear Run Mining LLC, dated September 7, 2026, especially pp. 3 and 5.

3. MSHA Report of Investigation, Young Mine, Nyrstar Tennessee Mines, fatal fall of roof or back accident of July 12, 2025, especially PDF p. 6.