RecordStrata.ai

Method / Note 1

How RecordStrata Builds a Deep-Read

A source-first workflow for turning fragmented investigation and enforcement records into auditable case intelligence.

At a glance

  • The source collection is defined before the analysis begins.
  • The as-published document is preserved and identified by version and file hash.
  • Agency findings, interview statements, analyst synthesis, silence, and uncertainty remain distinct.
  • Dataset joins are documented and ambiguous matches are not silently accepted.
  • A human reviewer approves the evidentiary labels, quotations, joins, and final publication.

A Deep-Read is a method, not a longer summary

A conventional summary reduces a document. A RecordStrata Deep-Read reconstructs a record. It identifies the incident, preserves the source edition, builds the chronology, extracts the agency findings, separates contributing and non-contributing factors, maps equipment and management knowledge, identifies enforcement actions, and states what the record does not establish.

That distinction matters because a fatality investigation is rarely self-contained. The report may describe the accident and print the enforcement language, while proposed penalties, contest history, docket disposition, controller history, and comparable incidents reside in separate datasets or later proceedings. A Deep-Read does not pretend that one document answers every question. It shows what each source can support and where the next join must occur.

The governing discipline is simple: do not make the record say more than it says, and do not leave its material limits hidden.

The ten-stage workflow

  1. Define the research unit

    Identify the incident, mine, operator, contractor, equipment line, cited standard, or enforcement question that the Deep-Read will address. The research unit determines which sources are in scope and which are not.

  2. Register the source collection

    Record the filename, source location, retrieval date, page count, publication or revision date, and cryptographic hash. Preserve the source exactly as obtained.

  3. Establish canonical identity

    Use stable identifiers – especially mine ID, contractor ID, accident date, operator name, and report number – to distinguish the event from similarly named mines or entities.

  4. Segment the document

    Separate overview, accident narrative, investigation, discussion, equipment, examinations, training, root causes, conclusion, enforcement actions, and appendices without losing page boundaries.

  5. Build the chronology

    Reconstruct the event sequence before interpreting causation. Record who did what, when, and on what stated basis. Keep reported times, video-derived times, and inferred intervals distinct.

  6. Classify the evidence

    Mark agency findings, interview accounts, documented conditions, corrective actions, analyst synthesis, unresolved questions, and source silence using controlled evidence labels.

  7. Map equipment and responsibility

    Identify the operator, contractor, controller, manufacturer, model, component, maintenance history, examinations, training, supervision, and any prior notice reflected in the source.

  8. Construct the enforcement map

    List each order, citation, safeguard, standard, recipient, and stated basis. Then identify the documented join keys needed to reach violation, penalty, docket, or decision records.

  9. Test the limits

    Identify contradictions, facts that could not be determined, matters the report did not address, and questions that require another source. Do not convert absence into a negative finding.

  10. Verify and release

    Verify every quotation and page anchor, confirm critical joins, review the analytical language, and release the report only after human approval.

1. Define the source collection before drawing conclusions

The starting point is not a theory of fault. It is a source manifest. The manifest identifies the documents and structured records the analysis may use. For a fatality matter, the initial collection commonly includes the as-published investigation report and the corresponding Part 50 accident record. Depending on the assignment, it may later expand to violations, penalty dockets, decisions, mine and controller histories, equipment records, or comparable incidents.

This scope control prevents two common errors. First, it prevents an analyst from importing unverified facts merely because they appear plausible or are available elsewhere. Second, it makes every limitation visible. A reader can see whether an answer rests on the investigation report alone or on a larger linked record.

2. Preserve the as-published source

The original PDF is retained as an immutable source object. Extracted text, normalized fields, and later annotations are separate representations. They can be corrected when an extraction error is found; the original source is not overwritten.

The two published sample Deep-Reads identify the report number, page count, and MD5 hash of the edition used. That practice is not decorative. It allows a later reviewer to determine exactly which edition supported the analysis, particularly where an agency replaces, corrects, or republishes a report.

3. Establish canonical identity and preserve disagreement

Structured Part 50 accident records are treated as canonical for core event identity fields such as accident date, mine identity, and classification. The investigation report supplies the narrative, findings, quotations, and detailed context. When the two disagree, both values are preserved. The canonical field governs the normalized record, while the printed value remains visible as a source fact.

4. Build the chronology before interpreting causation

Chronology is the control structure for the analysis. A Deep-Read distinguishes time stamps printed in the report, times derived from video, statements from interviews, and intervals calculated from those facts. It does not merge them into one narrative voice.

5. Separate the agency finding from the analytical frame

A Deep-Read may organize the record around themes such as management knowledge, prior notice, coordination failure, equipment posture, or allocation of responsibility. Those themes are RecordStrata synthesis. The underlying causal finding remains attributed to MSHA and is reproduced or paraphrased with a page reference.

This distinction is visible in the sample reports. In Nevada, the phrase "management-knew record" is an analytical frame supported by the report’s statements about a procedure management chose not to follow, daily management travel through the area, and a longstanding unreported camera condition. In Bear Run, "the dual citation" is an analytical frame built from the agency’s parallel citations to the operator and contractor. The labels help the reader see the pattern; they do not replace the official finding.

6. Treat factor verdicts as separate evidentiary outcomes

The Deep-Read does not reduce every examined factor to "present" or "absent." It preserves the agency’s actual verdict: contributed, did not contribute, compliant, could not determine, or no verdict stated. These outcomes are not interchangeable.

For example, the Bear Run report states that investigators could not determine whether wind contributed. It separately states that certain Part 48 training gaps did not contribute. The first issue remains open; the second received an express non-contribution finding. A Deep-Read carries that distinction into every downstream table, summary, and Ask the Record answer.

7. Map equipment without manufacturing a defect theory

Equipment is analyzed at the level the source supports: manufacturer, model, component, condition, maintenance history, examination, reported defect, testing, corrective action, and agency conclusion. The presence of equipment in an accident does not itself establish a product defect.

Bear Run illustrates the discipline. The shaker and dewatering screens were identified, but the causal findings concerned blocking, task training, and compliance with the Surface Safety Handbook. The telehandler was not identified by manufacturer or model. The Deep-Read therefore described an uncontrolled disassembly and source silence about the telehandler; it did not invent an equipment-defect claim.

8. Construct the enforcement map and document the join

Each enforcement action is treated as its own record: instrument, statutory authority, cited standard, recipient, and stated basis. If citation or order numbers are absent from the investigation report, the Deep-Read says so. It then identifies the join keys that should be used to locate the structured violation or penalty record.

In Nevada, the proposed join uses Mine ID 26-02573, the cited standard, and the accident-date window. In Bear Run, the operator-side actions must be keyed to Mine ID 12-02422, while the contractor-side actions must be keyed to Contractor ID C3953. Keeping those tracks separate is essential because the same standards were cited against different recipients and their penalty proceedings may run separately.

9. State what the record does not tell the reader

Every Deep-Read contains an express limitations section. The purpose is not caution for its own sake. It prevents the reader from mistaking a narrow source for a complete factual universe.

  • If the report does not print penalty amounts, the Deep-Read directs the reader to the violations and penalty datasets.

  • If investigators could not determine whether a condition contributed, the question remains open.

  • If a manufacturer or model is not identified, the field is marked not stated – not unknown in the world.

  • If accident damage prevented testing, the report cannot support an affirmative claim about the pre-accident condition.

  • If the report is silent about a party, the silence is not treated as clearance.

10. Human review is the release gate

Automation can accelerate text extraction, field suggestions, quotation matching, retrieval, and comparison. It does not release the report. A human reviewer confirms the source edition, chronology, evidence labels, quotations, page anchors, joins, contradictions, limitations, and analytical language before publication.

The two sample Deep-Reads state the standard directly: quotations were machine-verified against the source PDF and page-anchored automatically, and the extraction record was then reviewed line by line by a human reader. The value of automation lies in consistency and coverage; the value of review lies in judgment, restraint, and accountability.

The standard Deep-Read structure

Section Purpose
Case at a glance Identifies the report, mine, operator, contractor, sector, event date, principal enforcement posture, and exact source edition.
What happened Reconstructs the event chronology and separates source facts from RecordStrata synthesis.
What the agency found Presents causal findings and root causes with page-level support.
Factors examined Preserves the agency’s verdict for each factor, including non-contribution and inability to determine.
Equipment and OEM posture Explains what is known about the equipment, testing, maintenance, component conditions, and any limits on defect analysis.
Enforcement map Lists instruments, standards, recipients, stated basis, and join keys.
What the record does not tell you Makes silence, conflict, missing fields, and unresolved questions explicit.
Dataset note Explains canonical fields, source joins, publication lag, quotation verification, and review status.

What a Deep-Read does not do

  • It does not determine civil liability, damages, admissibility, or the legal effect of an agency finding.

  • It does not convert every citation into proof of causation.

  • It does not treat a search hit as a verified fact.

  • It does not resolve a contradiction merely to create a cleaner narrative.

  • It does not imply that the source collection is complete when later proceedings or outside records may exist.

  • It does not replace counsel, an expert, or independent verification against the original record.

The result

A useful Deep-Read is concise enough to guide an early case decision and rigorous enough to survive close review. Its principal deliverable is not the prose alone. It is the evidentiary architecture underneath the prose: a defined source collection, preserved editions, structured fields, controlled labels, documented joins, page anchors, and a visible account of what remains unresolved.

The report tells the reader what matters. The source architecture lets the reader verify why it matters.

Source notes

The examples in this Method Note are drawn from the following records. Page references are to the cited as-published or Deep-Read edition.

1. "A Berm That Management Chose Not to Build," MSHA Fatality Deep-Read Series No. 1, Nevada Gold Mines, LLC, dated September 7, 2026, especially pp. 1-4.

2. "Eleven Minutes: A Coordination Failure, Cited Twice," MSHA Fatality Deep-Read Series No. 2, Peabody Bear Run Mining LLC, dated September 7, 2026, especially pp. 1-5.