Dataset Verification

Corpus v3.0 — Publication Snapshot

Designed to Detect, Not to Correct is computed from a frozen, immutable snapshot of the Inside the Work oversight corpus: publication version 3.0, frozen July 6, 2026. This page publishes the verification record referenced in Appendix A — the snapshot's SHA-256 hashes, the 95-column schema, and the denominator register — so the freeze can be verified without access to the proprietary dataset itself.

Frozen Files

oig-dataset-PUBLICATION-FROZEN-v3.0-20260706.tsv — 1,219 report records × 95 columns

SHA-256: 8bb50bbdb23802feb4632028b36ec6378177e4187f282285a485684519bcbbbd

oig-recommendations-PUBLICATION-FROZEN-v3.0-20260706.tsv — 3,437 recommendation records

SHA-256: 8d7df5a675be9d6897b2f648ffb58ab3e161f328cdae51525fc2693a9bef7801

The frozen files are never edited. Corrections create working versions with changelogs and a new manifest at the next freeze. Any distributed copy can be authenticated by computing its SHA-256 hash and comparing it to the values above; any byte-level change produces a different hash.

The 95-Column Schema

Roughly fifty core classification fields across five domains, extended by provenance, source-verification, and derived fields that bind every row to its source document. Locked in fixed order:

Report_ID, OIG_Office, Report_Number, Report_Title, Publication_Date, Fiscal_Year_Issued, Audit_Type, Entity_Level, Entity_Name, Scope_Type, Facility_or_VISN, Program_Area_Primary, Program_Area_Secondary, Recommendation_Count, Repeat_Finding, Prior_Report_Referenced, Contract_Involved, Contractor_Name, IT_System_Involved, Primary_System_Name, Secondary_System_Name, Dollar_Impact_Stated, Dollar_Impact_Value, Financial_Impact, Improper_Payment_Explicit, Financial_Reporting_Impact, Appropriations_Issue_Type, GreenBook_Relevant, GreenBook_Component, Failure_Mode, Governance_Failure, Governance_Failure_Type, Monitoring_Breakdown, Accountability_Failure, Control_Design_Failure, Control_Execution_Failure, Data_Reliability_Issue, Documentation_Gap, Staffing_or_Resource_Factor, Training_Gap, Contractor_Oversight_Failure, Systemic_Risk, Escalation_Level, CFO_Relevance, CFO_Function_Affected, Executive_Relevance_Level, Root_Cause_Primary, Governance_Pattern_Primary, Confidence_Level, Content_Value_Score, Final_Decision, Content_Citations, Period_Audited_Start, Period_Audited_End, Detection_Lag_Years, Report_Is_FollowUp, Prior_Report_Linked, Prior_Report_IDs, Recurrence_Confirmed, Time_Since_Prior_Years, Condition_Summary, Criteria_Type, Criteria_Citation, Cause_Narrative, Effect_Narrative, Sample_Size, Error_Rate, Projected_Population, Dollar_Context_Type, Dollar_Recoverable, Consequence_Type, Population_Affected, Context_Flags, Acting_Official_Present, Mgmt_Response_Present, Overall_Concurrence, OIG_Response_Adequacy, Governance_Pattern_Confirmed, Pattern_Tier, Failure_Layer, Maturity_Stage, Pass1_Corrections, Pass2_Confidence, Pass2_Source_Quotes, Disposition_Tag, Source_File_ID, Source_File_Name, Source_Match_Method, Source_File_Hash, Concurrence_Note, Response_Adequacy_Note, Provenance_Note, Dollar_Recoverable_Note, Root_Cause_Note, Pattern_Note

The Denominator Register

Every analytic subpopulation in the book, with its exclusion rule.

1,219 — all reports in the frozen corpus.

1,006 — core report records: reports whose confirmed-pattern classification is anything other than No Adverse Finding; the 213 no-adverse-finding reports are analyzed separately.

1,001 — the index calibration set: core report records with a mapped control family and a classified pattern tier; five core reports lack a mapped family.

990 — failure-mode-coded report records: sixteen core reports lacked the detail to classify; unclassifiable is distinct from the coded value Other (five reports), which denotes a determinable mode outside the seven named categories.

947 — the two-tier comparison set: 609 structural plus 338 transactional; fifty-nine core report records are hybrid (55) or untiered (4) and are excluded from structural-versus-transactional contrasts.

695 — non-follow-up core report records. Follow-up here is the coded flag for a report commissioned against a linked prior report — distinct from the Follow-Up Review report type in the corpus profile (16 reports); the 94 percent recurrence figure for follow-ups rides on the flag, not the type.

690 — report records carrying at least one extracted recommendation.

576 — report records with at least one concurred recommendation (recommendations file).

549 — report records whose report-level management response was full concurrence (Concur or Concur All).

409 — report records with the dollar-impact indicator flagged; 406 of them carry a parsable positive dollar value, which is why two figures appear in different places.

273 — report records with a recoverable-dollar figure.

258 — report records with a curated, finding-attributable dollar value under the Gate 6 classification.

88 — report records with at least one recommendation in the legacy Closed/Resolved status aggregation.

46 and 42 — those 88 report records scoring at or above and below 60 on the index.

21 — closed-and-recurred report records carrying an attributable dollar value at the $10 billion per-finding cap.

Method

The corpus was assembled by exhausting the public report libraries of the issuing offices within the inclusion criteria, seeded from VA OIG and expanded across agencies — a census of what those libraries publish that meets the criteria, not a sample. Each report was processed through a five-pass extraction engine into the 95-column record with controlled vocabularies, then frozen; the publication snapshot is immutable and hashed, and every figure in the book computes from the frozen files. Full methodology, coding disclosures, and inter-rater agreement coefficients appear in Appendix A of the book.

Verification questions: joseph@insidethework.com