Every CPS investigation of alleged child abuse or neglect closes with a disposition summary — the caseworker's written account of what was reported, what was found, and how the case is being resolved. Michigan closes about 78,141 such investigations a year; 1,250,271 across the period this dataset covers.
What's in a disposition summary that isn't in the administrative record
Every case already carries a structured administrative row — case ID, dates, county, allegation type, disposition category. Those fields answer who, when, and where. They don't answer what did the caseworker actually see.
That part lives in the narrative:
- Whether a parent's substance use is current or in remission — not just whether "substance abuse" was checked on a form.
- Whether firearms are accessible in the home, and to whom.
- Whether housing instability is a soft mention or edging into homelessness.
- Whether domestic violence is between caregivers, against a partner, or witnessed by a child.
- The order, severity, and co-occurrence of concerns — details the structured record can't hold.
Narrative Insights reads every one of those narratives with small, local AI and extracts key risk indicators as structured data for analysis. The administrative record is not replaced — it is enriched with the qualitative context that was previously locked in prose.
- Investigation summaries enriched
- 1,250,271
- Per year on average
- ≈78,141
- Years of coverage
- 16
- Michigan counties
- 83
The volume above is the AI-solved problem — every narrative gets read. The mention rates on the tiles below focus on substantiated CPS cases (Categories I–III, n = 289,280), where the program-relevant signal lives. All-cases comparisons are shown beside each rate.
At a glance
Headline mention rates by topic — tap any tile to drill into the maps, trends, and top counties.
Highest in Ionia at 59.8%. rose from 39.0% (2010) to 41.6% (2024).
Highest in Marquette at 17.3%. held steady around 9.0%.
Highest in Lenawee at 39.1%. rose from 21.1% (2010) to 35.2% (2024).
Highest in Muskegon at 12.6%. rose from 5.0% (2010) to 8.2% (2024).
Highest in Muskegon at 13.8%. rose from 6.0% (2010) to 8.1% (2024).
Highest in Sanilac at 8.6%. rose from 6.4% (2010) to 12.3% (2024).
The problem this solves
A team of human reviewers cannot read ~312 disposition summaries every working day. No structured field captures whether a parent's substance use is current or in remission, whether firearms are accessible, whether housing instability is at the edge of homelessness. It's all in the prose. Left untouched, that information stays locked in free text and never surfaces in any dashboard, any report, or any program evaluation.
Small local AI reads every narrative — all 1,250,271 — and turns those free-text observations into structured data analysts can query, aggregate, and cross-tabulate. Program analysts then work with the enriched structured data alongside the fields they already have.
How this works — and doesn't
Five things to know about how Narrative Insights reads these narratives.
Not the giant AI systems in the news. Each risk signal is handled by a small, specialized program trained to spot one specific thing in a narrative. Small tools are enough because each job is well-defined.
Every piece of the pipeline runs on the lab's own computers. No narrative is ever sent to a cloud service, a commercial AI provider, or across any outside network. The data never leaves the building.
Because we don't rely on massive data centers, running the whole pipeline uses about as much electricity as a household appliance for a few hours. Comparable cloud-AI processing at this scale uses orders of magnitude more energy.
Every indicator has a written definition ("codebook") that human reviewers wrote. The AI applies that human-written standard consistently across every case. It's a tool for extraction at scale, not a decision-maker.
Every indicator is measured against a set of cases that human coders labeled first. If the AI doesn't agree with the human standard, it doesn't ship. If it drifts, it's pulled and re-trained.
Names, addresses, phone numbers, and case numbers are removed from every narrative before any AI touches it. Only aggregate county-by-year numbers appear in the dashboard; per-case labels stay inside existing data protections.
See the Methodology for how each indicator is built, checked, and released.
Explore the risk indicators
Methodology
How to use the tabs
Every topic tab packages three views of an indicator together: where in Michigan (a county map), year by year (a trend line), and top counties (ranked bar chart with a statewide reference). Tap any county on the map to swap the trend chart for that county's history vs. statewide.
Every rate on this dashboard is a mention rate — the share of case narratives in which the indicator was recorded. It is not a population-wide prevalence and it is not a clinical diagnosis. See Methodology for the full framing.