CRIMINAL HISTORY'S EXPIRATION DATEDATA EXPLORER

How fast does a criminal charge's predictive power actually fade?

Explore the decay pattern behind a 16.6-million-person, 40-year study of recidivism risk — how the model is built, how the pattern holds across groups, and how it converges back to baseline over time.

28%
Predicted recidivism risk in year 1 after a charge
−51%
Drop in that risk by year 7
9–10
Years until risk converges to the population baseline
63%
Of people charged were never charged again
FIGURE 1

Why the model matters: GLM vs. GLMM

One person can rack up several charges over the study's 40 years — the average is 4.28 per person. A standard model (GLM) treats every person-year as independent and ignores that clustering; a mixed model (GLMM) accounts for it with a random effect per person. Modeling that clustering changes the picture substantially.

Why this matters

The widely repeated claim that roughly two-thirds of people released from prison reoffend traces back to older cohort studies that treat every charge as an independent, uncorrelated event. Those models can't tell one person with ten charges apart from ten different people with one charge each — a small share of highly active "frequent fliers" ends up inflating the apparent share of the whole population that reoffends, and slows the apparent decay of risk over time. The GLMM corrects for this by explicitly modeling each person's charges as clustered together, and the result is a lower starting risk and a faster decline — a more accurate picture of how risk actually behaves for a typical person. That gap is most visible above in the probability view; switching to relative risk rescales both curves to the same starting point, which is useful for comparing decay shape but hides just how different the two models' real-world intercepts are.

FIGURE 2

Compare decay curves across groups

Every dimension below shows the same underlying pattern: groups start at different levels of risk, but risk fades at a similar rate for almost everyone. Switch views to see it both ways — as raw year-1 probability, or normalized so every line starts at the same point.

FIGURE 3

Convergence to the general population base rate

Each state's decay curve eventually meets the ordinary background rate of a new charge in that state's general population — the point where a prior charge stops meaningfully elevating someone's risk. The dashed line marks each state's real base rate; the marker on each curve is that state's real, published convergence year.

State base rates & convergence years — Table 7
StateBase rateConvergence year95% CI

Does this pattern hold across historical periods?

The study spans 1980–2020. Comparing charges from 1980–2000 to charges from 2001–2020 (each restricted to a common 20-year follow-up window) shows more recent charges start at lower baseline risk and fade faster.

Why this matters

Both the starting point and the decay rate shift with the era a charge was filed in: charges from 2001–2020 start lower and fade faster than charges from 1980–2000. That's consistent with recidivism risk genuinely declining over the study period, not just an artifact of which years happen to be in the sample. Practically, it means estimates built on more contemporary data are likely a more accurate picture of current recidivism risk than estimates anchored to older cohorts — a risk model trained mostly on 1980s–90s charges will tend to overstate both how likely someone is to reoffend today and how long that elevated risk lasts.

How these curves were built

Figure 1 (GLM vs. GLMM): Table 2 reports both specifications' lookback odds ratios (GLM 0.44, GLMM 0.39, each per 1-SD/4.85-year increase in lookback), converted here into an illustrative per-year decay rate. Year-1 starting probabilities (GLM “exceeding 20%,” GLMM “approximately 5%”) are taken directly from the paper's own description of Figure 4. The GLMM's lower, steeper curve reflects a subject-specific (conditional) prediction once each person's repeat charges are modeled as clustered rather than independent — the paper retains the GLMM for every subsequent analysis for this reason.

Figure 2 (group comparisons): decay rate derived from the published interaction-model coefficients in Tables 3–5 and Appendix Tables E-1/E-2. Year-1 starting probability: each group's real published odds ratio applied multiplicatively, in odds-space, onto the study's year-1 population probability of 28%: odds_group = odds_reference × OR, then converted back to a probability. Where the paper reports a genuine null (violent vs. non-violent offense type), the lines are drawn essentially on top of each other.

Figure 3 (convergence): base rates and convergence years are Table 7's real, published values, not illustrative — unchanged from the paper. Because the convergence analysis is estimated on a different scale (population-average GLMM predictions, the same lower-baseline specification as Figure 1's GLMM curve) than Figure 2's charts (anchored to the 28% observed population curve for cross-group comparability), each state's curve here is scaled so it passes exactly through that state's own real base rate at its own real convergence year, using the same real per-state decay-rate multiplier from Table 3 as Figure 2's states chart. The historical-period comparison uses Table 6: year-1 probability from the additive model's main-effect odds ratio (0.31) applied the same way as every Figure 2 dimension; decay rate from the interaction model's odds ratios (reference-period slope 0.54, interaction term 0.67), estimated on Washington data only.

All charts on this page are illustrative: built from real, published model estimates, but not literally the paper's fitted values at every year for every subgroup. They're meant to make the paper's findings something you can see and compare yourself.