Retention policies for training data

How long to keep the data a model was trained on, who decides, and how to delete it without losing the audit trail.

Elev8 Team 2 min read

Training data is the dataset most teams forget to put a retention date on. The source records have one. The copy that was pulled, cleaned and joined to train a model usually does not. It sits in a bucket under a service account until an auditor, or a subject access request, finds it.

A model does not need its training data forever

Once a model is trained, evaluated and signed off, the training set has two jobs left: rebuild the model if it ever has to be, and show an auditor what went in. Neither needs the raw rows kept indefinitely. Keep a versioned snapshot for as long as the model is in production, plus the period your regulator expects for the decisions it made. Then delete it.

  • Live model: keep the exact snapshot, read-only
  • Retired model: keep the snapshot for the decision-retention period, commonly six to seven years in UK financial services
  • Experiments that never shipped: delete within 90 days

Tie the retention date to the model, not the bucket

Set the date when the model goes live, in the model registry, next to its version. The deletion job reads it from there. When the model is retrained, the new snapshot gets its own date and the old one starts its countdown. Retention set on the storage bucket instead is how a live model’s data vanishes, or a retired model’s stays.

What to keep after deletion

A record that the snapshot existed: its hash, its row count, the register rows it drew on and the date it was deleted. That is what answers “what was this model trained on” after the data has gone.

Delete the rows. Keep the receipt.

Elev8 delivery team
model: claims-triage v3
trained: 2026-03-14
snapshot: ml-train/claims-triage/v3 (sha256 a91f…)
retain_until: live + 7y
deleted: pending

Keep reading

All resources