Menu
OEE Data Retention Strategy: How Long to Keep Raw vs Aggregated and Why It Matters Later

OEE Data Retention Strategy: How Long to Keep Raw vs Aggregated and Why It Matters Later

Keep too little and ML training is impossible. Keep too much and storage costs explode. A practical retention strategy for OEE data.
OEE Data Retention Strategy: How Long to Keep Raw vs Aggregated and Why It Matters Later

Key takeaways

  • Data retention strategy = the rules for how long to keep raw vs aggregated OEE data.
  • Raw data: 30-90 days at full resolution is typical for active analysis.
  • Downsampled data: 1-3 years of minute-level aggregates for trend analysis.
  • Long-term aggregates: 5-7 years for benchmarking and audit.
  • Decide retention before deployment. Retroactively reconstructing data is impossible.

Short answer: An OEE data retention strategy specifies how long to keep raw sensor data, downsampled aggregates, and long-term summaries. Keep too little and future ML training and trend analysis are impossible. Keep too much and storage costs explode. A working pattern: 30-90 days raw, 1-3 years minute aggregates, 5-7 years long-term summaries. Decide before deployment because reconstruction is impossible. See also OEE vs Utilization.

Why retention matters

Three audiences need OEE data at different resolutions:

  • Operations. Real-time and last-30-days, full resolution.
  • Reliability and improvement teams. 1-3 years at meaningful resolution for trend analysis.
  • Future ML and benchmarking. Years of historical data, possibly downsampled.

A single retention policy cannot serve all three. A tiered policy can.

The tiered retention pattern

Tier 1: Raw, full resolution. 30-90 days. Used for active troubleshooting and short-term analysis. Storage cost real but bounded.

Tier 2: Minute-level aggregates. 1-3 years. Used for trend analysis, root cause investigation, OEE Pareto over time.

Tier 3: Hour-level aggregates. 5-7 years. Used for long-term benchmarking, year-over-year comparison.

Tier 4: Daily aggregates. Permanent. Used for executive reporting and audit.

What goes in each tier

Raw:

  • Every PLC tag at native cadence.
  • Every operator entry with timestamp.
  • Every reason code event.
  • Every quality measurement.

Minute aggregates:

  • OEE per minute.
  • Availability, Performance, Quality per minute.
  • Cycle counts per minute.
  • Downtime events with reason.

Hour aggregates:

  • OEE per hour.
  • Shift summary metrics.
  • Pareto-level loss categorization.

Daily aggregates:

  • Daily OEE.
  • Daily output.
  • Daily defects.

Why short-term raw is so valuable

Troubleshooting requires raw resolution. A 30-second cycle anomaly that occurred two weeks ago needs to be visible. Aggregated data loses the resolution to find it.

The 30-90 day window covers most active investigations. Older raw data is rarely accessed but expensive to keep.

Why long-term aggregates matter

Year-over-year comparison, benchmarking, ML training all need long history. But not at raw resolution. Daily or hourly aggregates retain the trend while reducing storage by orders of magnitude.

How to set tier boundaries

  1. Identify access patterns. What is queried, how often, at what resolution.
  2. Estimate storage cost per tier. Raw is expensive; aggregates are cheap.
  3. Balance cost against utility. Most plants land near the 30-90 day raw, 1-3 year minute pattern.
  4. Lock the policy. Retroactive changes lose data.

Common mistakes

1. Keeping raw forever. Storage cost grows without bound. Rarely accessed.

2. Aggregating too aggressively. Daily aggregates lose the resolution to investigate.

3. No downsampling pipeline. Manual aggregation breaks; automated tiered pipeline is essential.

4. Forgetting regulatory requirements. Some industries require longer retention for compliance.

The regulatory angle

Some industries (pharma, food, automotive) have batch record retention requirements that extend years beyond operational needs. Check before setting the policy.

Cloud vs on-prem implications

Cloud storage is cheap and elastic; on-prem storage is fixed-cost. The retention policy economics differ:

  • Cloud: retention tier costs are visible monthly. Easy to extend.
  • On-prem: upfront hardware cost. Extending retention is a capital decision.

Both work; the right tier policy considers the storage model.

What changes with ML in mind

ML training needs history with labels. If you anticipate training models on past data:

  • Keep raw longer (1+ years if possible).
  • Tag events with downstream outcomes (defects, failures) for supervised learning.
  • Preserve raw alongside aggregates so future models can use original signals.

This adds cost but enables future capabilities. Decide upfront.

Common mistakes

1. Setting policy without consulting future use. Cannot reconstruct what was not kept.

2. No periodic review. Storage cost drifts; access patterns change.

3. Mixed tiers in one storage. Tiers should map to different storage cost classes.

4. Skipping data quality validation before aggregating. Bad raw data produces bad aggregates that cannot be fixed.

How a modern OEE platform handles this

A modern OEE platform implements tiered retention automatically, with policy configurable per data type. Aggregation pipelines are managed by the platform.

Fabrico's OEE module supports configurable tiered retention with automated downsampling pipelines and explicit per-data-type retention policies.

See how Fabrico captures this automatically, explore OEE for manufacturing or book a demo.

Related reading

Frequently asked questions

How long should I keep raw OEE data?

30-90 days for most use cases. Longer if ML training is anticipated.

Should I keep all data forever?

Storage cost makes this impractical at scale. Aggregate older data instead.

Can I reconstruct lost raw data from aggregates?

No. Aggregation loses information that cannot be recovered.

What is a reasonable storage budget?

Cloud time-series storage is typically a small fraction of OEE platform cost. On-prem requires upfront capital.

Should retention differ by data type?

Yes. PLC tag streams need different retention than reason codes or operator entries.

Latest from our blog

Define Your Reliability Roadmap
Validate Your Potential ROI: Book a Live Demo
Define Your Reliability Roadmap
By clicking the Accept button, you are giving your consent to the use of cookies when accessing this website and utilizing our services. To learn more about how cookies are used and managed, please refer to our Privacy Policy and Cookies Declaration