Automation Glossary • Tiered Historian Retention

What Is Tiered Historian Retention (Hot, Warm, Cold)?

Merobix Engineering • • 7 min read

Historian data is not all worth the same to keep close at hand. Yesterday's readings are queried constantly and need to come back instantly; data from three years ago is touched rarely and can afford to be slow to retrieve. Tiered retention exploits that difference by storing data on progressively cheaper, slower media as it ages. This guide explains the hot, warm, and cold tiers, how a retention policy moves data between them, and the trade-off between query speed, fidelity, and cost that tiering is designed to manage.

Back to Blog

Tiered Historian Retention in one line: Tiered historian retention stores time-series data on different classes of storage according to its age and how often it is accessed. Recent data sits in a hot tier on fast storage at full resolution, mid-age data moves to a warm tier that is often aggregated or compressed, and old data lands in a cheap cold archive, balancing query speed and fidelity against storage cost.

The Hot, Warm, and Cold Tiers

The hot tier holds the most recent data and is optimized for speed. This is where the last stretch of history lives at full resolution, because it is the data operators trend live, the data an alarm investigation reaches for, and the data most queries hit. It sits on the fastest storage available so those queries return immediately, and it keeps every sample so there is no loss of detail for the period people examine most closely. The cost per unit of storage here is the highest, which is acceptable precisely because the volume of recent data is comparatively small and its value is high.

The warm tier holds mid-age data that is accessed less often. As data ages out of the hot window, it is moved to storage that is cheaper and slower, and it is frequently transformed along the way, either compressed or reduced to aggregates such as periodic minimums, maximums, and averages rather than kept as every raw sample. The bet is that for data that is weeks or months old, most questions are about trends and summaries rather than the exact value at a specific instant, so some fidelity can be traded for a large reduction in footprint.

The cold tier is the deep archive for data that is rarely touched but must be retained, often for compliance or long-term analysis. It uses the cheapest storage, accepting that retrieval is slow and possibly involves a restore step, because the data may go untouched for long periods. Whether the cold tier keeps raw samples or only aggregates is a policy decision that depends on why the old data must be kept: a regulatory requirement to preserve raw values pushes toward archiving them cheaply, while a purely analytical need may be satisfied by aggregates alone.

How a Retention Policy Moves Data Between Tiers

The movement between tiers is driven by a retention policy, a set of rules that says how long data stays in each tier and what happens to it as it graduates to the next. A typical policy defines a hot window during which everything is kept at full resolution on fast storage, a warm window during which data is retained but often downsampled or compressed, and a cold window or indefinite archive for the oldest data. As data crosses each boundary with age, the policy transforms and relocates it automatically, so the tiers stay in their intended roles without manual intervention.

Because the policy governs both time and transformation, it encodes the organization's judgment about what matters. A policy that keeps full resolution longer favors fidelity at higher cost; one that aggregates aggressively favors cost at the expense of fine detail in older data. Policies can also differ by tag, so a critical measurement subject to regulatory scrutiny might be kept raw far longer than a routine utility point that can be summarized quickly. The policy is where those choices are made concrete and applied consistently.

The practical effect is a lifecycle that runs on its own once configured. New data lands hot, ages into warm where it is thinned, and eventually settles cold, and the historian serves each query from whichever tier holds the relevant data. A well-designed policy makes this invisible to users, who simply query a time range and get results, with only the speed and, for older ranges, the resolution reflecting which tier answered. The art is setting the windows and transformations so the boundaries fall where the access patterns actually change.

Query Speed, Fidelity, and Cost in a Cloud Historian

Tiering exists to manage a three-way trade-off. Keeping everything hot at full resolution forever would give the fastest queries and perfect fidelity but at a cost that grows without limit as data accumulates. Throwing old data away would be cheap but would lose history that may be needed. Tiering threads between these by matching the storage class to the data's actual value: fast and full-resolution where access is frequent and detail matters, cheap and summarized where it is not. Each tier is a deliberate point on the curve between speed, fidelity, and cost.

A cloud historian sharpens why this matters. In a platform like Merobix, data from many sites across oil and gas, water, power, and manufacturing accumulates continuously, so without tiering the volume of full-resolution data would grow relentlessly and every query would pay for storage no one needs to be fast. Tiering lets the recent operational data that operators actually watch stay instant, while years of history remain available at lower cost and, where appropriate, lower resolution. The platform serves a live trend and a multi-year lookback from different tiers, transparently, from the same browser.

The trade-off also shapes what users should expect. A query over the last day comes back instantly from the hot tier at full detail; a query spanning years may be slower and may return aggregates rather than every sample, because that older data lives in warm or cold storage. Understanding this helps set retention windows sensibly: the hot window should comfortably cover the period of routine investigation, the warm window should preserve enough resolution for the trending people do over months, and the cold tier should satisfy whatever long-term or compliance need justifies keeping the data at all.

Frequently Asked Questions

What do hot, warm, and cold mean for historian data?

Hot refers to recent data kept at full resolution on fast storage for instant queries, warm refers to mid-age data moved to cheaper storage and often compressed or reduced to aggregates, and cold refers to old data placed in a cheap archive that is slow to retrieve. The names describe how close the data is kept to fast access, which tracks how often it is used. Data moves from hot to cold as it ages.

Does tiering lose data resolution?

It can, by design, in the warmer and colder tiers. Recent data in the hot tier is kept at full resolution, but as data ages a retention policy may downsample it to aggregates such as minimums, maximums, and averages to save space, so very old data may no longer have every raw sample. Whether raw resolution is preserved in older tiers is a policy choice, often driven by compliance needs versus cost.

How do I decide the retention windows?

Set the hot window to comfortably cover the period over which people routinely investigate at full detail, the warm window to keep enough resolution for the months-long trending they do, and the cold tier to satisfy whatever long-term or regulatory need justifies keeping the data at all. Policies can also vary by tag, keeping critical or regulated points raw longer than routine ones. The goal is to align the tier boundaries with how access patterns actually change with age.

From Definitions to a Live Dashboard

Merobix reads your field devices into a cloud SCADA - the real thing behind these terms, live in days from any browser.

Request a Free Demo +1 (903) 307-7300
More in Automation Glossary
Historian Storage Sizing  •  Raw vs Aggregate Retention  •  PTP (IEEE 1588)  •  GPS-Disciplined Clock  •  Leap Second  •  UTC vs Local Time Storage  •  All Automation Glossary →
Free SCADA operator training
Merobix University - 70 video lessons & 261 quiz questions, from first login to compliance reporting. No demo call required.
Start free →