- Zen IT Technologies
- Technical notes
- You cannot investigate what you did not retain
You cannot investigate what you did not retain
Jonny Flaks, Founder & Principal Architect
Technical note in Security Readiness & Response
Almost every organization can say that it has logs. Far fewer can say how far back a particular question can still be answered, which is the only form of the statement that helps once something has gone wrong.
Default retention windows are set by the people who build the platform, and they are set around platform operation, storage cost and licensing. They were not designed around the questions an investigation eventually asks. Nobody chose them as an evidence policy, but that is what they become.
Detection date is not start date
An investigation arrives at the same question early: when did this actually begin?
The date something was noticed and the date it started are rarely the same, and the gap can be weeks. An account is found sending mail it should not be sending. That is today. The questions that follow are when the credential was first used by someone else, where it was used from, whether the same pattern appears on other accounts, and what was reached in between.
Those questions reach backwards. If the sources that could answer them have already aged the evidence out, the investigation does not fail loudly. It simply stops being able to say anything, and the report ends up describing the detection rather than the incident.
Part of the gap is structural. Detection tends to fire on a later stage of activity, because the earlier stages look ordinary: a successful sign-in, a mailbox rule, a file read by an account entitled to read it. The behavior that trips a rule is often several steps downstream of the behavior that started it, and those earlier steps are exactly the ones that were never interesting enough to keep for long.
The practical consequence is that the investigation horizon is not the retention setting of the platform that raised the alert. It is the shortest useful window across every source the question depends on.
Four clocks, not one horizon
It helps to stop thinking about retention as a single organizational number and start thinking about the separate clocks that actually exist:
- identity sign-in and administrative activity
- endpoint telemetry
- mail transport and message evidence
- application-side audit events
Each has its own window. Each is frequently tied to a license tier, a product configuration or a storage setting, and each can usually be extended at a price. What almost never exists is a document stating the four together as one joined investigation horizon.
The important point is not a universal number, because there is not one. It is that these clocks are set independently, often by different people at different times, and that an organization can hold a year of one and two weeks of another while describing both as logging.
The first useful piece of work is usually not extending retention. It is writing the four numbers down next to each other, at which point the shortest one tends to explain something that had previously been confusing.
Those numbers differ for reasons that have nothing to do with risk. The platforms were bought at different times by different people. One was migrated and inherited the defaults of the new tenant. One was tuned down after a storage bill. Retention drifts the way any setting drifts when no single person owns the whole picture.
The summary outlives the detail
Within a single source there is a second horizon, and it is easy to miss.
A detection record is small. The telemetry surrounding it is not. Process trees, command lines, parent and child relationships and detailed network activity are expensive to keep, so they are frequently kept for a shorter period than the alert that points at them.
That produces a particular kind of dead end. The high-level record confirms that something was detected on a given host on a given day, and the material that would establish what actually ran has already gone. You can prove the event happened without being able to describe it.
A detection that was contained and closed needs little history. The one that turns out to be the first visible symptom of something older is the one where the surrounding context has usually expired.
The response is not to keep everything for a year. It is to decide, in advance, which systems and which classes of activity justify deeper telemetry for longer, and to accept a shorter window everywhere else knowingly rather than by omission.
Design retention by question, not by product
The more durable approach is to start from the questions that must remain answerable and work back to the sources that answer them, rather than starting from the products and accepting whatever window each one arrived with.
Once the question is written down, the trade-off becomes explicit. Where native retention is shorter than the question requires, the options are deliberate export to secondary storage or an explicit acceptance of the horizon. Both are legitimate. Only one of them is usually a decision anybody made.
Retention is also, frequently, a function of the commercial arrangement rather than a security setting. A tier determines how long telemetry is held. A plan includes a longer audit window. A consolidation moves a workload onto a subscription with different defaults. Vendors differ and the details change, so this is a per-platform check rather than an assumption.
The pattern is procedural. A procurement decision or a platform consolidation can shorten the investigation horizon without anybody experiencing it as a change to a security control, and it will not appear in a change record as one. It runs the other way too: a tier bought for an unrelated feature can lengthen an audit window that nobody then thinks to use. Retention belongs on the list of things checked whenever a platform's commercial arrangement changes.
The questions themselves tend to be stable across organizations even when the platforms are not. Who had access to this, and when did that access change. What did this device do. Whether this message arrived, and what happened to it afterwards. Who altered a control. A retention design built around those four questions survives a platform migration in a way that a design built around a particular product's settings page does not.
The question also exposes redundancy. The same fact often exists in more than one place: a message may be evidenced at the gateway and in the mailbox, a sign-in may appear in the identity platform and in the application that was reached. Where a cheap source holds a fact for longer than the expensive one, that is worth knowing before an incident rather than during it.
The ninety-day test
Pick an ordinary date roughly ninety days ago. Not an incident, and not a date anyone has a reason to remember. Then try to answer four questions about it.
- Who signed in to a particular account, and from where?
- What did a particular endpoint execute?
- Was a particular message delivered?
- Who changed a significant setting in a business application?
Whatever can no longer be answered is not a gap in the plan. It is the real investigation horizon, and it is the one that will apply on the day it matters.
Ninety days is the test date rather than a recommendation. The right window depends on what the organization needs to be able to reconstruct, and for some questions it is far shorter or far longer. The value of the exercise is that it converts an assumption into a measurement, at a time when the answer is merely useful rather than urgent.
Explore this expertise: Security Readiness & Response