Memory hook: Raw evidence, normalized fields, useful context.
Must remember
- Prioritize logs by detection use cases and asset risk: identity, administrative activity, endpoint, network, application and cloud findings. Estimate volume and retention before enabling every verbose source.
- Ingestion transports events; parsing extracts meaning; normalization maps it to the Unified Data Model (UDM). A successfully delivered raw log can still be unusable for a rule if fields are wrong.
- Validate timestamps, event types, principal/target fields, network addresses and identifiers. Time-zone errors can break correlation; parser changes need representative regression samples.
- Use supported parser extensions or custom parsing when necessary, keeping the raw evidence available for investigation. Monitor rejected events, ingestion delay and schema changes.
- Entity data describes users/assets and enriches event data about activity. Aliasing fields link identities across sources; poor matching can merge unrelated users or split one attacker into many entities.
- Enrich with asset criticality, ownership, identity context and relevant intelligence. Labels and risk context should have provenance and freshness, not become permanent unverified truth.
Review details
UDM separates event metadata and normalized nouns/actors. Principal is the initiating actor and target the recipient/object in the event's context; observer/intermediary fields serve different roles. Mapping the victim into the actor field can invert a detection. Validate raw message → parsed timestamp/type → normalized entities → rule field references with known samples.
Entity context enriches activity with ownership, criticality and identity relationships. Aliasing can link a hostname, account and other identifiers, but NAT, reused IPs and shared accounts make careless matching dangerous. Parser extensions should be tested on expected variations and rejected cases; successful transport does not imply the rule can use the event.
Choose under exam pressure
| Requirement | Choice and reason |
|---|---|
| Rule sees no username although raw logs contain one | Inspect parser mapping and UDM principal/target fields. |
| High ingestion cost with little detection value | Review source priority, filtering and retention against concrete use cases. |
Traps
- Raw log arrival is not proof of correct normalization.
- An IP address alone may not uniquely identify a user or device over time.
Active recall
1. What is UDM for?
A common event/entity representation that enables consistent search and detection across sources.
2. Why preserve raw events?
To validate parsing and retain original evidence for investigation.
3. What can a timezone error cause?
Missed or false correlations across events that actually occurred together.
4. How test a parser change?
Replay representative samples and verify required fields and downstream detections.
5. Why track entity aliases?
To connect the same asset/user across identifiers without incorrect merging.