Skip to main content
  1. Runbook/
  2. Security/

Security Incident Triage

Time guide

  • Initial checks: approximately 10 to 20 minutes
  • Detailed investigation: approximately 20 to 60 minutes

Purpose
#

Validate alerts, establish scope, classify severity and preserve investigation evidence.

Triage workflow
#

  1. Record the detection time, source and identifier.
  2. Validate source health and timestamp accuracy.
  3. Establish the affected account, asset and scope using sanitised references.
  4. Review related identity, endpoint, network and change evidence.
  5. Classify the outcome and confidence.
  6. Escalate, contain or close according to authorised procedures.
  7. Record lessons learned and detection improvements.

Read-only checks
#

Use the authorised security consoles and read-only queries available for the affected platform. Record every query and preserve timestamps.

Expected result
#

The relevant services and resources should report healthy or ready states, expected dependencies should be present, and recent warning events should not show a persistent unresolved fault.

Troubleshooting
#

If a check fails:

  • Confirm names, namespaces, contexts and permissions.
  • Review recent events and service logs.
  • Compare the live state with the documented desired state.
  • Check storage, DNS, network and dependent services.
  • Preserve evidence before restarting, deleting or recreating anything.

Safety notes
#

These entries begin with read-only checks. Treat restart, delete, restore, reconcile, exposure and security-containment actions as state-changing operations. Confirm scope, impact, authorisation and rollback before proceeding. Replace all private names, addresses, identifiers and credentials with placeholders before publication or sharing.

Related entries#

  • Kubernetes Cluster Health Checks
  • Flux Reconciliation Commands
  • GitOps Repository Structure
  • Security Incident Triage