Time guide
- Initial checks: approximately 10 to 20 minutes
- Detailed investigation: approximately 20 to 60 minutes
Purpose#
Validate alerts, establish scope, classify severity and preserve investigation evidence.
Triage workflow#
- Record the detection time, source and identifier.
- Validate source health and timestamp accuracy.
- Establish the affected account, asset and scope using sanitised references.
- Review related identity, endpoint, network and change evidence.
- Classify the outcome and confidence.
- Escalate, contain or close according to authorised procedures.
- Record lessons learned and detection improvements.
Read-only checks#
Use the authorised security consoles and read-only queries available for the affected platform. Record every query and preserve timestamps.
Expected result#
The relevant services and resources should report healthy or ready states, expected dependencies should be present, and recent warning events should not show a persistent unresolved fault.
Troubleshooting#
If a check fails:
- Confirm names, namespaces, contexts and permissions.
- Review recent events and service logs.
- Compare the live state with the documented desired state.
- Check storage, DNS, network and dependent services.
- Preserve evidence before restarting, deleting or recreating anything.
Safety notes#
These entries begin with read-only checks. Treat restart, delete, restore, reconcile, exposure and security-containment actions as state-changing operations. Confirm scope, impact, authorisation and rollback before proceeding. Replace all private names, addresses, identifiers and credentials with placeholders before publication or sharing.
Related entries#
- Kubernetes Cluster Health Checks
- Flux Reconciliation Commands
- GitOps Repository Structure
- Security Incident Triage