Time guide
- Initial checks: approximately 10 to 20 minutes
- Detailed investigation: approximately 20 to 60 minutes
Purpose#
Classify informational, low, medium, high and critical alerts consistently.
Severity model#
- Informational: useful context with no immediate security impact.
- Low: limited risk or low-confidence suspicious activity.
- Medium: credible suspicious activity or a meaningful control failure.
- High: strong evidence of malicious activity, privilege abuse or material service risk.
- Critical: active or confirmed compromise with severe impact or urgent threat.
Read-only checks#
Use the authorised security consoles and read-only queries available for the affected platform. Record every query and preserve timestamps.
Expected result#
The relevant services and resources should report healthy or ready states, expected dependencies should be present, and recent warning events should not show a persistent unresolved fault.
Troubleshooting#
If a check fails:
- Confirm names, namespaces, contexts and permissions.
- Review recent events and service logs.
- Compare the live state with the documented desired state.
- Check storage, DNS, network and dependent services.
- Preserve evidence before restarting, deleting or recreating anything.
Safety notes#
These entries begin with read-only checks. Treat restart, delete, restore, reconcile, exposure and security-containment actions as state-changing operations. Confirm scope, impact, authorisation and rollback before proceeding. Replace all private names, addresses, identifiers and credentials with placeholders before publication or sharing.
Related entries#
- Kubernetes Cluster Health Checks
- Flux Reconciliation Commands
- GitOps Repository Structure
- Security Incident Triage