Monitoring & Observability
Reference guide
Read only the references needed for the current request:
- The Three Pillars — And How They Connect: references/the-three-pillars-and-how-they-connect.md
- Structured Logging That Actually Helps: references/structured-logging-that-actually-helps.md
- Prometheus: PromQL Deep Dive: references/prometheus-promql-deep-dive.md
- Grafana: Dashboard as Code: references/grafana-dashboard-as-code.md
- OpenTelemetry: Auto-Instrumentation: references/opentelemetry-auto-instrumentation.md
- Distributed Tracing: Practical Patterns: references/distributed-tracing-practical-patterns.md
- SLOs, SLIs, and Error Budgets: references/slos-slis-and-error-budgets.md
- On-Call and Incident Response: references/on-call-and-incident-response.md
- Severity: Critical: references/severity-critical.md
- Symptoms: references/symptoms.md
- First Response (< 5 minutes): references/first-response-5-minutes.md
- Diagnosis: references/diagnosis.md
- Mitigation: references/mitigation.md
- Escalation: references/escalation.md
- Timeline: references/timeline.md
- Root Cause: references/root-cause.md
- What Went Well: references/what-went-well.md
- What Went Wrong: references/what-went-wrong.md
- Action Items: references/action-items.md
- Lessons Learned: references/lessons-learned.md
- Datadog vs Self-Hosted: Decision Matrix: references/datadog-vs-self-hosted-decision-matrix.md
- Quick Reference: Essential Queries: references/quick-reference-essential-queries.md
- Checklist: Production Observability: references/checklist-production-observability.md