Module 13 of 13

Observability

Lessons

About This Module

Everything up to this point helps you build and ship software — this final module covers how you actually know it's working once it's live. It starts with observability fundamentals: the three pillars of logs, metrics, and traces, and why they answer different questions when something goes wrong. From there it covers centralized logging with the ELK/EFK stack and collecting metrics with a time-series database like Prometheus.

It also covers building dashboards with tools like Grafana, following a request across services with distributed tracing, setting up alerting that wakes the right person at the right time, defining SLIs, SLOs, and error budgets, and running effective incident response and postmortems once something does break.

This is the final module in the DevOps course — once you've watched these lessons, you'll have covered the full path from a Linux terminal to a fully observable, automated production system.

Lessons

8 videos
01

Observability Fundamentals: Logs, Metrics & Traces

The three pillars of observability, what question each one is best at answering, and how they work together when debugging an incident.

02

Centralized Logging with the ELK/EFK Stack

Shipping logs from every service to a single searchable place instead of SSHing into individual machines to grep files.

03

Metrics & Time-Series Databases: Prometheus

Collecting and storing numeric metrics over time with Prometheus, and querying them with PromQL to spot trends and anomalies.

04

Visualization & Dashboards: Grafana

Turning raw metrics into dashboards that give a team a shared, at-a-glance view of system health.

05

Distributed Tracing

Following a single request as it moves across multiple services, and using traces to pinpoint exactly where latency is coming from.

06

Alerting & On-Call Practices

Writing alerts that fire on symptoms that matter, avoiding alert fatigue, and structuring an on-call rotation that's sustainable.

07

SLIs, SLOs, SLAs & Error Budgets

Defining what "reliable enough" means for a service, and using an error budget to balance shipping new features against stability.

08

Incident Response & Postmortems

Running a clear-headed incident response, and writing blameless postmortems that turn an outage into lasting process improvements.