Module 14 of 15

Monitoring & Observability

Lessons

About This Module

Once your application is running in the cloud, you need to know whether it is healthy. Is it responding quickly? Are errors rising? Is a server running out of memory? Monitoring answers these questions by collecting and watching data about your systems. Observability goes a step further: it gives you enough data to work out why something is going wrong, even a problem you didn't predict.

You'll learn the three main kinds of data, metrics, logs and traces, and how they work together. Then you'll use popular tools: Prometheus and Grafana to collect metrics and build dashboards, Amazon CloudWatch for monitoring on AWS, and OpenTelemetry for gathering telemetry in a standard way. You'll also see how alerts tell you about a problem before your users do. Watch the videos in order, then move on to Cost & Reliability.

Lessons

5 videos
01

Observability Explained

A short introduction to how metrics, logs and traces help you understand what your systems are doing.

02

Observability Crash Course

Monitoring versus observability, logs, metrics, traces, OpenTelemetry and SLOs, with an incident walkthrough.

03

Prometheus and Grafana

Collect metrics with Prometheus and turn them into dashboards with Grafana.

04

Monitoring AWS with Amazon CloudWatch

Metrics, alarms, dashboards and logs in the monitoring service built into AWS.

05

Getting Started with OpenTelemetry

What OpenTelemetry is and how to use it to get observability across a distributed system.