Once your application is running in the cloud, you need to know whether it is healthy. Is it responding quickly? Are errors rising? Is a server running out of memory? Monitoring answers these questions by collecting and watching data about your systems. Observability goes a step further: it gives you enough data to work out why something is going wrong, even a problem you didn't predict.
You'll learn the three main kinds of data, metrics, logs and traces, and how they work together. Then you'll use popular tools: Prometheus and Grafana to collect metrics and build dashboards, Amazon CloudWatch for monitoring on AWS, and OpenTelemetry for gathering telemetry in a standard way. You'll also see how alerts tell you about a problem before your users do. Watch the videos in order, then move on to Cost & Reliability.
A short introduction to how metrics, logs and traces help you understand what your systems are doing.
Monitoring versus observability, logs, metrics, traces, OpenTelemetry and SLOs, with an incident walkthrough.
Collect metrics with Prometheus and turn them into dashboards with Grafana.
Metrics, alarms, dashboards and logs in the monitoring service built into AWS.
What OpenTelemetry is and how to use it to get observability across a distributed system.