0:00 / 0:00
Loading...

AWS re:Invent 2025 - From Metrics to Management: Practical Observability on EKS (DEV202)

Watch Topic Details
Introduction to Observability
  • Observability is about extracting actionable insights to assess and improve application performance, health, and behavior.
  • It should be a continuous cycle involving detection, investigation, remediation, and assessment.
Importance of Observability
  • Enables gaining insight into application health.
  • Facilitates better troubleshooting of issues.
  • Enhances customer experience.
  • Helps control costs.
Observability Signals
  • Metrics: Identifying problems.
  • Traces: Locating problems.
  • Logs: Understanding the cause of problems.
  • Profiles: Analyzing code behavior.
Amazon Managed Service for Prometheus
  • Fully managed, serverless, Prometheus-compatible monitoring service.
  • Provides high availability and multi-AZ support.
  • Uses Prometheus data model and PromQL for querying.
  • Can be integrated with OpenTelemetry for metrics collection.
Best Practices for Amazon Managed Service for Prometheus
  • Use private link and IAM for secure data transfer.
  • Reduce costs by adjusting scrape interval and relabel config.
  • Set appropriate retention periods for data.
  • Ensure high availability by running multiple containers.
  • Use consistent identifiers for data correlation.
Amazon Managed Grafana
  • Fully managed analytics and visualization platform.
  • Used for creating dashboards, alerting, and managing metrics and traces.
  • Works in conjunction with Amazon Managed Service for Prometheus.
  • Best practices include creating clear, specific dashboards and setting appropriate retention periods.
Demo of Observability Setup
  • Demonstration of setting up observability using Amazon Managed Service for Prometheus and Amazon Managed Grafana.
  • Involves creating scrapers, workspaces, and configuring data sources.
  • Emphasis on following best practices for configuration and visualization.

Description

This talk shows how to bake observability into Amazon EKS so you can see exactly how each deployment impacts your SLOs. I’ll cover setting up Prometheus and Grafana, cutting noise with smarter instrumentation, and turning metrics into release decisions. You’ll leave with configs and patterns you can drop straight into your own clusters using best practice.

Learn more:
AWS re:Invent: https://go.aws/reinforce.
More AWS events: https://go.aws/3kss9CP

Subscribe:
More AWS videos: http://bit.ly/2O3zS75
More AWS events videos: http://bit.ly/316g9t4

ABOUT AWS:
Amazon Web Services (AWS) hosts events, both online and in-person, bringing the cloud computing community together to connect, collaborate, and learn from AWS experts.
AWS is the world's most comprehensive and broadly adopted cloud platform, offering over 200 fully featured services from data centers globally. Millions of customers—including the fastest-growing startups, largest enterprises, and leading government agencies—are using AWS to lower costs, become more agile, and innovate faster.

#AWSreInvent #AWSreInvent2025 #AWS

s Spotlight
or a Approve
i Improve
or r Reject