CloudWatch has been the default place AWS customers look for metrics, logs and alarms for over a decade. Amazon CloudWatch Omni, announced by AWS in September 2026, is the newest evolution of that service — an AI-powered observability workspace that watches applications and AI agents together, discovers your architecture automatically, and brings an AI agent into the investigation alongside your team. This guide covers what it is, how it works end to end, and where it fits next to the CloudWatch you already run.
Amazon CloudWatch Omni is an AI-powered observability experience, built on top of Amazon CloudWatch, for monitoring applications and AI agents — separately or together — in one workspace. Instead of living inside the AWS Management Console, Omni is reached through a dedicated sign-in URL for your organization, using the identity provider your team already manages.
Rather than replacing CloudWatch, Omni sits on top of it. Your existing metrics, logs, alarms, dashboards and Logs Insights queries keep working exactly as they do today. What Omni adds is a reimagined, AI-assisted way to explore that same telemetry: automatically discovered service topology, plain-language querying, and an AI agent that can join an investigation the moment something goes wrong.
If you want the fundamentals of the underlying signal model first, our guide on designing observability with metrics, logs and traces is a useful primer before diving into Omni specifically.
Two trends collided to make traditional CloudWatch feel incomplete for a growing number of teams:
CloudWatch Omni addresses both by discovering topology and adjusting alarms automatically instead of requiring manual upkeep, and by scoring agent output quality with evaluators that run continuously in production, not just checking whether a request completed.
The key difference from traditional CloudWatch is organizational as much as technical: CloudWatch is a set of console pages, APIs and data stores you query yourself; Omni is a standalone, AI-assisted workspace built on that same data, organized around your team's applications and agents rather than around individual metrics and log groups.
Omni is organized around three concepts that determine where your telemetry lives and who can see it:
End to end, the flow looks like this: your applications and agents send telemetry the same way they do today (the CloudWatch agent, AWS SDKs, or an OpenTelemetry pipeline) → that telemetry lands in a Dataset scoped to a Space → Omni's discovery layer builds a topology and computes golden metrics from it automatically → engineers work with that topology through the Omni web UI or an IDE extension → when something needs investigating, the AWS DevOps Agent joins the session and correlates signals across the whole graph.
Omni discovers your services and the dependencies between them directly from telemetry, with no manual tagging or configuration. As you deploy new services, the topology and its alarms update automatically — a meaningful shift from the usual pattern of dashboards quietly drifting out of sync with the architecture they're supposed to represent.
For every discovered service, Omni reports the RED metrics — Request rate, Error rate and Duration — as its golden metrics, computed automatically rather than requiring you to build them into a dashboard by hand.
You can ask a question about your applications in plain English, and Omni writes the underlying query for you — in SQL or PromQL — and answers from your own telemetry. This lowers the bar for exploring unfamiliar services: an engineer new to a system can ask "which service is driving the latency increase in checkout" without first learning that service's specific query syntax or dashboard layout.
The AWS DevOps Agent is enabled by default in every Omni investigation session. It correlates events across services, traces likely root-cause paths through the dependency graph, and suggests next steps alongside your team rather than replacing the team's judgment. Investigation history is captured automatically as a Thread, so when an incident escalates, the next responder joins with full context already in front of them instead of starting from a blank dashboard.
Because a Space's Dataset correlates logs, metrics and traces together, an investigation doesn't require switching between separate tools for each signal type — you can move from a metric anomaly to the trace that explains it to the log line that names the exact error, inside one workspace.
For agentic workloads, Omni records every step of an agent's execution as a structured, hierarchical trace — LLM calls, tool invocations and reasoning steps — so you can see exactly where behavior diverged from what you expected. Each span in that trace exposes inputs, outputs, token usage and latency, and a Compare mode lets you place two traces side by side to see how a different prompt, model or configuration changed the outcome.
Supporting tools built around these traces include a Session Explorer for reviewing full multi-turn conversation histories, and an Agent Topology view that visualizes an agent system's architecture — sub-agents, tools and how they connect.
Omni scores agent responses with evaluators — 17 built-in evaluators covering dimensions like coherence, helpfulness, faithfulness and routing correctness, plus support for bringing your own custom criteria scored by a judge model. Scoring runs both continuously against live production traffic and on demand against a curated dataset, and the scores sit alongside latency, errors and token usage on the same trace. This is what catches the failures that error rates miss entirely: a syntactically perfect, completely wrong answer.
A Playground lets you test different system prompts and model configurations side by side, Prompt Management versions and tracks those configurations over time, and an Experiments view runs the same golden dataset against two agent variants to compare their evaluation scores, latency and token usage before you ship a change.
Omni supports agents built with LangChain, LangGraph, CrewAI, the OpenAI Agents SDK, Strands and the Vercel AI SDK, in both Python and TypeScript, plus native observability for agents built on Amazon Bedrock AgentCore.
Omni is built on OpenTelemetry rather than a proprietary agent format. Telemetry you already send to CloudWatch appears in Omni with nothing to reconfigure, and any workload instrumented with OpenTelemetry — using OpenInference for AI-specific spans and the AWS Distro for OpenTelemetry (ADOT) — can send data straight to an OTLP endpoint. In practice, this means no re-instrumentation for existing CloudWatch or OpenTelemetry users, and it means Omni works the same way whether a workload runs on Lambda, ECS, EKS, or infrastructure outside AWS entirely.
A meaningful part of Omni's design is that agent debugging doesn't require leaving the editor. IDE extensions are available for VS Code, Cursor and Kiro, and the extension is free to use — an agent developer can start sending telemetry and inspecting traces without an AWS account at all. Kiro adds auto-instrumentation that detects the agent framework in use and configures tracing automatically, while a Cloud Login flow connects the local IDE session to an AWS account when a developer is ready to move from local traces to the full Omni workspace, with a playground and evaluators a click away from any trace.
Because a Space maps to exactly one account and region, that boundary — rather than a filter applied at query time — is what separates one team's telemetry from another's. An organization observes multiple accounts and regions by creating multiple Spaces inside the same Domain, and Omni maps topology and golden metrics across a single account or hundreds of them, across multiple regions, without manual configuration or tagging. Where you already centralize telemetry across accounts and regions using CloudWatch's own centralization rules, that centralized data appears in a Space like any other telemetry — Omni doesn't aggregate across accounts and regions on its own.
Omni also extends beyond AWS: connectors bring in telemetry from other environments, and AWS specifically highlights cross-cloud visibility that includes Azure workloads — useful for organizations running a genuinely mixed AWS/Azure estate that want one topology view instead of two separate monitoring stacks.
Every engineer reaches Omni through a single URL for the organization's Domain, authenticated through enterprise SSO via IAM Identity Center — including providers like Okta and Microsoft Entra ID — rather than through individual AWS Console credentials. A Space then groups the applications a specific team owns along with the telemetry associated with them, so access control follows team and application boundaries rather than raw account permissions. Combined with Threads for shared investigation history, this is what lets an investigation hand off cleanly between engineers, or between an on-call shift and the next one, without losing context.
At general availability (announced September 23, 2026), CloudWatch Omni is available in three AWS Regions: US East (N. Virginia), US West (Oregon) and Europe (Ireland). Getting started follows three steps:
Pricing is metered and detailed on the CloudWatch pricing page — worth modeling the same way you would evaluate any new telemetry or observability spend before rolling it out broadly.
For startups and small teams building on generative AI, the free IDE extension removes the usual observability bootstrap cost — there's no AWS account setup or dedicated observability engineer needed just to see what an agent is actually doing during development. For teams already on CloudWatch, Omni is an incremental step, not a migration.
For enterprises, the value leans more organizational: enterprise SSO and Identity Center integration, Spaces that map cleanly onto team and application ownership, multi-account and multi-region topology without a manual tagging project, and cross-cloud visibility for organizations that genuinely run mixed AWS/Azure estates. Adopting a new observability surface well — deciding what belongs in which Space, how alerts route into existing on-call tooling, and how this fits your current CI/CD and DevOps practice — is exactly the kind of rollout work a managed DevOps partner can help get right the first time.
CloudWatch Omni is a strong fit when a team already relies on CloudWatch and is expanding into agentic or generative AI features, when engineers are spread across multiple AWS accounts and regions and want one topology view instead of several, when incident response regularly involves handing an investigation off between people or teams, or when "did this change make the agent worse" needs a repeatable answer instead of manual spot-checking.
It matters less for a single small account with no AI agents and a already-mature, fully tuned Prometheus/Grafana or equivalent stack that the team has no appetite to change — in that case, the migration cost may outweigh what Omni adds on day one.
Consider a support-chat agent that starts drawing complaints after a prompt update. In Omni, the workflow looks roughly like this:
Amazon CloudWatch Omni is less a new monitoring tool than a rethink of how observability data gets used — automatically discovered instead of hand-maintained, correlated across applications and AI agents instead of split across separate stacks, and worked on with an AI agent in the loop rather than alone at 2 a.m. with a dozen open dashboard tabs. For teams already running AWS workloads or shipping their first agentic features, it's worth evaluating now rather than after the next incident makes the gap obvious.
Getting the rollout right — deciding Space boundaries, wiring alerts into existing on-call tooling, and folding evaluation-driven development into an existing CI/CD pipeline — is exactly the kind of cloud monitoring and managed DevOps work our team helps clients get right the first time.
No. Omni is built on CloudWatch and extends it — existing alarms, dashboards, APIs and console workflows continue to work unchanged.
Observability
A complete guide to Amazon CloudWatch Omni, AWS's AI-powered observability platform for applications and AI agents, built on OpenTelemetry.
More on this and related topics.
Modern systems need more than basic monitoring. Learn how metrics, logs and traces work together to provide useful operational visibility.
A practical guide to designing cloud infrastructure with the right balance of reliability, security, scalability and operational control.
How modern engineering teams can automate build, test, security and deployment workflows while keeping releases consistent and recoverable.