Talk · AWS Community Day Adria 2026

Trusting the Unpredictable: Observability, Evaluation and Explainability for Agentic AI on AWS

Time and room follow soon

Abstract

Your agent demo worked perfectly. Then you shipped it and it started making decisions you can’t explain, calling tools you didn’t expect, and failing in ways no test caught. Welcome to production!

The gap between an impressive prototype and a system you’d actually put in front of users comes down to one question: how do you know when an agent deserves your trust?

This session moves past building agents to the harder problem of operating them. We’ll trace the shift from MLOps and LLMOps to AgentOps, and why accuracy alone tells you almost nothing about whether an autonomous system is behaving well in production.

Using Amazon Bedrock AgentCore as the backbone, we’ll work through three pillars of trustworthy agents: observability, explainability, and evaluation, grounded in real traces from an agent that misbehaves. You’ll see how a single failed run looks through each lens: the execution trace that exposes a bad tool call, the reasoning path that explains why the agent chose it, and the evaluation score that catches a quality regression before users do. Along the way we’ll show where AgentCore Observability, AgentCore Evaluations, Bedrock Guardrails, and CloudWatch GenAI Observability fit, and where the work is still yours.

Whether you’re an architect, developer, or AI practitioner, you’ll leave with a practical framework, and concrete examples, for designing agents that are not just intelligent, but observable, explainable, and trustworthy.

Contact Us

Credits

This website uses the open source AWS Community Day Template built by AWSug.nl hosted on Amazon CloudFront and Amazon S3. The website uses bootstrap and hugo.