Ship Responsibly: Evaluation and Safety in Azure AI Foundry
You have built an AI-powered application - now how do you know it actually works? Traditional software has unit tests and integration tests, but AI applications are probabilistic, context-dependent, and capable of producing harmful content. In this session, we tackle the hardest problem in AI engineering: measuring quality and ensuring safety before you ship. Using Azure AI Foundry’s built-in evaluation framework, we will run quality evaluators (groundedness, relevance, coherence, fluency) and safety evaluators (violence, hate, self-harm, protected material, jailbreak detection) against a real application. You will see how to build test datasets, run evaluations locally and in the cloud, interpret scorecards, and integrate evaluation gates into your CI/CD pipeline so that a regression in quality or safety blocks deployment automatically. We will also cover Foundry’s runtime content filters, the Risks and Safety monitoring dashboard, and how to build a responsible AI practice that goes beyond tooling - including human oversight, incident response, and transparency. Whether you are a developer, tech lead, or architect, you will leave with a practical playbook for shipping AI applications that your users can trust.
This session is rated level 200 -"Basic Concepts".
Reference Links
Interested in this Talk?
Would you like this talk at your event? You can send me an email. If you use Sessionize you can view the talk on Sessionize.Share on
Bluesky Facebook LinkedIn Reddit XLike what you read?
Please consider sponsoring this blog.