Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Agentic AI Evaluation
Explore practical methods for measuring agentic AI performance, safety, and transparency using multi‑dimensional dashboards, emphasizing real‑world applicability and continuous governance.
Agentic AI Evaluation
Agentic AI refers to systems that can plan, decide, and act autonomously across multiple steps.
Evaluating such systems is harder than testing traditional AI because their behavior changes with context.
Standard benchmarks often fail to capture real-world complexity and tool use.
New methods combine task success rates, reasoning quality, and adaptability measures.
Human-centric factors like safety, transparency, and ethics are equally important.
Industry tools now provide multi-dimensional evaluation dashboards for agents.
Governance and monitoring are critical for safe deployment.
This talk will present key evaluation dimensions and emerging best practices.
Attendees will learn how to balance performance metrics with trust and accountability.
The goal is to make agentic AI both effective and responsible in real-world use.
Compose Email
Loading recent emails...