Anthropic engineers on building evaluation harnesses, testing...

• 02:18 - Moving from text output to agent outcome evaluation
• 05:43 - Architecture and mechanics of an Eval Harness
• 10:48 - Converting production failures into test tasks
• 14:52 - Code-based checks vs LLM-as-a-judge calibration
• 17:09 - Structuring Regression & Capability test suites
• 24:32 - Q&A with Anthropic Applied AI engineers
in this 45-minute deep dive, Anthropic Applied AI engineers alongside Notion's PM break down how to systematically test autonomous systems from scratch.
Agent Harness + Regression Suites + End State Verification = Reliable Agents
Watch the full session today, then build your own evaluation suite.

• 00:00 - The hype vs reality of AI coding agents
• 02:09 - Why the "Product Management Bottleneck" is getting worse
• 03:26 - The rise of high context, generalist micro teams
• 06:15 - Using coding agents to rapidly assemble building blocks
• 13:40 - Why bottom up AI innovation isn't enough for the enterprise
• 27:01 - The next $100M problem: Unstructured data architecture for agents
• 29:28 - Why NoSQL is speeding up agentic iteration
This 30-minute discussion will save you $1000 in paid courses on agentic workflow strategy.
Watch it today, then level up your skills