Evaluations are the single most reliable indicator of the health and long term viability of any gen AI project. As a Principal Applied AI Architect for AWS, I've had the opportunity to look at over 100 different attempts at evaluation frameworks over the last few years.
In this talk I share some stories about the best and worst, and then distill the 7 most common elements I've seen in successful evaluations.
Slides at https://d2ot4ns4zf41bm.cloudfront.net...