In this recorded session, we host Yixin Hu (VU Amsterdam) and Thomas Hulard (McDermott), who showcase the implementation of an evaluation framework for RAG within an engineering services company.
They present the challenges of evaluating RAG and share the solution that was selected, along with its results.
Resources
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
https://arxiv.org/abs/2306.05685
ARAGOG: Advanced RAG Output Grading
https://arxiv.org/abs/2404.01037v1
LlamaIndex documentation:
Answer Relevancy and Context Relevancy Evaluations
https://docs.llamaindex.ai/en/stable/...
Correctness Evaluator
https://docs.llamaindex.ai/en/stable/...
Faithfulness Evaluator
https://docs.llamaindex.ai/en/stable/...