If you thought that RAG was only possible with text based documents like boring pdfs, think again! In this video we go over all you need to know about setting up a RAG based application that allows querying and context over photos! We take advantage of lesser known CLIP models and new ChromaDB integrations to walk through an end to end example of setting up your own multimodal retrieval augmented generation flow.
Code Available Here: https://github.com/ALucek/multimodal-rag
Resources:
ChromaDB: https://www.trychroma.com/
CLIP OpenAI Blog Post: https://openai.com/index/clip/
OpenCLIP: https://github.com/mlfoundations/open...
Fashionpedia Dataset: https://huggingface.co/datasets/detec...
LangChain multimodal prompts: https://python.langchain.com/v0.2/doc...
LAION: https://laion.ai/blog/laion-5b/
Chapters:
00:00 - Intro & Overview
00:54 - Context: What are CLIP Models?
03:21 - MultiModal RAG Diagram
03:40 - Choosing a Dataset: Fashionpedia
05:14 - Loading & Preparing the Dataset
06:55 - Context: OpenCLIP Embedding Models
08:09 - Setting Up Your VectorDB
09:01 - Adding Images to Your VectorDB
10:06 - Test: Querying the VectorDB
12:23 - Context: Choosing a Vision Model
13:39 - Setting Up RAG Prompt, Model, & Chain
16:01 - Formatting Pictures & Prompt for Vision Models
18:47 - Putting Everything Together
19:48 - Executing the Multimodal RAG Flow!
21:34 - Conclusion & Closing Thoughts
#multimodal #artificialintelligence #RAG