In this 4th video in the unstructured playlist, I will explain you how to extract metadata for better retrieval and also show you how to do better chunking.
80% of enterprise data exists in difficult-to-use formats like HTML, PDF, CSV, PNG, PPTX, and more. Unstructured effortlessly extracts and transforms complex data for use with every major vector database and LLM framework.
Link ⛓️💥
https://unstructured.io/
Code 👨🏻💻
https://github.com/sudarshan-koirala/...
------------------------------------------------------------------------------------------
Timestamps ⏰
00:00 Introduction
02:21 Semantic Search vs Hybrid Search
05:00 Setup
07:04 Find Elements associated with Chapters
10:00 Load documents into a ChromaDB
11:34 Hybrid Search with metadata
14:23 Chunking by Title
16:18 Conclusion
------------------------------------------------------------------------------------------
☕ Buy me a Coffee: https://ko-fi.com/datasciencebasics
✌️Patreon: / datasciencebasics
------------------------------------------------------------------------------------------
🤝 Connect with me:
📺 Youtube: / @datasciencebasics
👔 LinkedIn: / sudarshan-koirala
🐦 Twitter: / mesudarshan
🔉Medium: / sudarshan-koirala
💼 Consulting: https://topmate.io/sudarshan_koirala
#unstructureddata ##unstructuredio #metadata #chunking #llm #datasciencebasics