Metadata Extraction & Chunking Using Unstructured | ChromaDB

Опубликовано: 09 Февраль 2026
на канале: Data Science Basics
6,373
99

In this 4th video in the unstructured playlist, I will explain you how to extract metadata for better retrieval and also show you how to do better chunking.

80% of enterprise data exists in difficult-to-use formats like HTML, PDF, CSV, PNG, PPTX, and more. Unstructured effortlessly extracts and transforms complex data for use with every major vector database and LLM framework.

Link ⛓️‍💥
https://unstructured.io/

Code 👨🏻‍💻
https://github.com/sudarshan-koirala/...

------------------------------------------------------------------------------------------
Timestamps ⏰
00:00 Introduction
02:21 Semantic Search vs Hybrid Search
05:00 Setup
07:04 Find Elements associated with Chapters
10:00 Load documents into a ChromaDB
11:34 Hybrid Search with metadata
14:23 Chunking by Title
16:18 Conclusion

------------------------------------------------------------------------------------------

☕ Buy me a Coffee: https://ko-fi.com/datasciencebasics
✌️Patreon:   / datasciencebasics  

------------------------------------------------------------------------------------------
🤝 Connect with me:
📺 Youtube:    / @datasciencebasics  
👔 LinkedIn:   / sudarshan-koirala  
🐦 Twitter:   / mesudarshan  
🔉Medium:   / sudarshan-koirala  
💼 Consulting: https://topmate.io/sudarshan_koirala

#unstructureddata ##unstructuredio #metadata #chunking #llm #datasciencebasics