Wait... What REALLY Is A Vector Database?

Опубликовано: 02 Февраль 2026
на канале: Adam Lucek
6,082
258

Vector databases and similarity based search/retrieval have had a massive increase in popularity with the rise of language models and retrieval augmented generation (RAG) pipelines. Many people are using VDBs like pinecone, chroma, and mongodb, but how many of us actually know how they works? In this video we cover how vector databases operate, diving into what an embedding is, how similarity is calculated, indexing and searching algorithms, and an example of a full RAG flow.

Code: https://github.com/ALucek/embeddings-...

Resources:
Vector Embeddings for Developers: https://www.pinecone.io/learn/vector-...
What are Embeddings Blog: https://platform.openai.com/docs/guid...
Vector Database Blog: https://www.pinecone.io/learn/vector-...

Chapters:
00:00 - Introduction
01:19 - What is an Embedding?
02:56 - What is a Vector?
05:24 - Embeddings: One Hot Encoding
07:27 - Embeddings: Co-occurence Matrix
09:43 - Embeddings: Modern Embedding Models
12:47 - Embeddings: Why + Setup
14:17 - Visualization: 1 Dimension
16:28 - Visualization: 2 Dimensions
18:31 - Visualization: 3 Dimensions
20:32 - Calculating Similarity: Why it’s Important
21:28 - Calculating Similarity: Euclidean Distance
24:02 - Calculating Similarity: Dot Product
25:56 - Calculating Similarity: Cosine Similarity
29:18 - Semantic Retrieval Overview
31:53 - Algorithms: Introduction
34:09 - Algorithms: K-Dimensional Trees
37:48 - Algorithms: Locality Sensitive Hashing
39:39 - Algorithms: Product Quantization
44:30 - Algorithms: Inverted File Index
46:19 - Algorithms: Hierarchical Navigable Small World
49:23 - RAG: Process Overview
51:33 - RAG: Chunking & Splitting
52:45 - RAG: ChromaDB Setup
54:20 - RAG: LLM Setup
55:00 - RAG: Full Pipeline
56:31 - Additional Embedding Use Cases
59:02 - Closing Thoughts

#ai #programming #data