Jan Čurn - How to feed LLMs with data from the web | WebExpo 2024

Опубликовано: 27 Апрель 2026
на канале: Apify
931
29

All major generative AI models have been trained using data scraped from the web. Applications of large language models (LLMs) often extract web data to provide up-to-date context using Retrieval Augmented Generation (RAG). Unfortunately, reliably collecting online data at scale is challenging due to issues like blocking, dynamic content rendering, and the sheer volume of data. In this talk, Jan will explain how you can establish an efficient web data extraction pipeline, clean the HTML to circumvent the “garbage in, garbage out” problem, and demonstrate how to use this in an LLM application. The demo uses Apify's Website Content Crawler https://apify.com/apify/website-conte... - a specialized crawler built for the LLM and RAG use cases.

This talk was presented at the WebExpo Conference in Prague on May 30, 2024 🎤

Big thanks to the WebExpo team for allowing us to publish this recording 🤝🏻

📲 Follow Jan:   / jancurn     / jancurn  
🌍 Get your ticket for the next WebExpo: https://webexpo.net/
🔍 Watch our webinar on feeding your LLMs with web data:    • Web Scraping Data for Generative AI - Lear...  

More AI-related resources from Apify 🧑‍💻
🧠 Explore tools we offer to ingest entire websites and feed data for AI/LLM: https://apify.it/3zPzrYT
🛍️ Browse AI scrapers and automation tools: https://apify.it/3LDM3oQ
🤩 Learn more about Apify: https://apify.it/46hlxLr

Follow us 🤳
  / apifytech  
  / apify  
  / apifytech  
  / discord  

#webscraping #webexpo #llm #ai