#01 · web data extraction, no LLM — live log

Опубликовано: 30 Сентябрь 2026
на канале: Olssigo
32
1

Build-in-public log #01.

Structured data extraction from real web pages — record detection + field inference running live, no selectors, no per-site rules, no LLM. Multiple site types (shop / board / forum / etc). Unedited, real time, before any manual correction.

A raw progress log, not a demo reel — I'm stacking these as the model improves.

Useful as a front-end for RAG / data pipelines — feeding listing, catalog, and monitoring data into downstream systems.

Why no LLM: the model classifies the DOM nodes already on the page instead of generating text — so it can't invent a value that isn't there, and there's no per-page token cost. LLMs are more flexible; this trades that for predictability.

Not yet: article/blog body text and detail-page (link-through) extraction — repeating record structures only. Shown honestly.

#webscraping #dataextraction #structureddata #webdata #rag #dataengineering #datapipeline #webautomation #machinelearning #graphneuralnetwork #noLLM