Retrieval moves inside a loop the model steers: it searches, reads what came back, says what is still missing, and picks the next query itself. That one change — not a bigger retrieval budget — is the whole difference between search and deep search.
If you have shipped retrieval-augmented generation, the dial you would reach for is top-k. It is the one thing here that does not matter. RAG retrieves once against a fixed k and then answers; a deep search agent decides query N+1 after reading result N, and decides for itself when to stop. Anthropic's own sub-agent prompt has to spell out "start broad, then narrow", because agents that open with a specific query find nothing and thrash.
The finished agent is 158 lines. This episode writes the first 108 of them, and there are only four things in the file: one Python list of messages, which is the context window and not a database; a search tool that is a name, a sentence of description and a schema holding one string; four sentences of system prompt that carry the whole broad-then-narrow strategy; and an exit that is a single line — if the reply is not a tool call, break. The turn limit above it is a safety net, not a decision. Nothing in the file forces the agent to stop searching, and nothing forces it to keep going either.
Everything here has a cost. Anthropic measured a single research agent at roughly four times the tokens of an ordinary chat turn, before any of the machinery the later episodes add. And the version that works is still broken: the entry point passes the whole question in as one string, so where you see three parts — the Mac mini, the card, and what "better" is even being measured by — the code sees one question and one message list. In the run recorded in the repo, on a small local model, it issued fourteen searches and every one was the same comparison reworded. One loop cannot hold the whole question; it answers all of it, thinly. Episode 2 splits it.
Chapters
0:00 Intro — The whole agent is 158 lines
0:11 The question we keep asking
0:32 The dial you'd reach for
0:52 The whole map
1:13 Retrieval moves inside the loop
1:34 Region one, in Python
1:47 One list, three kinds of entry
2:06 The tool is a description
2:28 Broad, then narrow — and it's a prompt
2:47 The exit is one line
3:10 What one loop costs
3:30 It's finished, and it runs
3:42 What you read, what the code reads
4:01 Get it and run it
4:16 Read the trace
4:37 One loop, one list
The code
https://github.com/ofrik/load-bearing — this series lives under deep-search-agent/, one directory per episode. Episode 1 is ep01/deep_search.py, the 108 lines built here. It is composed for this series against the documented Anthropic Messages API surface; it is not an excerpt from anyone's production system.
Sources
How we built our multi-agent research system (Anthropic): https://www.anthropic.com/engineering...
Building multi-agent systems: when and how to use them (Anthropic): https://claude.com/blog/building-mult...
Tool use overview (Anthropic docs): https://platform.claude.com/docs/en/a...
Open Deep Research (LangChain): https://www.langchain.com/blog/open-d...
open_deep_research, the implementation (LangChain, GitHub): https://github.com/langchain-ai/open_...
Load Bearing — how it works, and what it costs you. Mechanism-first breakdowns for developers who can read code but are new to the specific topic. Every number comes with the condition it was measured under.
#DeepResearch #AIAgents #LoadBearing