What Happens When You Press Enter on ChatGPT

Опубликовано: 28 Сентябрь 2026
на канале: Beneath the Prompt
660
9

You press Enter on ChatGPT. Under a second later, the first word appears. Here is what your prompt actually did in that gap — every cable, every chip, every step.

It crossed fiber at the speed of light, entered a building that draws more power than a small town, became thousands of numbers per word, and was processed by a chip so hot that coolant is pumped onto its lid. Seventeen minutes, one continuous path, no gaps left as magic. If you use this thing every day and still feel one layer removed from it, this is the layer.

═══════════════════════════════
📌 CHAPTERS
═══════════════════════════════
0:00 Under one second
0:32 The keystroke — 8 bytes leave your keyboard
0:53 DNS and TLS — finding the server, encrypting the request
1:24 Crossing the internet — packets, routers, fiber at light speed
2:01 The physical scale — CPU vs GPU, 700 W per chip, liquid cooling
3:08 Two separate networks — Ethernet vs the accelerator fabric
3:58 You're not alone — continuous batching and the safety filter
5:36 Tokenization — text becomes numbers
6:04 Embedding — one word becomes thousands of numbers
6:45 The Transformer — mixture of experts, a reported ~1.8 T parameters
7:58 Attention — every token looks at every other token
9:35 The hardware reality — FlashAttention and FP8
10:43 Eight accelerator devices — all-reduce and NVLink
12:07 Prefill — the whole prompt in one pass, and the KV cache
12:37 Decode — one token at a time, and why the chip sits idle
14:24 The KV cache problem — why long chats slow down
15:03 Streaming back — server-sent events and the typewriter effect
15:45 The full picture — 3 to 4 seconds, start to finish
16:19 What happens every time you press Enter
16:46 Corrections welcome

═══════════════════════════════
📚 WHAT YOU'LL LEARN
═══════════════════════════════
• The physical route of a prompt: keyboard → browser → DNS/TLS → 10–15 routers → Cloudflare → a private lane into the data center
• Why an H100 draws 700 W, a server 10 kW, a rack 40–100 kW — and why the coolant goes straight to the chip
• Continuous batching: how one accelerator device serves 64 prompts at once
• Mixture of experts: why, by the widely reported estimate, only ~280 B of GPT-4's ~1.8 T parameters wake up for your token
• FlashAttention: changing where the math happens, not the math
• Prefill vs decode — and why the chip sits over 99 % idle while it answers you
• The KV cache: why long conversations slow down, and why context has a limit
• The whole timeline: ~30 ms network · 50–200 ms queue · 80–300 ms prefill · 2–3 s decode

═══════════════════════════════
⚠️ ESTIMATES, AND THREE CORRECTIONS
═══════════════════════════════
The GPT-4 architecture figures (≈1.8 T parameters, ≈280 B active per token, 16 experts, ≈120 layers) come from a July 2023 SemiAnalysis report working from unnamed sources. OpenAI has never confirmed them, and the GPT-4 technical report withholds architecture and model size on purpose.

Before launch I re-checked every number in the film against the sources. Three didn't hold up: 12,288 ("twelve thousand numbers per token") is GPT-3's published width, not GPT-4's; the memory latency at 9:39 is hundreds of nanoseconds, not microseconds; and published output speeds run roughly 30 to 135 tokens per second, not 80 to 100, and not per chip. All three are corrected in full in the pinned comment. The hardware specifications — 700 W, 3.35 TB/s, NVLink, TLS, BPE, FlashAttention — come from official documentation and peer-reviewed papers.

═══════════════════════════════
⚠️ A NOTE ON THE H100
═══════════════════════════════
The H100 is a Hopper chip from 2022, and NVIDIA has moved through Blackwell to Rubin since. It is the one used here because it is the best-documented accelerator in public — the power draw and the memory bandwidth are on the manufacturer's own spec sheet. The argument does not rest on the part number: hundreds of watts per chip, coolant run straight to the die, and memory bandwidth rather than arithmetic as the limit have each gotten more extreme since, not less.

If you spot an error, comment. Corrections get pinned and the description gets updated.

The photographic shots — the fiber, the data-center exterior and interior, the chip under coolant — are AI-generated images, as are the music and the narration voice. Every diagram and animation was built from scratch for this video, and the script is original and human-checked against the sources below. No real company's facility is shown.

═══════════════════════════════
🔗 LINKS
═══════════════════════════════
🔔 Subscribe for the next one — inside the training run:    / @beneaththeprompt  
📄 Full source list (25 primary sources): pinned comment
🌐 Every source, the corrections log and the newsletter: https://beneaththeprompt.com

#ChatGPT #GPT4 #HowAIWorks #DataCenter #Transformers #MachineLearning #BeneathThePrompt