This paper reveals how model size fundamentally changes attention patterns during in-context learning, with smaller models focusing on key features for robustness while larger models spread attention across more features, making them paradoxically more sensitive to noise despite their greater capacity.
Related Videos
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
DeepSeek-R1: • DeepSeek-R1 - Paper Walkthrough
The Illusion of Thinking: • The Illusion of Thinking - Paper Walkthrough
SAM2: • SAM2: Segment Anything in Images and Video...
Google's AlphaEvolve: • Google's AlphaEvolve - Paper Walkthrough
Chain-of-Verification (COVE) Reduces Hallucination in Large Language Models: • Chain-of-Verification (COVE) Reduces Hallu...
Why Language Models Hallucinate: • Why LLMs Hallucinate
Transformer Self-Attention Mechanism Explained: • Transformer Self-Attention Mechanism Visua...
Jailbroken: How Does LLM Safety Training Fail? - Paper Explained: • Jailbroken: How Does LLM Safety Training F...
How to Fine-tune Large Language Models Like ChatGPT with Low-Rank Adaptation (LoRA): • Low-Rank Adaptation (LoRA) Explained
Multi-Head Attention (MHA), Multi-Query Attention (MQA), Grouped Query Attention (GQA) Explained: • Multi-Head Attention (MHA), Multi-Query At...
LLM Prompt Engineering with Random Sampling: Temperature, Top-k, Top-p: • LLM Prompt Engineering with Random Samplin...
Contents
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
00:00 - Intro
00:24 - Main contributions
00:45 - Evaluation dataset
01:06 - First setting
02:07 - Second setting
02:36 - Main findings
03:20 - Limitations
03:38 - Outro
Follow Me
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
🐦 X: @datamlistic https://x.com/datamlistic
📸 Instagram: @datamlistic / datamlistic
📱 TikTok: @datamlistic / datamlistic
Channel Support
▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬▬
The best way to support the channel is to share the content. ;)
If you'd like to also support the channel financially, donating the price of a coffee is always warmly welcomed! (completely optional and voluntary)
► Patreon: / datamlistic
► Bitcoin (BTC): 3C6Pkzyb5CjAUYrJxmpCaaNPVRgRVxxyTq
► Ethereum (ETH): 0x9Ac4eB94386C3e02b96599C05B7a8C71773c9281
► Cardano (ADA): addr1v95rfxlslfzkvd8sr3exkh7st4qmgj4ywf5zcaxgqgdyunsj5juw5
► Tether (USDT): 0xeC261d9b2EE4B6997a6a424067af165BAA4afE1a
#llm #deepseek #reasoning #machinelearning #deeplearning #ai