Open Science and LLMs, with Dr. Valentina Pyatkin

Опубликовано: 14 Август 2026
на канале: Women in AI Research WiAIR
385
8

Can open-source large language models really outperform closed ones like Claude 3.5? 🤔

In this episode of the Women in AI Research podcast, Jekaterina Novikova and Malikeh Ehghaghi engage with Valentina Pyatkin https://valentinapy.github.io/, a postdoctoral researcher at ‪@allenai‬.

We dive deep into the future of open science, LLM research, and extending model capabilities.

🔑 Topics we cover:

• Why open-source LLMs sometimes beat closed models
• The value of releasing datasets, recipes, and training infrastructure
• The role of open science in accelerating NLP innovation
• Insights from Valentina’s award-winning research journey

ToC:
00:00 Introduction - Valentina Pyatkin, Open Science and Language Models
02:52 Valentina's Journey and Mentorship in AI
06:18 Defining Success in Research
08:48 Visibility and Communication in Research
11:34 The Role of Open Science in AI Advancement
16:56 Exploring Tulu3 and Its Innovations
21:47 Data Ethics and Quality in AI Research
24:14 Community Impact and Adoption of Open Tools
26:06 Behavioral Focus in Post-Training Techniques
29:46 Complexity in Post-Training Methods
32:09 Performance Gains from Multi-Stage Post-Training
38:12 Understanding Verifiable Rewards in AI Training
39:39 The Shift from Traditional RLHF to RLVR
41:03 Persona-Driven Synthetic Data Generation
44:17 Advancements in Precise Instruction Following
51:39 Balancing Constraints and Response Quality
54:32 Challenges in Benchmarking AI Models
58:07 Generalization Failures in Evaluation Practices
59:38 RewardBench 2: A New Approach to Evaluating Reward Models
01:06:49 Insights from Open Research Practices
01:08:09 Future Directions in AI Research
01:09:57 Creating an Inclusive Research Environment

REFERENCES:
00:44 Valentina Pyatkin Google Scholar profile (https://scholar.google.com/citations?...)
13:40 Olmo: Accelerating the science of language models (https://arxiv.org/abs/2402.00838)
16:57 Tulu 3: Pushing Frontiers in Open Language Model Post-Training (https://arxiv.org/abs/2411.15124)
25:01 open-instruct (https://github.com/allenai/open-instruct)
43:55 Generalizing Verifiable Instruction Following (https://arxiv.org/abs/2507.02833)
59:45 RewardBench 2: Advancing Reward Model Evaluation (https://arxiv.org/abs/2506.01937)

🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.

WiAIR website:
♾️ https://women-in-ai-research.github.io

Follow us at:
♾️ LinkedIn:   / women-in-ai-research  
♾️ Bluesky: https://bsky.app/profile/wiair.bsky.s...
♾️ X (Twitter): https://x.com/WiAIR_podcast

#AI #LLM #OpenSource #WomenInAI #AIResearch #wiair #wiairpodcast