VideoGameBench: Can Vision-Language Models complete popular video games?

Опубликовано: 15 Май 2026
на канале: Richard Aragon
273
10

Link to Arxiv Research Paper: https://arxiv.org/pdf/2505.18134

This video discusses a research paper from Princeton University titled "Video Game Bench: Can Vision Language Models Complete Popular Video Games?" [00:02]. The presenter shares their long-standing interest in this field, even mentioning a past attempt to involve investors in a similar concept [00:17].

The central theme is the "Video Game Bench," a benchmark developed by researchers to evaluate the capability of Vision Language Models (VLMs) in playing and finishing a variety of video games [01:42].

Key highlights include:

Video Game Bench: This benchmark includes 23 different video games from handheld consoles like Game Boy and Game Boy Color, as well as PC games from Microsoft DOS [02:18, 05:25]. The objective for the VLMs is to complete the main goal of each game, such as defeating a final boss or finishing a campaign [05:37].
Performance on Benchmark: The study revealed that current VLMs performed poorly, with only two models demonstrating any game completion: GPT4.0 completed 0.9% of Pokémon Crystal, and Gemini 2.5 Pro completed 4.8% of Kirby's Dreamland DX [02:01, 06:27].
Rigorous Standards: The speaker emphasizes that this benchmark is considerably more demanding than other existing methods, like Cloud Play's AI. It intentionally omits helper tasks and tools that could give the models an unfair advantage [06:55, 07:01].
Video Games as AI Training Tools: The presenter strongly advocates for using video games as effective training models for AI systems [03:33, 09:02]. They also mention a business idea from two years prior that focused on using AI in video games to train AI models through reinforcement learning [03:46].
Future Prospects: Despite the current low performance, the speaker views the small percentage of completion in some games as a significant indicator of future development potential in this area [08:05]. They consider themselves an expert on this topic and express interest in an advisory role [11:02].
Games Included: Some of the games mentioned in the benchmark are Quake, Prince of Persia, Super Mario Land, Doom, Warcraft 2, Oregon Trail, XCOM, Scooby-Doo, Age of Empires, and Pokémon Red [05:56].
Framework Details: The benchmark is Python-based and uses emulators to run the games [05:49].