24949 подписчиков
98 видео
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling (Jan 2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction (Apr 2024)
Recursive Language Models (Dec 2025)
uGMM-NN: Univariate Gaussian Mixture Model Neural Network (Sep 2025)
Code as Agent Harness (May 2026)
$δ$-mem: Efficient Online Memory for Large Language Models (May 2026)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (June 2025)
Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining (Apr 202
Benchmarking Vision-Language Models on OCR in Dynamic Video Environments (Feb 2025)
Back to Basics: Let Denoising Generative Models Denoise (Nov 2025)
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent (Aug 2025)
FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space (May 2025)
LLMD: A Large Language Model for Interpreting Longitudinal Medical Records (Oct 2024)
Benchmarking On-Device Machine Learning on Apple Silicon with MLX (Oct 2025)
MEMO: Memory as a Model (May 2026)
Scalable watermarking for identifying large language model outputs (Oct 2024)
FAPO: Fully Automated Prompt Optimization of Multi-Step LLM Pipelines (Jun 2026)
Simplifying DINO via Coding Rate Regularization (February 2025)
Learn to Guide Your Diffusion Model (Oct 2025)
Kwai Keye-VL-2.0 Technical Report (Jun 2026)
U.S. Policies Unintentionally Accelerated China's Open AI Ecosystems (Jun 2026)
You Don't Need to Run Every Eval (Jun 2026)
Apple Neural Engine: Architecture, Programming, and Performance (Jun 2026)
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments (Jun 2026)
PixelFlow: Pixel-Space Generative Models with Flow (Apr 2025)
Generative Modeling with Bayesian Sample Inference (Feb 2025)