Infer() Summit 2026: Inference at Scale

Опубликовано: 30 Июль 2026
на канале: Momento
1,086
35

Inference is where AI systems meet production. The forces that govern it are ones you already know: routing, scheduling, memory pressure, cache eviction, tail latency, and state movement.

Infer() Summit is a one-day virtual event for engineers running inference at scale. June 25, 2026, 10:00 AM to 4:30 PM PT. Presented by Momento.

The day is built around three pillars:
• Engines — schedulers, KV caches, and the serving systems underneath them
• Operations — reliability, observability, and performance at scale
• Architectures — building inference platforms that run faster and cost less

Speakers: Hien Luu (host), Salina Wu (Pinterest), Harry Kim (NVIDIA), Ying Chen (Databricks), Philip Kiely (Baseten), Emilio Andere (Wafer), Khawaja Shams (Momento), Meryem Arik (Doubleword), Chenyang Zhao (SGLang), and Abi Aryan.

Sessions cover VLM serving on NVIDIA Dynamo, agent execution and scheduling, reliable LLM serving under variable demand, state-of-the-art inference on AMD, multi-stage decoding with SGLang Omni, and async agents in production.

Register: https://luma.com/cluvtqhu
#LLM #Inference #KVCache #AIInfrastructure #MachineLearning #vLLM #GPU #LLMServing