October 20–21 | San Jose, CA
PyTorch, a fully featured framework for building deep learning models, is distinctive for its excellent leverage of GPU acceleration. Open AI innovation is happening on PyTorch, and NVIDIA is foundational to that. We want our platform to help enable all of this innovation.
NVIDIA uses PyTorch, and we have many projects that extend PyTorch into different spaces. For example, NVIDIA TensorRT™-LLM for inference optimizations is built on PyTorch, while NVIDIA PhysicsNeMo extends PyTorch specifically for physics neural networks.
Join us at PyTorch 2026 to share how PyTorch is accelerating your research, discoveries, and data science.
Keynote With Ujval Kapasi
PyTorch Performance for Agents on Vera Rubin
Tuesday, October 20, 10:15–10:25 a.m. PT | Grand Ballroom
In this keynote, Ujval will explore what agentic workloads need from AI systems and how the PyTorch stack can address these needs.
Agentic systems introduce new bottlenecks and optimization opportunities. A single request can trigger long chains of reasoning, tool calls, code execution, data movement, and distributed inference. Training and inference serving software needs to be optimized across CPUs, GPUs, memory, and networking. Using NVIDIA Vera Rubin as a case study, he will show how co-design across hardware and software can improve end-to-end performance, efficiency, and developer experience for agents at scale.
Take a closer look at the scheduled NVIDIA speakers and sessions in this year’s program.
Check out the scheduled NVIDIA sessions, posters, meetups, and more in this year’s program.
10:15–10:25 a.m. PT
Grand Ballroom
Ujval Kapasi, VP, AI & HPC Frameworks and Libraries, NVIDIA
12:20–12:45 p.m. PT
Room 210AE
Marc Romeyn, ML Engineer, NVIDIA
12:20–12:45 p.m. PT
Room 210BF
Michael Goldfarb, Senior Deep Learning Performance Engineer, NVIDIA; Guray Ozen, Principal Compiler Engineer, NVIDIA
12:20–12:45 p.m. PT
Room LL20AB
Aastha Jhunjhunwala, Senior Solutions Architect, NVIDIA; Mark Moyou, Senior Solutions Architect, NVIDIA
2:15–2:40 p.m. PT
Room LL21ABC
Zhiyu Cheng, Engineering Manager, Model Optimizer, NVIDIA; Trenton Starkey, Product Manager, NVIDIA
11:45 a.m.–12:10 p.m. PT
Room LL21ABC
Elias Ellison, Software Engineer, NVIDIA; Daniel Galvez, AI Developer Engineer, NVIDIA
11:45 a.m.–12:10 p.m. PT
Room 210BF
Ke Wen, Software Architect, NVIDIA; Natalia Gimelshein, Software Engineer, Meta; Kapil Shama, Software Engineer, Meta
12:20–12:45 p.m. PT
Room 210AE
Chris Hoge, AI Platform Software, NVIDIA; Joseph Groenenboom, Principal ML Engineer, Red Hat
2:15–2:40 p.m. PT
Room 210AE
Susie Xia, Principal Engineer, NVIDIA; Ruijie Zheng, Research Scientist, NVIDIA; George Kurian, Distinguished Engineer, NVIDIA
2:50–3:15 p.m. PT
Room LL20CD
Itay Alroy, Senior Software Architect, NVIDIA
3:25–3:50 p.m. PT
Room LL20AB
Anjulie Agrusa, Senior ML Engineer, NVIDIA; Ryan Spring, Software Engineer, NVIDIA; Bruce Zitelli, Software Engineer, NVIDIA
4:55–5:20 p.m. PT
Room 210AE
Sreeram Potluri, Distinguished Software Architect, NVIDIA; Artem Polyakov, Senior Software Engineer, NVIDIA
4:55–5:20 p.m. PT
Room 210BF
Christine Cheng, Software Engineer, NVIDIA; Dylan Doblar, ML Infrastructure Engineer, NVIDIA
Tuesday, October 20 | 6 p.m. PT
“Day-0 RL at Scale for Emerging LLM and VLM Architectures”
Huiying Li, Shuang Yu
“Don't Just Trust the Model, Test the Physics: Evaluating PyTorch Models with PhysicsNeMo-CFD”
Kaustubh Tangsali
“Dynamo Snapshot: Autoscaling and Failure Recovery in Seconds Using GMS”
Schwinn Saereesitthipitak, Vikram Sharma Mailthod
“Efficient, Large-Scale LoRA and Fine-Tuning for Diffusion Models in PyTorch”
Pranav Prashant Thombre
“Elastic Endpoint P2P Transfers for Inference with NIXL”
Samuel Nordmann
“Exploiting CUDA Locality Domains in PyTorch on Blackwell and Beyond”
Matthias Jouanneaux
“Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML”
Meghan Cowan, Srinivas Sridharan
“FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning”
Jingqi Zhang, Zhaopeng Qiu, Shuang Yu
“From MP4 to Tensor: Optimizing Multi-Camera Video Data Loading for PyTorch Training”
Kyle Huang, Mihai Alden
“From Scheduled to Real Overlap: All-Gather/Reduce-Scatter in PyTorch Distributed”
Jeffrey Mahou
“Graphing the Ungraphable: Making MoE Training CUDA-Graphable in PyTorch”
Syed Ahmed
“Hardware-Accelerated Strided KV-Cache Transfers for Asymmetric LLM Inference”
Michal Shalev
“Inductor-TV: Formal Methods for the PyTorch Compiler”
Abhilash Majumder
“Maestro: Benchmarking Concurrent GPU Operations as They Run in Real Workloads”
Eyal Chocron
“Memory Management in CUDA Graphs”
Frank Lin
“Multi-Teacher On-Policy Distillation for SOTA Agentic AI Models”
Jiaqi Zeng
“Network Fault Tolerance for AI Workloads”
Evgeny Leksikov
“Production FP8 Training for Multimodal Transformers in PyTorch”
Iain Weissburg, Wei Chen
“Race-Free TorchInductor Artifact Reuse for Elastic Training”
Christine Cheng, Kwanghoon An
“Resilient PyTorch Distributed with NCCL Shrink and Revoke”
Bruce Chang
“Scaling Multi-Model PyTorch Workloads for Closed-Loop Robotics Evaluation”
Nan Zhu
“Shape-Stable Dynamic Control Flow in PyTorch CUDA Graphs”
Daniel Galvez
“StageFrontier: Always-On Stage Accounting for Distributed PyTorch Training”
Boram Yoon, Aaron Meng
“TorchTitan Performance Optimizations”
Elfie Guo, Syed Ahmed
“Using HugePages with PyTorch for Faster LLM Inference”
Ofer Achler
NVIDIA is a foundational sponsor and key participant at PyTorch Conference 2026, presenting keynotes, technical sessions, and poster exhibits highlighting GPU acceleration, distributed AI, and framework optimizations.
Reference: PyTorch Conference 2026 Event Details provides supporting information about this topic.
PyTorch Symmetric Memory allows developers to build fused compute-communication GPU kernels that bypass standard collective operations for faster distributed AI training.
Reference: PyTorch Symmetric Memory Session provides additional technical detail.
NCCL (NVIDIA Collective Communications Library) Extensions provide custom communication patterns designed to optimize inter-GPU data transfers in modern distributed AI workloads.
Reference: NCCL Extensions Overview provides supporting information about this topic.
Parameterized Dynamic Shape CUDA Graphs allow CUDA Graph execution in PyTorch to support varying tensor input dimensions without incurring full graph recapture overhead.
Reference: CUDA Graphs with Dynamic Shapes provides supporting information about this topic.
The PyTorch Ecosystem Working Group is a collaborative community initiative aiming to expand library support, tooling integration, and open-source contributions across the PyTorch framework.
Reference: PyTorch Ecosystem Working Group provides additional technical detail.
NVIDIA hardware and PyTorch software enable multimodal models to process text, image, and video tokens simultaneously with accelerated attention mechanisms and efficient precision scaling.
Reference: Modern Vision Language Models Session provides supporting information about this topic.
NVFP4 is an ultra-low precision FP4 quantization and data format supported on modern NVIDIA GPU architectures to enable efficient, high-throughput pretraining and inference for LLMs.
Reference: LLM Pretraining in NVFP4 provides supporting information about this topic.
Register now to join NVIDIA at PyTorch 2026.