Blog

A Ray-Focused Guide to PyTorch Conference North America

By September 30, 2026No Comments

Featured projects

TL;DR

In only a few weeks, PyTorch Conference North America 2026 will begin in San Jose, bringing together the open source AI community to share ideas and collaborate. If you’ve heard of Ray but never had a reason to dig in, this is the year as more teams seek to scale AI on top of their own secure and governed AI infrastructure. 

Ray is a distributed compute framework that helps teams scale any AI framework, library or model. It provides a consistent surface for developers to go from curating petabytes of multimodal data, to training a model across multiple nodes and then deploying it in production whether it’s on VMs or a Kubernetes-managed cluster. That coverage across the full AI lifecycle is also why Ray has become the standard substrate for post training and reinforcement learning (RL) at labs like Cursor, Microsoft AI and NVIDIA. Whether you’re new to distributed PyTorch or already running Ray in production, here’s your guide to the sessions worth blocking off on your schedule.

The open source AI compute stack

Keynote: Evolving Ray and Kubernetes Together for the AI Era
Ion Stoica, Anyscale / Databricks / UC Berkeley · Tue 9:35 AM · Grand Ballroom

Ray co-creator Ion Stoica will deliver a keynote on where Ray and Kubernetes are headed together in the AI era.

Crossing the Divide: Co-Evolving Kubernetes and Ray for the AI Era
Jago Macleod, Google, Ion Stoica, Anyscale / Databricks / UC Berkeley · Tue 2:50 PM · Room 210BF

The AI workload orchestration landscape is currently split across two massive, parallel universes: the CNCF (the bedrock of cloud native and modern infrastructure) and the PyTorch Foundation (the epicenter of AI/ML innovation). While these foundations operate independently, the end user does not have the luxury of choosing just one. PyTorch Foundation contains PyTorch, vLLM, Ray, and more, while CNCF owns Kubernetes, Envoy, OpenTelemetry, llm-d, containerd and more. To build, deploy, and scale modern AI applications, users require a seamless integration of projects from both ecosystems. This session will explore the critical bridge connecting these communities: the co-evolution of Kubernetes and Ray into a unified OSS AI Stack. We will discuss how Google and Anyscale, together with the broader community, are actively collaborating to build an open, vertically integrated stack that prevents ecosystem fragmentation.

Production-scale data curation and model training

Ray All the Way Down: A Heterogeneous, Elastic PyTorch Training Stack at LinkedIn
Tommy Li, Tao Huang, LinkedIn · Tue 3:25 PM · Room 210AE

If you want a concrete blueprint for scaling PyTorch past a single machine, this is the session to attend. LinkedIn’s team walks through a three-layer Ray-based training stack: a Ray-actor-based data loading library that streams Avro, Parquet, and Iceberg; a Ray Train framework layered with FSDP, HSDP, elastic scaling, and async checkpointing; and an on-demand cluster provisioning platform. The core lesson? Keep data loading off the GPUs by routing it through dedicated CPU nodes over a zero-copy bridge. This led to a 50–70% cut in dataloader memory and a scenario where they had to work around Ray to reclaim 24% of their step time. This use case is practical, specific, and reusable well beyond LinkedIn’s stack.

PyTorch-Native Feature Transformation and Training Framework for Uber Eats Recommendation
Peng Zhang, Ke Chen, Xandra Zhu, Uber · Oct 21 · 4:55–5:20 PM · Room LL20CD

Uber’s talk on migrating the Eats recommendation stack from TensorFlow/Horovod to PyTorch leans on Ray Data as a core piece of the redesign, replacing legacy Spark-based preprocessing with distributed feature-stat computation and zero-copy batch transformations. The results are striking: a 5x speedup and 90% lower memory usage for data transformation, plus a 20x improvement in training throughput.

Scaling the LLM Serving Layer with Ray

Ray isn’t just powering dedicated training sessions this year, it’s also showing up as critical infrastructure inside talks with a different primary focus:

Keeping GPUs Busy: High-Speed Storage for PyTorch via fsspec
Ankita Luthra, Trinadh Kotturu, Google · Tue 3:40 PM · Room LL21DEF

A bottleneck has shifted from compute to storage: GPUs sitting idle while waiting on legacy REST-based data access. The proposed fix, Rapid Storage, brings a high-throughput gRPC-based protocol to PyTorch via fsspec. According to the team, the payoff extends across the fsspec ecosystem – including Ray, alongside Dask, HF Datasets, and vLLM. If your Ray Data pipelines are storage-bound, this session is worth a visit. 

From PyTorch to Production: Serving a Physics-Constrained Generative Model with ONNX, Ray, and vLLM
Arun Sharma, University of Minnesota · Tue 2:50–3:15 PM · Room 210AE

A deep engineering narrative following a physics-constrained generative model – PC-RF, a conditional rectified-flow model for climate downscaling – from training through to a served, production stack. The talk covers exporting to ONNX two different ways, running the flow-matching sampler outside the graph, and enforcing physics constraints at inference time. Ray shows up as part of the training recipe used to scale the model, with the serving path fanning out into ONNX Runtime, Temporal, and a vLLM agent.

Post-training and RL with Ray

Building a Post-Training Platform for Teams That Don’t Own the Training Loop
Gaurav Arora, Shunyao Li, Eric Wang, Pinterest · Weds 3:25–3:50 PM · Room LL20CD

As post-training expanded across teams at Pinterest (SFT, DPO, GRPO on vision-language models), each built its own stack: different frameworks, data formats, and distributed strategies. Teams spent weeks integrating OSS frameworks with internal infra before training a single step. Our traditional ML platform (MLEnv) owned the PyTorch training loop and reached 95% adoption. Post-training broke that: the loop now lives inside fast-moving OSS frameworks releasing monthly with new techniques. We built PTEnv – a lifecycle harness for PyTorch-based post-training, owning everything around the loop (data, orchestration, scaling, eval, export) while delegating the inner loop to OSS frameworks. The same interface supports subprocess execution with Ray for RL and in-process with MS-Swift using PyTorch FSDP/DDP for SFT – framework choice is a config line, not a rewrite. This talk covers the harness architecture, production gotchas training VLMs at scale (weight sync bottlenecks, MoE workarounds), post training optimization on real workloads, and framework-agnostic eval with LLM-as-judge. A practical talk for platform teams supporting multiple post-training workloads.

SkyRL: Democratizing Scalable RL Training
Sumanth Hedge, Eric Tang, Anyscale · Oct 20 4:55 PM-5:20 PM · LL20A (Lower Level)

Agents have taken center stage in 2026, with ever-growing interest from companies in training custom agents with reinforcement learning. As agents shift to longer horizon, multi-turn interactions, the systems challenges and requirements on underlying training infrastructure have also evolved. SkyRL is a modular, performant RL library built to meet this growing demand for customization and scalability. This talk traces where SkyRL started – a set of modular APIs decoupling training, inference, and environments orchestrated with Ray – and where it is today. The session will cover: 

  • How SkyRL provides scalable fully async RL training on 350 billion+ parameter MoE models with Megatron and vLLM
  • SkyRL’s multi-tenant Tinker Engine, which allows researchers to efficiently use their own hardware for RL training while using Tinker’s flexible training APIs to iterate on recipes
  • SkyRL’s redesign towards HTTP-based APIs for scalable inference, and our contributions of native RL APIs to vLLM
  • Community recipes built on top of SkyRL including large MoE training on long-horizon tasks, custom recursive language models, and more.

Demo Theater

Evolving Ray Core for Post-Training at Scale
Mengjin Yan, (Ray Core), Josh Lee, Anyscale · Oct 21 3:55 PM-4:05 PM · Community Expo (Concourse Level)

Before a post-training job trains a single step, its trainers and inference engines must be scheduled onto the cluster, and every step after that, fresh weights must move between them.  Ray Core evolves directly from the workloads it powers, shaped by the post-training community’s core requirements for higher scalability and stronger performance. Continuous engineering enhancements to Ray Core address these needs directly, establishing it as a key open source runtime underneath post-training frameworks such as SkyRL, Miles, and NeMo RL. The session introduces the underlying building blocks of Ray Core and demonstrates how they map onto an active post-training loop. Key scalability achievements are examined, highlighting how Ray scales scheduling across 10,000-node clusters, alongside performance optimizations like Ray Direct Transport, which moves PyTorch tensors directly between GPUs. The discussion concludes with an overview of the future development roadmap for Ray Core and practical ways for community members to contribute.

Meet the Developers

Tues ·  3:50 PM-4:25 PM · Community Expo

Join Ray experts for an interactive “Meet the Developers” session for a chance to learn in a small group setting and ask your most pressing questions.

Learn more about Ray at PyTorch Conference North America

Taken together, these sessions tell a coherent story: Ray has moved from a scaling library you reach for to foundational infrastructure spanning orchestration, training, data loading, storage, and observability. If you’re building AI infrastructure on PyTorch, don’t miss the Ray sessions at PyTorch Conference North America.

View the complete schedule here

Register for PyTorch Conference North America today