Featured projects
TL;DR
The Accelerator Integration Working Group plays a vital role in standardizing how new hardware architectures connect to the open source AI ecosystem. As compute platforms diversify across cloud, edge, and specialized silicon, the Accelerator Integration Working Group establishes clear, vendor-neutral integration mechanisms that eliminate custom code patches and reduce ecosystem fragmentation. In this blog, we cover what the group achieved in H1 2026, including key infrastructure improvements, unified testing workflows, and reference backends across the entire PyTorch platform.
Introduction
The PyTorch hardware ecosystem keeps diversifying, with a growing range of AI compute platforms adopting native framework integration. This expanding adoption creates a shared community challenge: streamlining integration paths, lowering complexity, standardizing integration mechanisms, and establishing reliable, unified validation infrastructure for heterogeneous compute platforms.
That is the core mission of the Accelerator Integration Working Group. With a long‑term vision for a scalable, inclusive PyTorch hardware ecosystem, cross‑community contributors work on upstream framework improvements, shared tooling, and standardized reference implementations.
In 2026 H1, we made tangible progress across multiple core workstreams: refactored test suites for broader cross‑backend reuse, the Cross‑Repository CI Relay (CRCR) mechanism, profiling support for PrivateUse1‑based backends, ongoing maturation of OpenReg as the reference backend, distributed capability support, compiler backend integration, and enhanced CI infrastructure and visibility dashboards. This post outlines these major achievements and their impact on PyTorch integration for compute platforms across the broader ecosystem.
Cross-Repository CI Relay
Authors: Subin George, Jiahao Chen, Jiahao Tan, Jiawei Li, Jewel K. M.
Goals
The PyTorch community has introduced Cross-Repository CI Relay (CRCR), a new system designed to close a long-standing visibility gap in its CI ecosystem. PyTorch sits at the center of a large network of dependent projects – accelerator like Intel XPU, AMD ROCm, Apple MPS, Qualcomm AI Engine and various accelerator integrated via PrivateUse1 mechanism, as well as ecosystem libraries such as vLLM, SGLang, and Hugging Face Transformers. While PyTorch’s own upstream CI is mature and extensive, it runs within pytorch/pytorch repo, leaving downstream repositories without a standard way to know when to test against upstream changes, report results back, or correlate failures across projects. This introduced a significant coordination challenge: PyTorch contributors could not tell whether their PRs would impact other backends prior to merge, and hardware maintainers could only offer feedback once changes had already been committed.
Benefits
CRCR solves this with a fully automated pipeline. When a PR is opened or a commit is pushed to pytorch/pytorch, a webhook triggers dispatch events to all registered downstream repositories in parallel.

Figure 1: CRCR Architecture
Each repo runs its own CI workflow and reports status back via an authenticated callback, using a GitHub OIDC token that cryptographically verifies the calling repository’s identity. Results then flow into the PyTorch CI HUD (hud.pytorch.org/crcr) within seconds, giving maintainers a single dashboard for both in-tree and cross-repository CI health.

Figure 2: CRCR Hud Dashboard
The system uses a tiered allowlist with four participation levels (L1–L4), letting downstream repos progress from simple dispatch notifications, to full HUD reporting, to non-blocking and eventually blocking check runs on upstream PRs. Security is enforced through five layers of validation – OIDC identity verification, allowlist authorization, rate limiting, state-machine checks, and strict separation of trusted versus self-reported data – ensuring that even a compromised downstream repo can only affect its own displayed CI status, not PyTorch’s actual build infrastructure or merge decisions. Onboarding requires minimal effort from downstream maintainers: Just an allowlist entry and a lightweight workflow file, with no custom authentication code needed.
Test Refactoring
Authors: Riya Punia, Tanmay Kumar, Jiahao Chen, Jiahao Tan, Jiawei Li
Goals
Validating a new accelerator backend against PyTorch’s existing test suite is one of the most effective ways to ensure implementation correctness and catch upstream changes early. PyTorch maintains over 600,000 test cases covering operators, autograd, profiling, distributed training, and more – a massive validation surface that, in principle, any backend should be able to leverage. In practice, however, many of these tests were written for a specific set of accelerators. Device strings, profiler activities, memory APIs, and skip decorators are hardcoded throughout, tightly coupling test logic to specific backends rather than expressing it in device-agnostic terms. Before this refactor, the only option was to maintain custom test cases via patches. Every PyTorch upgrade demanded significant manual effort, and this duplicated work was further exacerbated by PyTorch’s faster release cadence.
In H1 2026, the working group launched a systematic effort to decouple PyTorch’s test suite from specific hardware. This work is tracked through a central tracking issue, an RFC on test case refactoring, and an RFC on test class classification, with 259 items tracked in the PyTorch Test Refactoring project.

Figure 3: Test Refactoring Architecture
The effort has three dimensions:
Device-agnostic test migration: Replacing hardcoded device references (device=”cuda”, torch.cuda.synchronize(), @onlyCUDA) with parameterize equivalents (device=device, getattr(torch, device.type), instantiate_device_type_tests()). Each test class is classified as accelerator-unrelated (CPU-only, targeting generic features like the Dispatcher), accelerator-agnostic (device-parameterized, runs across all backends), or accelerator-specific (tied to a particular backend’s internal APIs). During H1, contributors migrated over 276 test files spanning dynamo, profiler, nn modules, linalg, optimizers, convolutions, serialization, multiprocessing, and dataloader.
Hardware classification metadata: Building on the test class classification proposed by Alban, the group implemented a hw_classification class attribute (GENERIC, DEVICE_GENERIC, CUDA, XPU, MPS) and a –hw-classification flag across all execution paths — unittest, pytest, subprocess, parallel, XML output, and run_test.py forwarding — enabling CI runners to select the correct test subset for their hardware automatically.
CI guardrails: A HW_CLASSIFICATION linter enforces that every new test class declares its classification. Existing unclassified files (1,191 at launch) are allowlisted and being driven to zero incrementally —matching the phased rollout agreed upon with PyTorch maintainers, with a future AI linter layer planned for more nuanced enforcement.
Benefits
For accelerator developers, these changes fundamentally shift the onboarding experience. Instead of patching numerous test files to adapt them for their hardware, vendors can rely on instantiate_device_type_tests() to automatically generate test variants for their backend. Tests that were previously invisible to non-CUDA backends – because they were gated behind @onlyCUDA or hardcoded device=”cuda” – now run on any registered accelerator. A single device-agnostic test produces validated configurations across CUDA, MPS, XPU, ROCm, and PrivateUse1 backends simultaneously.
For the PyTorch project itself, the hw_classification system provides the infrastructure foundation, which proposes class-level hardware and frequency metadata to reduce double-testing, improve CI observability, and rationalize test scheduling. The refactoring work done in H1 – migrating tests into properly classified classes – is a strict prerequisite for that vision.
For the broader ecosystem, the multi-level test skipping mechanism (by feature, class, test, and operator) combined with hardware classification gives backend vendors fine-grained control over which tests to run and which to skip based on their implementation maturity – without modifying upstream test code.
Next Steps
In H2, the priority is driving the unclassified allowlist toward zero, extending device-agnostic coverage to distributed, JIT, and autograd modules, and wiring hw_classification into CI scheduling for automatic test selection per runner. The group is also exploring a declarative per-operator capability registry that lets accelerators declare supported dtypes and precision overrides without modifying upstream OpInfo entries.
Profiling
Authors: Vishal Goyal, Anisha Kushwaha
Goals
Integrating a new hardware accelerator with PyTorch profiling requires more than operator-level timing. Vendors need a clear contract for kernel-level timelines, memory and runtime events, and CPU–device correlation in Chrome/Perfetto traces. Without a working out-of-tree reference, backend developers are left with two incomplete options: the legacy ProfilerStubs / KINETO_PRIVATEUSE1_FALLBACK path, which only records coarse per-operator device time, or reverse-engineering in-tree Kineto plugins such as CUPTI and XPUPTI, which conflate vendor-specific tooling with the integration patterns themselves.
In H1 2026, profiling followed OpenReg’s usual minimal, stub-like approach: not a production profiler, but just enough to demonstrate the integration mechanics. Building on the PyTorch core registration API (REGISTER_PRIVATEUSE1_PROFILER), which lets out-of-tree backends plug an IActivityProfiler into Kineto without further changes to PyTorch or Kineto, we landed an OpenReg reference stub stack (tracer plus IActivityProfiler / session, RFC #177978) covering session lifecycle, activity types, and correlation IDs, and backed it with end-to-end session-lifecycle and fallback tests using torch.profiler and documentation for both the legacy stubs path and the Kineto plugin path.
Benefits
For accelerator developers, the OpenReg profiling path provides a concrete starting point for kernel-level bring-up. Instead of reverse-engineering CUPTI-style plugins inside Kineto or settling for coarse ProfilerStubs timing, new accelerator teams can follow a single, minimal stub implementation that shows how an out-of-tree backend registers an IActivityProfiler, manages session lifecycle, and wires the correlation-ID call path – then replace the stub bodies with their own device profiling SDK (PTI) calls, without modifying PyTorch or Kineto. The legacy stubs fallback remains available when only operator-level timing is needed.
For the broader PyTorch ecosystem, this turns PrivateUse1 profiling into a validated integration surface rather than a per-vendor addition to Kineto’s source tree. End-to-end torch.profiler tests catch registration, session-lifecycle, and fallback regressions early, and the stub reference stays aligned with the documentation so the workflow can be trusted for any backend.
OpenReg
Authors: Mansi Agarwal, Jiahao Chen, Jiahao Tan
Goals
Integrating a new hardware accelerator with PyTorch requires implementing a wide surface of functionality – device registration, operator dispatch, stream and event management, autograd integration, and more. Without a working reference, backend developers often resort to reverse-engineering production backends, which conflates device-specific optimizations with the integration patterns themselves.
OpenReg is PyTorch’s in-tree reference backend for PrivateUse1-based accelerator integration. It is not a production backend – it is a minimal, CPU-backed implementation that isolates the integration mechanics from hardware complexity. In 2026 H1, we focused on three areas: expanding coverage of core integration patterns (device registration, operator dispatch, streams, events), tightening the correspondence between OpenReg’s implementation and the accelerator integration documentation, and strengthening OpenReg’s role as a validation backbone for PrivateUse1 integration paths through testing and CI.
Benefits
For accelerator developers, OpenReg provides a concrete starting point for bring-up work. Instead of piecing together integration patterns from scattered examples and production backends, new accelerator teams can follow a single, minimal implementation that demonstrates how an out-of-tree backend is expected to interact with PyTorch, from device module registration through operator dispatch and runtime behavior.
For the broader PyTorch ecosystem, OpenReg serves as a regression guard. Exercised through CI, it validates that PrivateUse1 integration mechanisms remain stable as PyTorch evolves, catching breakages before they reach downstream hardware vendors. The tighter alignment between OpenReg and the accelerator integration documentation also means that the documented workflow can be trusted – if it works in OpenReg, it works for your backend. It also provides the common foundation for the more advanced integration paths covered later in this post like distributed training, compile backend support, and profiling all built on top of the base patterns that OpenReg establishes and validates.
Next Steps
In H2, OpenReg will continue to expand into additional PyTorch subsystems, while validating that PrivateUse1 integration paths remain stable across releases.

Figure 4: OpenReg Integration Architecture
Distributed Support
Authors: Mansi Agarwal, Atharva Kshirsagar
Goals
Distributed training is a core requirement for large-scale model development, but integrating a new accelerator with PyTorch’s distributed stack requires navigating a broad surface area, including custom transport, ProcessGroup registration, collective dispatch, and multi-process coordination. Without a clear reference, that path can be difficult for backend authors to follow.
In 2026 H1, the working group began addressing this gap by building OCCL (OpenReg Collective Communications Library), a minimal reference implementation of a custom c10d backend for OpenReg (RFC #176877). Rather than targeting production performance, OCCL is designed to make the main integration points explicit. It shows how a backend can register a ProcessGroup, dispatch collectives, and manage Work completion semantics in a way that aligns with PyTorch’s distributed abstractions (#171250, #183257).
This effort also helped improve the surrounding integration experience. Alongside OCCL, the contributors from working group added the documentation to the accelerator integration guide (#182308) and supported upstream improvements in PyTorch’s distributed infrastructure, including better support for Python-based backend implementations and device-generic distributed test selection.
Benefits

Figure 5: Distributed OCCL Architecture
For accelerator developers, OCCL provides a much clearer starting point for adding distributed support to a new backend. Until now, vendors often had to study production backends such as NCCL to infer the c10d integration contract. While those backends are essential in practice, they are tightly coupled to hardware-specific communication libraries and are not ideal as learning-oriented references.
OCCL makes that contract easier to understand by presenting a standalone, hardware-independent path through the core pieces of distributed integration. Backend authors can follow a simpler model for how ProcessGroup registration, collective execution, and completion semantics fit together inside PyTorch. In combination with the related upstream improvements, this gives new accelerators a more reliable foundation for bringing up distributed training support within PyTorch’s existing infrastructure.
Compile Backend Support
Authors: Parshant Sharma
Compiler integration is one of the important pieces for any accelerator that wants to work well within PyTorch. Integrating a new accelerator with torch.compile requires navigating a broad surface area, including graph capture and device management in Dynamo, and optionally scheduling, fusion, and code generation in Inductor for vendors who want to leverage its optimization passes. Without a clear reference, backend authors are often left studying production integrations like CUDA and Triton to piece together the registration contract. That path can be difficult to follow because it mixes hardware-specific optimizations with the underlying integration mechanics. This work, outlined in RFC #181093, addresses that gap.
Rather than targeting production performance, the work is designed to make the Dynamo and Inductor integration points explicit. Using OpenReg, it shows how a device backend can register with both layers of the compiler stack, validate graph capture, and generate fused kernels, all without modifying upstream PyTorch code.
Beyond the code itself, this work helped surface gaps in the existing documentation and developer experience. The working group contributed a compiler integration guide to PyTorch’s accelerator docs, covering both Dynamo and Inductor paths with step-by-step instructions that reference the OpenReg implementation as a concrete example.
Benefits
This work provides a much clearer starting point for adding compiler support to a new backend. Until now, vendors often had to study CUDA and Triton codepaths to infer the registration contracts. While those backends are essential in production, they are tightly coupled to hardware-specific scheduling and codegen strategies.
Through the Dynamo backend and Inductor integration, backend authors can see how Dynamo graph capture, Inductor scheduling, wrapper codegen, and device operation overrides fit together inside PyTorch. This gives out-of-tree backends a more reliable foundation for bringing up torch.compile support, from basic graph capture all the way through fused kernel generation, within PyTorch’s existing infrastructure.
PyTorch Additional Platform Page
Authors: Jiahao Chen, Jiawei Li
Goals
For a long time, users of the PyTorch community were missing an official channel to get clear information regarding PyTorch support for a new backend. The PyTorch Additional Platforms page is now the official, foundation-governed home on pytorch.org for compute platforms, linked directly from the main install page every PyTorch user already visits. To keep the page meaningful rather than a self-reported free-for-all, the Accelerator Integration Working Group runs it under a formal admission process.
Applying is deliberately simple: open an Additional Compute Platform Application issue, provide evidence (links, docs, CI dashboards, security policy etc.) for each requirement, and submit.
In H2, we’re actively encouraging more accelerator vendors, especially those who’ve already done the integration work described elsewhere in this recap, to apply and get their platform in front of the full PyTorch user base. Full details on requirements and the review process are in the Admission Process document or raise questions in the Working Group Slack channel.
Conclusion
The progress achieved by the Accelerator Integration Working Group in H1 2026 demonstrates the power of vendor-neutral collaboration in open source AI. By establishing standardized testing, profiling, compiler support, and distributed execution pathways, the community has built a foundation that allows hardware innovation to flourish without fracturing the framework. As silicon architectures continue to evolve, these standardized integration mechanisms ensure that new compute platforms can achieve production readiness seamlessly.
To explore the ongoing workstreams or get involved in defining next-generation hardware enablement for PyTorch, visit the Accelerator Integration Working Group repository at https://github.com/pytorch-fdn/accelerator-integration-wg or join the community discussion on the PyTorch Technical Advisory Council working group page at https://pytorch.org/working-groups/.
Acknowledgements
We sincerely appreciate the generous support and guidance from PyTorch maintainers and community members throughout 2026H1. Special thanks to @alband, @afrittoli, @atalman, @jbschlosser, @marco-s, @matthew-d-white,@mikaylagawarecki, @ZainRizvi, @zxiiro … for their invaluable feedback and continuous help. We are also grateful to all community contributors for their participation and assistance throughout this work.
Find more information about Accelerator Integration Working Group: https://github.com/pytorch-fdn/accelerator-integration-wg
PyTorch TAC Working Groups: https://pytorch.org/working-groups/