Featured projects
TL;DR
PyTorch’s Cross-Repository CI Relay (CRCR) gives out-of-tree accelerators a clean, scalable way to plug into upstream CI – and it leaves each backend free to decide which of PyTorch’s tens of thousands of tests actually matter for its hardware, and how to keep that answer current as PyTorch, the backend, and the test suite all evolve. This post walks through how Torch Spyre built on top of that foundation to reach CRCR’s L2 integration, using an agentic pipeline that selects and re-buckets tests from code and execution evidence, a declarative YAML framework that adapts upstream tests without patching them, and a set of workflow patterns for reliability and reproducibility at scale. Most of this is applicable for any generic privateuse1 device, and we’re looking to work with the pytorch community to make it reusable and upstream where relevant.
Introduction
PyTorch’s Cross-Repository CI Relay (CRCR) closes the coordination gap between PyTorch and out-of-tree (OOT) accelerator repositories by providing a standardized way to trigger downstream CI on upstream changes, surface results directly to PyTorch reviewers, and present regressions into a unified view through PyTorch CI CRCR HUD. This is done by allowing OOT accelerator repositories to receive dispatches from pytorch/pytorch and report results back to the HUD with minimal integration burden consisting of an allowlist entry, a workflow that listens for repository_dispatch, and performs a composite callback action. Further, CRCR provides four levels of integration, enabling downstream projects to incrementally onboard from receiving upstream dispatch notifications to reporting results through HUD and ultimately participating in non-blocking or blocking upstream PR validation.
The interesting engineering starts one step later, in the decisions the relay leaves to each backend: what are the test candidates, which dispatches deserve a build, which of PyTorch’s tens of thousands of tests are meaningful on your hardware, how to adapt tests written for CUDA without forking them, and what “green” is allowed to mean. Those are the right decisions to own, because only the backend can make them. This post describes the mechanisms we built to make them, consisting of an agentic test-selection pipeline, a declarative test-reuse framework, and reliable workflows. These together took Torch Spyre to L2 integration. None of it is Spyre-specific by design, that is the accelerator sits behind the config, not inside the CI logic, so rather than leave these as patterns for each backend to reimplement, we intend to work with the PyTorch team to upstream the reusable parts, with PyTorch OpenReg as the natural reference point.
Three Evolving Candidates under Test: OOT Accelerator PyTorch Backend, PyTorch Core, Test Suite

Figure 1: Testing Surface in PyTorch and Torch Spyre
Challenge 1: Deciding what to test: three moving targets, and which combination to run
Often, OOT accelerators have to deal with three evolving candidates to test after receiving a dispatch from PyTorch CRCR. First, an OOT accelerator PyTorch backend codebase, such as Torch Spyre. As the PyTorch backend for the IBM Spyre Accelerator, Torch Spyre relies on deep integration through PyTorch’s OOT extension interfaces. This exposes a broad testing surface across hooks into PyTorch core, including operators and runtime behaviour. The second candidate is the PyTorch core itself, where changes to runtime components, or other core functionality can introduce regressions in the OOT backend. Finally, there is the evolving test suite, where tests could have been modified, potentially uncovering regressions. Regressions could be introduced by any of these candidates. Testing the OOT accelerator backend against the PyTorch test suite helps surface regressions across all three candidates.
Knowing the three candidates leads directly to the next decision: which combination of them to actually run. When a dispatch arrives, the OOT accelerator backend may have progressed since the last one, PyTorch core has new commits, and the test suite itself may have changed. The primary combination is the latest of all three that is the newest backend code, the newest PyTorch core, and the test suite as it stands at the time of the dispatch. That is what makes a regression attributable: when the only thing that moved since the last green run is upstream, a new failure points at upstream.
Optionally, this can be expanded to include other useful combinations, such as testing a range of OOT accelerator backend versions against a moving PyTorch core and test suite. This can help uncover forward- and backward-compatibility issues for each OOT accelerator backend version. Further, each combination can be reported to HUD as a separate job, making the results and regressions associated with each combination easier to track.
Taming the PyTorch Test Suite
Challenge 2: Identifying the candidate tests for your OOT accelerator backend

Figure 2: Agentic Workflow for Test Selection
PyTorch has tens of thousands of tests. Manually identifying which ones are relevant to your backend doesn’t scale and keeping that selection current as both PyTorch and your backend evolve is worse. Hence, we built an agentic pipeline to do it.
The pipeline runs in four stages (refer Figure 2). The first two stages build context; the last two make decisions: first from code, then from execution.
- High-level selection. A first agent narrows the search space from the whole test tree to the folders and top-level files worth analysing, using two inputs: which PyTorch extension points the backend actually hooks into, and plain-language preferences about scope. A backend that does not hook
torch.dynamo, or that is not ready to take on autograd, says so here and skips those trees entirely. - Repository memory. A Repository Memory Generator turns that shortlist into a queryable index which includes symbols, files, LLM-written summaries and per-test embeddings. This lets the selection agent in the next phase to reason over the whole candidate set at once instead of being limited to whatever fits in a single context window.
- Low-level selection. The second agent then reads the backend’s own codebase, its documentation, and metadata (e.g., supported operators), then queries the repository exact tests to run and buckets them into
mandatory_successorskip. Every decision is recorded with the reasoning behind it. - Refinement from real execution. The selected tests run on real hardware. The agent takes a second pass using execution logs that leads to catching runtime failures and numerical differences that static analysis misses. The output is a corrected, re-bucketed test set with updated reasoning.
This pipeline currently produces a per-file config for each upstream test file we run, covering thousands of individually named test cases, and a merge run exercises them on real hardware. Coverage grows automatically enabling a new op in the backend to admit every test that was waiting on it, with no manual list maintenance.
Auditability matters: Every bucketing choice is written back into the config as a comment explaining why. When someone asks “why is this test skipped?”, the answer is in the file, not in a model’s memory. The agent proposes; the config is the reviewable artifact of record.
Furthermore, the agentic pipeline can be customized for specific use cases. Two such use cases are discussed below.
Use Case A: Enabling a New Op
When a new op is enabled in the backend, which tests should I turn on?
This involves preparing a test-to-op mapping, where, for each test, we record the operators it exercises and build a reverse index that maps each operator to the complete list of tests that exercise it. The Repository Memory Generator captures this metadata during a one-time execution pass:
- Eager path: captured via TorchDispatchMode
- Compile path: captured via TORCH_LOGS
However, the tests captured through this process cannot be added directly to the test suite. Although a test may exercise a desired operator, it may also exercise other PyTorch features that are either not relevant to the OOT accelerator or are not currently supported. Therefore, the operator information, along with the identified test files, is passed to the low-level selection agent. The agent uses this information to further filter the candidates and generate the final set of tests and files that can be added to the test suite.
Result: Enabling one op surfaces all relevant tests without any manual curation.
Use Case B: Upgrading PyTorch Versions
When upgrading to a new PyTorch version, how can test selection be updated without reprocessing everything?
Across PyTorch versions, the test suite can evolve as new tests are added and existing tests are modified or removed. These changes can be incorporated into the repository memory to reflect the updated test suite. The low-level selection agent can then be run only on the delta that is the newly added and modified tests, to identify the changes relevant to the OOT accelerator and generate a revised test suite. This avoids reprocessing the entire test suite while keeping the selected tests aligned with the updated PyTorch version.
Result: Version upgrades become config diffs, not full reprocessing. For example, upgrading from PyTorch 2.13 to 2.14 meant evaluating ~4000 changed tests instead of tens of thousands.
Challenge 3: Adapting the candidate tests
PyTorch tests are often parameterized across dtypes, shapes, and operators. Skipping an entire test because one dtype isn’t supported is too coarse, in such cases, only that dtype is excluded while the rest run. Deeper adaptations (e.g., removing CUDA-specific hardcoding) traditionally require patching upstream. To avoid that, Torch Spyre maintains a declarative test reuse framework that expresses all adaptations through YAML config.
The framework provides three controls:
- Parameter-level knobs: Subselect dtypes, exclude specific shapes, add coverage for a dtype of interest.
- Outcome bucketing: Assign each test to mandatory_success (must pass), xfail (expected to fail), xfail_strict (expected to fail, and an unexpected pass is treated as failure), or skip.
- Capability-driven inclusion: Declare supported ops/dtypes globally; enabling a new op automatically admits all tests waiting on it.
Concretely, this is what the declarative interface looks like in practice. A file entry sets a default policy for anything not explicitly listed, then lists tests in buckets, and an edits: block adapts individual cases without touching upstream source:
test_suite_config:
labels: [trunk]
files:
- path: ${TORCH_ROOT}/test/test_view_ops.py
unlisted_test_mode: skip # default policy for unlisted tests
tests:
# Basic tensor metadata manipulation which is fully supported on Spyre.
- names:
- TestOldViewOps::test_broadcast_tensors
- TestViewOps::test_contiguous_self
mode: mandatory_success
# bool/int64 fail only because test setup uses aten::random_.from.
- names:
- TestOldViewOps::test_broadcast_to
mode: mandatory_success
edits:
dtypes:
exclude:
- name: bool
- name: int64
global:
supported_dtypes: [{name: bfloat16}, {name: float16}, {name: float32}]
supported_ops:
- name: _scaled_mm
dtypes: [{name: bfloat16}]
Given these additional user-facing controls, the agentic workflow can be extended to generate the corresponding configuration files that can be consumed by the OOT test reuse framework. The low-level agent along with test repository memory can then capture not only test outcomes (e.g., mandatory_success), but also identify supported dtypes and other test parameters, enabling tests to be bucketed at the individual test-case level.
Four properties of the schema do most of the work:
- A declared default.
unlisted_test_modesets the outcome for anything unlisted, so a test added upstream tomorrow has a defined result instead of breaking CI the day it lands. - Four outcome buckets.
mandatory_success,xfail,xfail_strictandskip. The strict variant earns its place: an unexpected pass surfaces as a signal to promote the test, not silence. - Per-case adaptation.
edits:applies to individual dtypes, ops and modules, so an unsupported dtype costs one dtype rather than the whole test. - Capability declared once.
global.supported_opsandsupported_dtypesare declared per file, so enabling an operator in the backend is a one-line change that admits every test waiting on it.
How it works: The framework patches upstream’s @ops, @modules, and @dtypes decorators at collection time, emitting pytest marks (op__, dtype__). This means the same config drives both adaptation and selection: -m op__add runs exactly the tests exercising that operator.
Key guarantee: The upstream test tree stays pristine. No patches against PyTorch’s tests. A version bump becomes a config delta, not a merge conflict. The framework targets the generic privateuse1 device, so nothing is hardware-specific.
CRCR Integration into Existing Workflows

Figure 3: Workflow with CRCR Integration for Torch Spyre
Challenge 4: Consuming the dispatch: gate on the payload, resolve the SHA yourself
| Dispatch Field | Purpose |
|---|---|
| SHA | Identifies the exact upstream commit to validate. |
| PR Number | Associates downstream test results with a specific upstream PR. |
| Action | Specifies whether the PR was opened, updated (synchronize), or closed/merged. |
Base Branch (pull_request.base.ref) |
Identifies the target branch and enables filtering for main and release branches such as release/{version}. |
PR Label (pull_request.labels) |
Labels used with the pull request. Can be used to identify merged PRs through label “Merged”. |
Table 1: Key Dispatch Fields
Integrating CRCR starts with consuming the dispatch from PyTorch. The dispatch payload contains key pieces of information as shown in the Table 1. Dispatches can be filtered by consuming these pieces of information for various use cases. Filtering dispatches for pull requests merged to main requires action to be closed with a merge signal and base branch as main. The merge signal is indirect. PyTorch doesn’t use GitHub’s merge button that means PRs are closed, then PyTorchBot squash-merges and applies a “Merged” label. So detecting a landed PR requires checking for that label:
HAS_MERGED=$(echo "$PAYLOAD_JSON" | jq -r '.payload.pull_request.labels // [] | map(.name) | index("Merged") // ""')
Edge case: A race condition can cause the label to be applied after the dispatch, or a manual merge may skip the label entirely. Defensive handling (e.g., polling or a fallback heuristic) may be needed for critical workflows.
Apart from this dispatch-driven use case, there are two other use cases that may be of interest to OOT accelerators: testing against PyTorch nightly builds and testing against PyTorch releases.
Use Case C: Nightly Testing
CRCR doesn’t dispatch for nightlies, but HUD supports reporting nightly results. The workflow runs on a schedule and resolves the SHA directly as below:
NIGHTLY_SHA=$(curl -fsSL "https://api.github.com/repos/pytorch/pytorch/commits?sha=nightly&per_page=1" | jq -r '.[0].sha // empty')
if [ -z "${NIGHTLY_SHA}" ]; then
echo "::error::Could not resolve HEAD for pytorch/pytorch/nightly"
exit 1
fi
COMMIT_MSG=$(curl -fsSL "https://api.github.com/repos/pytorch/pytorch/commits/${NIGHTLY_SHA}" | jq -r '.commit.message')
SOURCE_SHA=$(echo "$COMMIT_MSG" | grep -oP '\(([a-f0-9]{40})\)' | tr -d '()')
if [ -z "${SOURCE_SHA}" ]; then
echo "::error::Could not extract source SHA from latest nightly commit"
exit 1
fi
echo "Resolved pytorch/pytorch@nightly -> source ${SOURCE_SHA}"
Use Case D: Release Testing
No dispatch is sent when a PyTorch release is published. Release testing is triggered manually via workflow_dispatch, passing the release branch (e.g., release/v2.14). SHA resolution follows the same pattern as nightly.
Note: HUD does not currently have a dedicated view for release test results.
Challenge 5: Enhancing workflow efficiency: test splitting, parallelization, and build-once-test-many
Thousands of tests need to run fast. Splitting happens at three levels:
- Feature-level: Group tests by area (Inductor, operators, eager, etc.)
- Duration-constrained: Subdivide each feature group to stay under a max job duration (e.g., 30 min)
- Balanced bucketing: Distribute individual tests across splits to avoid stragglers
Building this requires a one-pass timing run to measure per-test and per-feature durations. The result: parallel jobs that finish together, not one job holding up the rest.
Build-once, test-many: Build PyTorch and backend wheels once, publish as workflow artifacts, consume across all test splits. This avoids redundant builds and guarantees all jobs test the same artifacts.
Challenge 6: Enhancing workflow reliability

Figure 4: Retry Patterns for Workflow Resilience
The final state of the workflow in HUD should focus on surfacing regressions introduced by the evolving PyTorch and OOT accelerator codebases, rather than workflow stability failures such as transient infrastructure issues that may not be meaningful to the broader community.
While infrastructure and platform layers could be equipped with multiple reliability measures to improve workflow stability, developers can further strengthen the workflow by adopting additional reliability patterns, as shown in Table 2. The patterns listed here are not intended to be exhaustive, but provide a practical baseline for OOT accelerators to build reliable CRCR workflows.
| Reliability Dimension | Pattern | Impact from CRCR PoV | Implementation Detail |
|---|---|---|---|
| Fault tolerance | Parallel and isolated execution pattern | Non-faulty test jobs continue to complete and report their status even when another job fails. | GitHub Actions matrix jobs with pods/containers at the platform layer for execution isolation. |
| Fault tolerance | Continue on error pattern | Failures in non-critical paths do not cause the overall workflow to be reported as failed in HUD. | Use GitHub Actions continue-on-error for non-critical jobs or steps. |
| Resilience | Retry pattern | Transient failures can be recovered without surfacing persistent failures in HUD. | As shown in Figure 4, use workflow-level retries and finer-grained in-job retries based on test logs and failure heuristic. |
| Resilience | Logging and failure classification | Distinguishes actionable test regressions from transient infrastructure or platform failures not recoverable from retry pattern. | Verbose logging with summaries can help classify failures (that passed through retry pattern) and act for recovery. |
| Consistency | Build-once, use-many | Parallel jobs test the same PyTorch and OOT accelerator artifacts, reducing inconsistencies across splits. | Build wheels once, publish them as GitHub workflow artifacts, and consume the same artifacts across all test jobs. |
Table 2: Workflow Reliability Patterns
Challenge 7: CRCR Callbacks with Matrix Jobs
Github provides the ability to run a set of jobs based on a combination of variables and refers to that as a matrix strategy. Individual jobs as part of a matrix strategy are called legs.
The constraint: A matched in_progress/completed callback pair must originate from the same job. CRCR uses check_run_id to track retries, and each matrix leg gets its own check_run_id. This means callbacks cannot be sent from a parent job that spawns a matrix.
Two workarounds:
Option A: Per-leg callbacks. Each matrix job sends its own in_progress and completed. Simple, but can clutter HUD for large matrices and may not be ideal for dynamic jobs.
Option B: Dispatch + poll. The parent job dispatches a separate workflow containing the matrix, then polls for completion. This pattern also applies to heterogeneous CI such as triggering Jenkins or other platforms from a GitHub Actions parent job.
Challenge 8: Making a run reproducible weeks later
Debugging a failure from two weeks ago requires knowing exactly what ran. Capture this metadata as workflow artifacts:
- Dispatch payload (PR number, SHA, action)
- PyTorch and backend wheel versions (or commit SHAs)
- Runtime environment (Python, CUDA, OS, CPU arch)
- Per-test results (pass/fail/skip, duration)
Combined with build-once artifacts, this ensures any run can be reproduced or bisected weeks later.
Conclusion
CRCR handles distribution and reporting. The decisions it leaves to the downstream repo are the right ones to own: which dispatches matter, which tests are meaningful, and what “green” means on a given platform.
Owning those decisions meant keeping three moving targets in sync: the backend, PyTorch core, and the test suite. That’s where the engineering in this post went: an agentic test-selection pipeline, a declarative test-reuse framework, build-once-promote-many, and layered retries.
None of this is hardware-specific. The test-reuse framework targets the generic privateuse1 device. We intend to work with the PyTorch team to upstream it — most naturally alongside Torch OpenReg, the in-tree reference backend that accelerator authors already use to learn the extension points.
If any of these patterns would help your integration, we’d like to hear which ones matter most before shaping that proposal. Reach out by opening an issue at torch-spyre/torch-spyre.
Acknowledgements
We would like to thank the IBM Spyre Team for their support in this effort.