Skip to content

PyTorch 2.11 Full Inductor Unit Tests #5

Description

@naromero77amd

gfx1250 PyTorch Inductor Merged Outcome

Updated: 2026-08-13

Important

The table now includes complete LLVM b010a18d results for test_torchinductor_opinfo.py
(3,674 / 3,674) and the 20 MI308X-clean follow-up suites (4,985 / 4,985).
Other included suites retain their latest exact-node classifications from the August 10 merge.
test_mps_basic.py and test_xpu_basic.py are permanently excluded;
their 50 nodes remain struck through for history and are omitted
from all totals and rates.

Overall Result

  • Discovered nodes: 22,962
  • Excluded nodes: 50
  • Included nodes used in totals and rates: 22,912
  • Explicit PASSED outcomes: 19,124 (was 17,633)
  • Pass rate — baseline (2026-08-06) (passed + skipped + xfailed) / included nodes: 21,150 / 22,912 = 92.31%
  • Pass rate — latest (2026-08-13) (passed + skipped + xfailed) / included nodes: 22,645 / 22,912 = 98.83%
  • Unresolved outcomes (failed + error + timed out + missed): 267 / 22,912 = 1.17% (was 1,762 / 22,912 = 7.69%)
  • Merge rule: the latest terminal classification overrides by exact pytest node ID; source configuration and prior state are retained as provenance.

Suite Summary

Struck-through suites are excluded from the included total and will not be updated again.
test_cudacodecache.py and test_extension_backend.py remain in the observed totals below but are pruned from follow-up rerun scope as pre-existing MI308X failures.

GitHub Markdown does not support row background colors, so the table uses a color-coded priority column and bold suite names.
Precedence is P1 > P2 > P3: 🔴 P1 = MI308X-clean suite with an MI450/gfx1250 timeout/miss; 🟡 P2 = remaining MI308X-clean suite with an MI450/gfx1250 failure/error; 🟣 P3 = remaining suite with more than 10 unresolved outcomes.

Priority Suite Total Passed Skipped Xfailed Failed Error Timed Out Missed
test_alignment.py 12 12 0 0 0 0 0 0
test_analysis.py 28 28 0 0 0 0 0 0
test_aot_inductor.py 950 457 480 3 0 0 10 0
test_aot_inductor_arrayref.py 314 129 182 3 0 0 0 0
test_aot_inductor_custom_ops.py 35 35 0 0 0 0 0 0
🟡 P2 test_aot_inductor_package.py 88 59 25 0 4 0 0 0
test_async_compile.py 8 8 0 0 0 0 0 0
test_augmented_graph_helper.py 20 20 0 0 0 0 0 0
test_auto_chunker.py 9 9 0 0 0 0 0 0
test_auto_functionalize.py 39 38 1 0 0 0 0 0
🟡 P2 test_benchmark_fusion.py 16 11 4 0 1 0 0 0
test_benchmarking.py 14 8 0 6 0 0 0 0
test_best_config.py 1 1 0 0 0 0 0 0
test_binary_folding.py 6 5 1 0 0 0 0 0
test_block_analysis.py 10 10 0 0 0 0 0 0
test_cache.py 725 725 0 0 0 0 0 0
test_caching.py 212 212 0 0 0 0 0 0
test_ck_backend.py 34 0 34 0 0 0 0 0
test_codecache.py 257 197 59 1 0 0 0 0
test_codegen_triton.py 1 1 0 0 0 0 0 0
test_collective_autotuning.py 2 0 2 0 0 0 0 0
test_combo_kernels.py 77 72 5 0 0 0 0 0
test_compile.py 10 10 0 0 0 0 0 0
test_compile_subprocess.py 936 890 40 0 0 0 6 0
test_compile_worker.py 16 16 0 0 0 0 0 0
test_compiled_autograd.py 875 866 4 5 0 0 0 0
test_compiled_optimizers.py 682 679 3 0 0 0 0 0
test_config.py 14 14 0 0 0 0 0 0
test_control_deps.py 4 4 0 0 0 0 0 0
test_control_flow.py 741 739 2 0 0 0 0 0
test_cooperative_reductions.py 163 163 0 0 0 0 0 0
test_coordinate_descent_tuner.py 5 5 0 0 0 0 0 0
test_cpp_wrapper_hipify.py 3 3 0 0 0 0 0 0
test_cpu_repro.py 752 747 5 0 0 0 0 0
test_cuda_repro.py 98 90 3 1 0 0 4 0
test_cudacodecache.py 3 0 0 0 2 0 1 0
test_cudagraph_trees.py 189 181 8 0 0 0 0 0
test_cudagraph_trees_expandable_segments.py 158 152 6 0 0 0 0 0
🟡 P2 test_custom_lowering.py 6 4 1 0 1 0 0 0
test_custom_op_autotune.py 4 3 0 0 0 0 1 0
test_custom_partitioner_fn.py 1 1 0 0 0 0 0 0
test_custom_post_grad_passes.py 6 6 0 0 0 0 0 0
test_cutedsl_grouped_mm.py 24 0 24 0 0 0 0 0
test_cutedsl_template.py 13 0 13 0 0 0 0 0
test_cutlass_backend.py 181 0 181 0 0 0 0 0
test_cutlass_evt.py 8 0 8 0 0 0 0 0
test_debug_trace.py 3 3 0 0 0 0 0 0
test_decompose_mem_bound_mm.py 37 35 2 0 0 0 0 0
test_dependencies.py 5 5 0 0 0 0 0 0
🔴 P1 test_deterministic.py 32 26 0 0 0 0 6 0
test_device_assert.py 8 8 0 0 0 0 0 0
test_distributed_patterns.py 20 20 0 0 0 0 0 0
test_efficient_conv_bn_eval.py 2 2 0 0 0 0 0 0
test_exc_lowering_stack_trace.py 2 2 0 0 0 0 0 0
test_extension_backend.py 1 0 0 0 1 0 0 0
test_external_callables.py 3 3 0 0 0 0 0 0
🔴 P1 test_flex_attention.py 786 756 27 1 0 1 1 0
test_flex_decoding.py 557 554 1 2 0 0 0 0
test_flex_flash.py 210 6 204 0 0 0 0 0
test_foreach.py 607 581 26 0 0 0 0 0
🟡 P2 test_fp8.py 202 102 55 0 45 0 0 0
test_fused_attention.py 116 114 2 0 0 0 0 0
test_fusion_regions.py 6 6 0 0 0 0 0 0
test_fuzzer.py 11 10 1 0 0 0 0 0
test_fx_fusion.py 4 4 0 0 0 0 0 0
test_fxir_backend.py 75 75 0 0 0 0 0 0
test_gpu_cpp_wrapper.py 297 296 1 0 0 0 0 0
test_gpu_select_algorithm.py 58 58 0 0 0 0 0 0
test_graph_transform_observer.py 1 1 0 0 0 0 0 0
test_group_batch_fusion.py 13 11 2 0 0 0 0 0
test_halide.py 4 0 4 0 0 0 0 0
test_helion_kernels.py 2 0 2 0 0 0 0 0
test_indexing.py 22 22 0 0 0 0 0 0
test_inductor_annotations.py 2 2 0 0 0 0 0 0
test_inductor_freezing.py 48 45 2 0 0 0 1 0
test_inductor_scheduler.py 10 10 0 0 0 0 0 0
test_inductor_utils.py 2 2 0 0 0 0 0 0
test_inplace_padding.py 9 9 0 0 0 0 0 0
test_inplacing_pass.py 23 23 0 0 0 0 0 0
test_kernel_benchmark.py 19 19 0 0 0 0 0 0
test_kernel_optimization.py 1 1 0 0 0 0 0 0
test_lookup_table.py 37 33 4 0 0 0 0 0
test_loop_ordering.py 70 70 0 0 0 0 0 0
🟣 P3 test_max_autotune.py 472 191 173 0 0 0 107 1
test_max_autotune_blackwell.py 106 2 104 0 0 0 0 0
test_mem_estimation.py 4 4 0 0 0 0 0 0
test_memory.py 8 8 0 0 0 0 0 0
test_memory_planning.py 4 4 0 0 0 0 0 0
test_metrics.py 6 6 0 0 0 0 0 0
test_minifier.py 14 14 0 0 0 0 0 0
test_minifier_isolate.py 2 1 1 0 0 0 0 0
test_minifier_utils.py 3 3 0 0 0 0 0 0
🟣 P3 test_mix_order_reduction.py 485 248 219 0 0 0 17 1
test_mkldnn_pattern_matcher.py 20 16 4 0 0 0 0 0
test_mmdecomp.py 28 28 0 0 0 0 0 0
test_move_constructors_to_gpu.py 7 6 1 0 0 0 0 0
test_mps_basic.py 46 24 0 0 20 0 2 0
test_multi_kernel.py 19 17 2 0 0 0 0 0
test_native_matmul.py 14 14 0 0 0 0 0 0
test_needs_exact_strides.py 2 2 0 0 0 0 0 0
test_nv_universal_gemm.py 23 3 20 0 0 0 0 0
test_online_softmax.py 31 31 0 0 0 0 0 0
test_op_completeness.py 5 4 1 0 0 0 0 0
test_op_dtype_prop.py 581 581 0 0 0 0 0 0
test_ordered_set.py 401 386 15 0 0 0 0 0
test_pad_mm.py 19 18 1 0 0 0 0 0
test_padding.py 55 46 9 0 0 0 0 0
test_pattern_matcher.py 63 63 0 0 0 0 0 0
test_perf.py 66 66 0 0 0 0 0 0
test_profiler.py 8 8 0 0 0 0 0 0
test_provenance_tracing.py 16 16 0 0 0 0 0 0
test_quantization.py 2 2 0 0 0 0 0 0
test_remote_cache.py 3 3 0 0 0 0 0 0
test_scatter_optimization.py 8 8 0 0 0 0 0 0
test_segmented_tree.py 12 12 0 0 0 0 0 0
test_select_algorithm.py 27 26 1 0 0 0 0 0
test_selective_lowering.py 2 2 0 0 0 0 0 0
test_smoke.py 3 3 0 0 0 0 0 0
test_snode_runtime.py 22 22 0 0 0 0 0 0
test_split_cat_fx_aten_passes.py 5 5 0 0 0 0 0 0
test_split_cat_fx_passes.py 11 11 0 0 0 0 0 0
test_static_triton_launcher.py 17 17 0 0 0 0 0 0
test_subgraph_choice.py 2 2 0 0 0 0 0 0
test_template_heuristics_registry.py 7 7 0 0 0 0 0 0
test_torchbind.py 16 16 0 0 0 0 0 0
🔴 P1 test_torchinductor.py 1041 993 42 0 4 0 2 0
test_torchinductor_codegen_config_overrides.py 4 4 0 0 0 0 0 0
🔴 P1 test_torchinductor_codegen_dynamic_shapes.py 1864 1465 200 195 2 0 2 0
test_torchinductor_dynamic_shapes.py 1933 1808 112 3 0 0 0 10
🔴 P1 test_torchinductor_opinfo.py 3674 2930 691 42 6 0 5 0
🟣 P3 test_torchinductor_strided_blocks.py 304 86 194 0 0 0 2 22
test_triton_extension_backend.py 3 3 0 0 0 0 0 0
test_triton_helpers.py 2 2 0 0 0 0 0 0
🟡 P2 test_triton_heuristics.py 13 11 1 0 1 0 0 0
test_triton_kernels.py 372 328 44 0 0 0 0 0
test_triton_syntax.py 1 1 0 0 0 0 0 0
test_triton_wrapper.py 2 2 0 0 0 0 0 0
test_unbacked_symints.py 34 34 0 0 0 0 0 0
test_utils.py 11 11 0 0 0 0 0 0
test_xpu_basic.py 4 4 0 0 0 0 0 0
Included total 22912 19124 3259 262 67 1 165 34

Execution, Improvement, and Provenance Notes

Latest MI308X-clean 20-suite full rerun

  • Completed: 4,985 / 4,985 nodes across 20 complete suites in 2h 51m 59s.
  • Outcomes: 4,315 passed, 404 skipped, 196 xfailed, 58 failed, 1 error, 11 timed out, and 0 missed.
  • Non-failing: 4,915 / 4,985 = 98.60%.
  • Unresolved classifications fell from 186 to 70: 127 previous problems became non-failing, while 11 previously non-failing nodes became problematic.
  • 11 / 20 suites now have no failed, error, timed-out, or missed outcomes: test_external_callables.py, test_kernel_benchmark.py, test_loop_ordering.py, test_multi_kernel.py, test_unbacked_symints.py, test_triton_kernels.py, test_gpu_cpp_wrapper.py, test_select_algorithm.py, test_native_matmul.py, test_pad_mm.py, test_pattern_matcher.py.
  • Execution used one full-file batch per suite, a 120s per-test timeout, a 12h per-file timeout, and no retries. The GPU-health monitor exited cleanly and the post-run GPU smoke test passed.

Earlier failed/error fix-validation rerun

These August 8-9 results remain the source for retained rows outside OpInfo and the 20 latest full-suite replacements.

  • Completed: 1,739 / 1,739 selected exact nodes.
  • Ordering: all 1,390 test_torchinductor_opinfo.py nodes first, then the other 349 failed/error nodes.
  • Window: 2026-08-08 01:03 CDT through 2026-08-09 09:06 CDT (about 32h 03m).
  • Outcomes: 1,281 passed, 3 xfailed, 26 failed, 253 timed out, and 176 missed.
  • The runner used one selected node per shard, a 120s per-test/file timeout, and no consecutive-outcome stop. Retries were 2 initially and changed to 0 after 1,410 completed nodes.

What improved in the earlier fix configuration

Note

These improvements were observed under a combined configuration: LLVM
850a2b1b, Triton PR #10721, HSA_HOTSWAP_ENABLE=1, and
TRITON_HIP_USE_EXPERT_SCHEDULING=0. The run was not a single-variable
experiment, so it does not isolate the causal contribution of each change.

  • 1,284 exact nodes transitioned from baseline FAILED/ERROR to PASSED/XFAILED: 1,281 passed and 3 xfailed.
  • Every definitive improvement was in test_torchinductor_opinfo.py. Its 1,390 prior errors became 1,281 passed, 3 xfailed, 6 timed out, and 100 missed.
  • All 60 selected OpInfo FFT nodes passed. This is broad batch confirmation of the COMGR translation path selected by HSA_HOTSWAP_ENABLE=1, beyond the earlier single-node check.
  • Complete baseline-error transition: 1,281 passed, 3 xfailed, 87 timed out, and 114 missed.
  • Complete baseline-failure transition: 26 remained failed, 166 timed out, and 62 missed. No prior failure reached a non-failing terminal state in this run.
  • The merged non-failing count increased by 1,284 nodes and the pass rate increased from 92.23% to 97.82% in the August 10 all-suite snapshot.

Exact result provenance

  • The current included table has 3,674 OpInfo outcomes and 4,985 outcomes from the 20-suite follow-up under LLVM b010a18d; the other 14,253 included outcomes are retained from the August 10 exact-node snapshot.
  • The 50 MPS/XPU nodes remain visible only as struck-through historical rows.
  • Before the OpInfo replacement and suite exclusions, the August 10 snapshot contained 14,897 original-run outcomes, 6,326 continuation outcomes, and 1,739 LLVM-850 fix-validation outcomes.
  • Exact prior-to-current transitions, source hashes, run configuration, and all improved node IDs are retained in failed_error_llvm850_expert0_hotswap1_20260808.provenance.json.
  • The retained source snapshot is merged_latest_results_20260809_llvm850_expert0_hotswap1.json; the replacement states are in opinfo_full_llvmb010_expert_default_hotswap1_sdma_unset_20260812.state.json and mi308x_clean20_full_llvmb010_expert_default_hotswap1_sdma_unset_20260812.state.json.

Unresolved outcomes and interpretation

  • The latest merge has 67 failed, 1 error, 165 timed out, and 34 missed nodes.
  • For retained rows that were not part of the two complete LLVM-b010 reruns, MISSED predominantly means pytest file collection exceeded the file timeout before the runner could identify the active exact node. It is not an assertion failure.
  • The only current ERROR is test_flex_attention.py::TestLearnableBiasesCUDA::test_flex_attention_with_dynamic_max_autotune_cuda, which aborted during the full rerun.
  • The triage below is intentionally narrower than the table: it includes only current timed-out/missed nodes from suites that were fully clean on MI308X.
  • That in-scope set has 16 timeouts and 0 misses across 5 suites.
  • The nine MI308X exact-node candidate suites contain 148 timeouts and 34 misses, but remain excluded until their 23 MI308X failing node IDs can be matched exactly.
  • Current failures and errors remain visible in the table but are outside this timeout/miss-focused triage.

Environment

  • GPU: AMD Radeon Graphics, gfx1250
  • Original baseline: PyTorch 2.11.0+rocm7.15.0a20260721, ROCm/HIP 7.15.0, Triton 3.8.0
  • Continuation: PyTorch 2.11.0+rocm10.1.0a20260803, ROCm/HIP 7.15.26305, Triton 3.8.0
  • Earlier fix-validation run: PyTorch 2.11.0+rocm10.1.0a20260803 (commit 2c6687ca5c33), ROCm/HIP 7.15.26305, Triton 3.8.0 (commit 475d7bc8e67c, PR #10721), LLVM 850a2b1b (PR #10739)
  • Earlier fix-validation environment: HSA_HOTSWAP_ENABLE=1, TRITON_HIP_USE_EXPERT_SCHEDULING=0
  • Current complete LLVM-b010 reruns: PyTorch 2.11.0+rocm10.1.0a20260803 (commit 2c6687ca5c33), ROCm/HIP 7.15.26305, Triton 3.8.0 (commit e9c329c8cb49), LLVM b010a18d
  • Current LLVM-b010 environment for OpInfo and the 20-suite follow-up: HSA_HOTSWAP_ENABLE=1; SDMA, expert scheduling, and CoExec variables were unset, so both gfx1250 schedulers used their default-enabled behavior.
  • Transformers 5.15.0.dev0 was preinstalled in /opt/venv before the 20-suite run.
  • Current OpInfo mean runner time: 1.787s per node.

Failure Clusters and Current Triage

Note

Inclusion rule: triage only current TIMED OUT or MISSED nodes from suites
that had no failed, error, timed-out, or missed outcomes on MI308X.
Current failures/errors are intentionally outside this section.
These timeout/miss nodes are 🔴 Priority 1.

The in-scope comparison covers the 21 MI308X-clean suites (8,659 nodes). It contains 16 gfx1250 timeouts and 0 misses across 5 suites.

  1. Pooling — 8 timeouts

    • Five OpInfo max_pool2d/max_pool3d cases timed out.
    • The CPU and GPU dynamic-shape avg_pool3d_backward2 cases timed out.
    • test_torchinductor.py::GPUTests::test_avg_pool3d_backward2_cuda timed out.
    • Every case reached the 120s limit; none was problematic on MI308X.
  2. Deterministic model execution — 6 timeouts

    • Five BertForMaskedLM training/inference precision cases and one DistillGPT2 training bfloat16 case timed out.
    • Transformers 5.15.0.dev0 was already installed, so these are execution timeouts rather than package-setup errors.
  3. Stable sort — 1 timeout

    • test_torchinductor.py::GPUTests::test_sort_stable_cuda timed out at 120s.
  4. FlexAttention block-mask recompilation — 1 timeout

    • test_create_block_mask_recompile_persistent_reduction_oob_cuda timed out at 120s.

Miss triage: there are no current misses in the 21 MI308X-clean suites.

Exact in-scope timeout/miss nodes (16)
  • TIMED OUTtest/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_inference_precision_amp
  • TIMED OUTtest/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_amp
  • TIMED OUTtest/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_bfloat16
  • TIMED OUTtest/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_float16
  • TIMED OUTtest/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_float32
  • TIMED OUTtest/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_DistillGPT2_training_or_inference_training_precision_bfloat16
  • TIMED OUTtest/inductor/test_flex_attention.py::TestFlexAttentionCUDA::test_create_block_mask_recompile_persistent_reduction_oob_cuda
  • TIMED OUTtest/inductor/test_torchinductor.py::GPUTests::test_avg_pool3d_backward2_cuda
  • TIMED OUTtest/inductor/test_torchinductor.py::GPUTests::test_sort_stable_cuda
  • TIMED OUTtest/inductor/test_torchinductor_codegen_dynamic_shapes.py::DynamicShapesCodegenCpuTests::test_avg_pool3d_backward2_dynamic_shapes_cpu
  • TIMED OUTtest/inductor/test_torchinductor_codegen_dynamic_shapes.py::DynamicShapesCodegenGPUTests::test_avg_pool3d_backward2_dynamic_shapes_cuda
  • TIMED OUTtest/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float16
  • TIMED OUTtest/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float32
  • TIMED OUTtest/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float64
  • TIMED OUTtest/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool3d_cuda_float32
  • TIMED OUTtest/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool3d_cuda_float64

Priority 2 queue

After Priority 1, the following MI308X-clean suites are 🟡 Priority 2 because MI450/gfx1250 has current failures/errors:

  • test_aot_inductor_package.py — failed 4, error 0
  • test_benchmark_fusion.py — failed 1, error 0
  • test_custom_lowering.py — failed 1, error 0
  • test_fp8.py — failed 45, error 0
  • test_triton_heuristics.py — failed 1, error 0

Priority 3 queue

The remaining rows with more than 10 total unresolved outcomes are 🟣 Priority 3:

  • test_max_autotune.py — unresolved 108
  • test_mix_order_reduction.py — unresolved 18
  • test_torchinductor_strided_blocks.py — unresolved 24

Deferred from triage: the nine suites that contain 23 MI308X failures currently have 148 gfx1250 timeouts and 34 misses. They are excluded until the MI308X failing node IDs are matched exactly; suite-level counts cannot prove that an individual gfx1250 problem is unique.

Recommended next step: rerun these 16 exact nodes in fresh processes with a 600s per-test timeout. No broad full-suite rerun is needed for this scoped triage.

The current OpInfo replacement is documented in this issue comment. The earlier failed/error rerun checkpoint is in this issue comment. Detailed reproduction steps and high-risk test notes are in this earlier comment.

The comments below are intermediate checkpoints and can be ignored.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions