You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The table now includes complete LLVM b010a18d results for test_torchinductor_opinfo.py
(3,674 / 3,674) and the 20 MI308X-clean follow-up suites (4,985 / 4,985).
Other included suites retain their latest exact-node classifications from the August 10 merge. test_mps_basic.py and test_xpu_basic.py are permanently excluded;
their 50 nodes remain struck through for history and are omitted
from all totals and rates.
Merge rule: the latest terminal classification overrides by exact pytest node ID; source configuration and prior state are retained as provenance.
Suite Summary
Struck-through suites are excluded from the included total and will not be updated again. test_cudacodecache.py and test_extension_backend.py remain in the observed totals below but are pruned from follow-up rerun scope as pre-existing MI308X failures.
GitHub Markdown does not support row background colors, so the table uses a color-coded priority column and bold suite names.
Precedence is P1 > P2 > P3: 🔴 P1 = MI308X-clean suite with an MI450/gfx1250 timeout/miss; 🟡 P2 = remaining MI308X-clean suite with an MI450/gfx1250 failure/error; 🟣 P3 = remaining suite with more than 10 unresolved outcomes.
Priority
Suite
Total
Passed
Skipped
Xfailed
Failed
Error
Timed Out
Missed
test_alignment.py
12
12
0
0
0
0
0
0
test_analysis.py
28
28
0
0
0
0
0
0
test_aot_inductor.py
950
457
480
3
0
0
10
0
test_aot_inductor_arrayref.py
314
129
182
3
0
0
0
0
test_aot_inductor_custom_ops.py
35
35
0
0
0
0
0
0
🟡 P2
test_aot_inductor_package.py
88
59
25
0
4
0
0
0
test_async_compile.py
8
8
0
0
0
0
0
0
test_augmented_graph_helper.py
20
20
0
0
0
0
0
0
test_auto_chunker.py
9
9
0
0
0
0
0
0
test_auto_functionalize.py
39
38
1
0
0
0
0
0
🟡 P2
test_benchmark_fusion.py
16
11
4
0
1
0
0
0
test_benchmarking.py
14
8
0
6
0
0
0
0
test_best_config.py
1
1
0
0
0
0
0
0
test_binary_folding.py
6
5
1
0
0
0
0
0
test_block_analysis.py
10
10
0
0
0
0
0
0
test_cache.py
725
725
0
0
0
0
0
0
test_caching.py
212
212
0
0
0
0
0
0
test_ck_backend.py
34
0
34
0
0
0
0
0
test_codecache.py
257
197
59
1
0
0
0
0
test_codegen_triton.py
1
1
0
0
0
0
0
0
test_collective_autotuning.py
2
0
2
0
0
0
0
0
test_combo_kernels.py
77
72
5
0
0
0
0
0
test_compile.py
10
10
0
0
0
0
0
0
test_compile_subprocess.py
936
890
40
0
0
0
6
0
test_compile_worker.py
16
16
0
0
0
0
0
0
test_compiled_autograd.py
875
866
4
5
0
0
0
0
test_compiled_optimizers.py
682
679
3
0
0
0
0
0
test_config.py
14
14
0
0
0
0
0
0
test_control_deps.py
4
4
0
0
0
0
0
0
test_control_flow.py
741
739
2
0
0
0
0
0
test_cooperative_reductions.py
163
163
0
0
0
0
0
0
test_coordinate_descent_tuner.py
5
5
0
0
0
0
0
0
test_cpp_wrapper_hipify.py
3
3
0
0
0
0
0
0
test_cpu_repro.py
752
747
5
0
0
0
0
0
test_cuda_repro.py
98
90
3
1
0
0
4
0
test_cudacodecache.py
3
0
0
0
2
0
1
0
test_cudagraph_trees.py
189
181
8
0
0
0
0
0
test_cudagraph_trees_expandable_segments.py
158
152
6
0
0
0
0
0
🟡 P2
test_custom_lowering.py
6
4
1
0
1
0
0
0
test_custom_op_autotune.py
4
3
0
0
0
0
1
0
test_custom_partitioner_fn.py
1
1
0
0
0
0
0
0
test_custom_post_grad_passes.py
6
6
0
0
0
0
0
0
test_cutedsl_grouped_mm.py
24
0
24
0
0
0
0
0
test_cutedsl_template.py
13
0
13
0
0
0
0
0
test_cutlass_backend.py
181
0
181
0
0
0
0
0
test_cutlass_evt.py
8
0
8
0
0
0
0
0
test_debug_trace.py
3
3
0
0
0
0
0
0
test_decompose_mem_bound_mm.py
37
35
2
0
0
0
0
0
test_dependencies.py
5
5
0
0
0
0
0
0
🔴 P1
test_deterministic.py
32
26
0
0
0
0
6
0
test_device_assert.py
8
8
0
0
0
0
0
0
test_distributed_patterns.py
20
20
0
0
0
0
0
0
test_efficient_conv_bn_eval.py
2
2
0
0
0
0
0
0
test_exc_lowering_stack_trace.py
2
2
0
0
0
0
0
0
test_extension_backend.py
1
0
0
0
1
0
0
0
test_external_callables.py
3
3
0
0
0
0
0
0
🔴 P1
test_flex_attention.py
786
756
27
1
0
1
1
0
test_flex_decoding.py
557
554
1
2
0
0
0
0
test_flex_flash.py
210
6
204
0
0
0
0
0
test_foreach.py
607
581
26
0
0
0
0
0
🟡 P2
test_fp8.py
202
102
55
0
45
0
0
0
test_fused_attention.py
116
114
2
0
0
0
0
0
test_fusion_regions.py
6
6
0
0
0
0
0
0
test_fuzzer.py
11
10
1
0
0
0
0
0
test_fx_fusion.py
4
4
0
0
0
0
0
0
test_fxir_backend.py
75
75
0
0
0
0
0
0
test_gpu_cpp_wrapper.py
297
296
1
0
0
0
0
0
test_gpu_select_algorithm.py
58
58
0
0
0
0
0
0
test_graph_transform_observer.py
1
1
0
0
0
0
0
0
test_group_batch_fusion.py
13
11
2
0
0
0
0
0
test_halide.py
4
0
4
0
0
0
0
0
test_helion_kernels.py
2
0
2
0
0
0
0
0
test_indexing.py
22
22
0
0
0
0
0
0
test_inductor_annotations.py
2
2
0
0
0
0
0
0
test_inductor_freezing.py
48
45
2
0
0
0
1
0
test_inductor_scheduler.py
10
10
0
0
0
0
0
0
test_inductor_utils.py
2
2
0
0
0
0
0
0
test_inplace_padding.py
9
9
0
0
0
0
0
0
test_inplacing_pass.py
23
23
0
0
0
0
0
0
test_kernel_benchmark.py
19
19
0
0
0
0
0
0
test_kernel_optimization.py
1
1
0
0
0
0
0
0
test_lookup_table.py
37
33
4
0
0
0
0
0
test_loop_ordering.py
70
70
0
0
0
0
0
0
🟣 P3
test_max_autotune.py
472
191
173
0
0
0
107
1
test_max_autotune_blackwell.py
106
2
104
0
0
0
0
0
test_mem_estimation.py
4
4
0
0
0
0
0
0
test_memory.py
8
8
0
0
0
0
0
0
test_memory_planning.py
4
4
0
0
0
0
0
0
test_metrics.py
6
6
0
0
0
0
0
0
test_minifier.py
14
14
0
0
0
0
0
0
test_minifier_isolate.py
2
1
1
0
0
0
0
0
test_minifier_utils.py
3
3
0
0
0
0
0
0
🟣 P3
test_mix_order_reduction.py
485
248
219
0
0
0
17
1
test_mkldnn_pattern_matcher.py
20
16
4
0
0
0
0
0
test_mmdecomp.py
28
28
0
0
0
0
0
0
test_move_constructors_to_gpu.py
7
6
1
0
0
0
0
0
—
test_mps_basic.py
46
24
0
0
20
0
2
0
test_multi_kernel.py
19
17
2
0
0
0
0
0
test_native_matmul.py
14
14
0
0
0
0
0
0
test_needs_exact_strides.py
2
2
0
0
0
0
0
0
test_nv_universal_gemm.py
23
3
20
0
0
0
0
0
test_online_softmax.py
31
31
0
0
0
0
0
0
test_op_completeness.py
5
4
1
0
0
0
0
0
test_op_dtype_prop.py
581
581
0
0
0
0
0
0
test_ordered_set.py
401
386
15
0
0
0
0
0
test_pad_mm.py
19
18
1
0
0
0
0
0
test_padding.py
55
46
9
0
0
0
0
0
test_pattern_matcher.py
63
63
0
0
0
0
0
0
test_perf.py
66
66
0
0
0
0
0
0
test_profiler.py
8
8
0
0
0
0
0
0
test_provenance_tracing.py
16
16
0
0
0
0
0
0
test_quantization.py
2
2
0
0
0
0
0
0
test_remote_cache.py
3
3
0
0
0
0
0
0
test_scatter_optimization.py
8
8
0
0
0
0
0
0
test_segmented_tree.py
12
12
0
0
0
0
0
0
test_select_algorithm.py
27
26
1
0
0
0
0
0
test_selective_lowering.py
2
2
0
0
0
0
0
0
test_smoke.py
3
3
0
0
0
0
0
0
test_snode_runtime.py
22
22
0
0
0
0
0
0
test_split_cat_fx_aten_passes.py
5
5
0
0
0
0
0
0
test_split_cat_fx_passes.py
11
11
0
0
0
0
0
0
test_static_triton_launcher.py
17
17
0
0
0
0
0
0
test_subgraph_choice.py
2
2
0
0
0
0
0
0
test_template_heuristics_registry.py
7
7
0
0
0
0
0
0
test_torchbind.py
16
16
0
0
0
0
0
0
🔴 P1
test_torchinductor.py
1041
993
42
0
4
0
2
0
test_torchinductor_codegen_config_overrides.py
4
4
0
0
0
0
0
0
🔴 P1
test_torchinductor_codegen_dynamic_shapes.py
1864
1465
200
195
2
0
2
0
test_torchinductor_dynamic_shapes.py
1933
1808
112
3
0
0
0
10
🔴 P1
test_torchinductor_opinfo.py
3674
2930
691
42
6
0
5
0
🟣 P3
test_torchinductor_strided_blocks.py
304
86
194
0
0
0
2
22
test_triton_extension_backend.py
3
3
0
0
0
0
0
0
test_triton_helpers.py
2
2
0
0
0
0
0
0
🟡 P2
test_triton_heuristics.py
13
11
1
0
1
0
0
0
test_triton_kernels.py
372
328
44
0
0
0
0
0
test_triton_syntax.py
1
1
0
0
0
0
0
0
test_triton_wrapper.py
2
2
0
0
0
0
0
0
test_unbacked_symints.py
34
34
0
0
0
0
0
0
test_utils.py
11
11
0
0
0
0
0
0
—
test_xpu_basic.py
4
4
0
0
0
0
0
0
Included total
22912
19124
3259
262
67
1
165
34
Execution, Improvement, and Provenance Notes
Latest MI308X-clean 20-suite full rerun
Completed: 4,985 / 4,985 nodes across 20 complete suites in 2h 51m 59s.
Unresolved classifications fell from 186 to 70: 127 previous problems became non-failing, while 11 previously non-failing nodes became problematic.
11 / 20 suites now have no failed, error, timed-out, or missed outcomes: test_external_callables.py, test_kernel_benchmark.py, test_loop_ordering.py, test_multi_kernel.py, test_unbacked_symints.py, test_triton_kernels.py, test_gpu_cpp_wrapper.py, test_select_algorithm.py, test_native_matmul.py, test_pad_mm.py, test_pattern_matcher.py.
Execution used one full-file batch per suite, a 120s per-test timeout, a 12h per-file timeout, and no retries. The GPU-health monitor exited cleanly and the post-run GPU smoke test passed.
Earlier failed/error fix-validation rerun
These August 8-9 results remain the source for retained rows outside OpInfo and the 20 latest full-suite replacements.
Completed: 1,739 / 1,739 selected exact nodes.
Ordering: all 1,390test_torchinductor_opinfo.py nodes first, then the other 349 failed/error nodes.
The runner used one selected node per shard, a 120s per-test/file timeout, and no consecutive-outcome stop. Retries were 2 initially and changed to 0 after 1,410 completed nodes.
What improved in the earlier fix configuration
Note
These improvements were observed under a combined configuration: LLVM 850a2b1b, Triton PR #10721, HSA_HOTSWAP_ENABLE=1, and TRITON_HIP_USE_EXPERT_SCHEDULING=0. The run was not a single-variable
experiment, so it does not isolate the causal contribution of each change.
1,284 exact nodes transitioned from baseline FAILED/ERROR to PASSED/XFAILED: 1,281 passed and 3 xfailed.
Every definitive improvement was in test_torchinductor_opinfo.py. Its 1,390 prior errors became 1,281 passed, 3 xfailed, 6 timed out, and 100 missed.
All 60 selected OpInfo FFT nodes passed. This is broad batch confirmation of the COMGR translation path selected by HSA_HOTSWAP_ENABLE=1, beyond the earlier single-node check.
Complete baseline-failure transition: 26 remained failed, 166 timed out, and 62 missed. No prior failure reached a non-failing terminal state in this run.
The merged non-failing count increased by 1,284 nodes and the pass rate increased from 92.23% to 97.82% in the August 10 all-suite snapshot.
Exact result provenance
The current included table has 3,674 OpInfo outcomes and 4,985 outcomes from the 20-suite follow-up under LLVM b010a18d; the other 14,253 included outcomes are retained from the August 10 exact-node snapshot.
The 50 MPS/XPU nodes remain visible only as struck-through historical rows.
Before the OpInfo replacement and suite exclusions, the August 10 snapshot contained 14,897 original-run outcomes, 6,326 continuation outcomes, and 1,739 LLVM-850 fix-validation outcomes.
Exact prior-to-current transitions, source hashes, run configuration, and all improved node IDs are retained in failed_error_llvm850_expert0_hotswap1_20260808.provenance.json.
The retained source snapshot is merged_latest_results_20260809_llvm850_expert0_hotswap1.json; the replacement states are in opinfo_full_llvmb010_expert_default_hotswap1_sdma_unset_20260812.state.json and mi308x_clean20_full_llvmb010_expert_default_hotswap1_sdma_unset_20260812.state.json.
Unresolved outcomes and interpretation
The latest merge has 67 failed, 1 error, 165 timed out, and 34 missed nodes.
For retained rows that were not part of the two complete LLVM-b010 reruns, MISSED predominantly means pytest file collection exceeded the file timeout before the runner could identify the active exact node. It is not an assertion failure.
The only current ERROR is test_flex_attention.py::TestLearnableBiasesCUDA::test_flex_attention_with_dynamic_max_autotune_cuda, which aborted during the full rerun.
The triage below is intentionally narrower than the table: it includes only current timed-out/missed nodes from suites that were fully clean on MI308X.
That in-scope set has 16 timeouts and 0 misses across 5 suites.
The nine MI308X exact-node candidate suites contain 148 timeouts and 34 misses, but remain excluded until their 23 MI308X failing node IDs can be matched exactly.
Current failures and errors remain visible in the table but are outside this timeout/miss-focused triage.
Environment
GPU: AMD Radeon Graphics, gfx1250
Original baseline: PyTorch 2.11.0+rocm7.15.0a20260721, ROCm/HIP 7.15.0, Triton 3.8.0
Current LLVM-b010 environment for OpInfo and the 20-suite follow-up: HSA_HOTSWAP_ENABLE=1; SDMA, expert scheduling, and CoExec variables were unset, so both gfx1250 schedulers used their default-enabled behavior.
Transformers 5.15.0.dev0 was preinstalled in /opt/venv before the 20-suite run.
Current OpInfo mean runner time: 1.787s per node.
Failure Clusters and Current Triage
Note
Inclusion rule: triage only current TIMED OUT or MISSED nodes from suites
that had no failed, error, timed-out, or missed outcomes on MI308X.
Current failures/errors are intentionally outside this section.
These timeout/miss nodes are 🔴 Priority 1.
The in-scope comparison covers the 21 MI308X-clean suites (8,659 nodes). It contains 16 gfx1250 timeouts and 0 misses across 5 suites.
Pooling — 8 timeouts
Five OpInfo max_pool2d/max_pool3d cases timed out.
The CPU and GPU dynamic-shape avg_pool3d_backward2 cases timed out.
test_create_block_mask_recompile_persistent_reduction_oob_cuda timed out at 120s.
Miss triage: there are no current misses in the 21 MI308X-clean suites.
Exact in-scope timeout/miss nodes (16)
TIMED OUT — test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_inference_precision_amp
TIMED OUT — test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_amp
TIMED OUT — test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_bfloat16
TIMED OUT — test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_float16
TIMED OUT — test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_float32
TIMED OUT — test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_DistillGPT2_training_or_inference_training_precision_bfloat16
TIMED OUT — test/inductor/test_flex_attention.py::TestFlexAttentionCUDA::test_create_block_mask_recompile_persistent_reduction_oob_cuda
TIMED OUT — test/inductor/test_torchinductor.py::GPUTests::test_avg_pool3d_backward2_cuda
TIMED OUT — test/inductor/test_torchinductor.py::GPUTests::test_sort_stable_cuda
TIMED OUT — test/inductor/test_torchinductor_codegen_dynamic_shapes.py::DynamicShapesCodegenCpuTests::test_avg_pool3d_backward2_dynamic_shapes_cpu
TIMED OUT — test/inductor/test_torchinductor_codegen_dynamic_shapes.py::DynamicShapesCodegenGPUTests::test_avg_pool3d_backward2_dynamic_shapes_cuda
TIMED OUT — test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float16
TIMED OUT — test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float32
TIMED OUT — test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float64
TIMED OUT — test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool3d_cuda_float32
TIMED OUT — test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool3d_cuda_float64
Priority 2 queue
After Priority 1, the following MI308X-clean suites are 🟡 Priority 2 because MI450/gfx1250 has current failures/errors:
test_aot_inductor_package.py — failed 4, error 0
test_benchmark_fusion.py — failed 1, error 0
test_custom_lowering.py — failed 1, error 0
test_fp8.py — failed 45, error 0
test_triton_heuristics.py — failed 1, error 0
Priority 3 queue
The remaining rows with more than 10 total unresolved outcomes are 🟣 Priority 3:
Deferred from triage: the nine suites that contain 23 MI308X failures currently have 148 gfx1250 timeouts and 34 misses. They are excluded until the MI308X failing node IDs are matched exactly; suite-level counts cannot prove that an individual gfx1250 problem is unique.
Recommended next step: rerun these 16 exact nodes in fresh processes with a 600s per-test timeout. No broad full-suite rerun is needed for this scoped triage.
gfx1250 PyTorch Inductor Merged Outcome
Updated:
2026-08-13Important
The table now includes complete LLVM
b010a18dresults fortest_torchinductor_opinfo.py(
3,674 / 3,674) and the 20 MI308X-clean follow-up suites (4,985 / 4,985).Other included suites retain their latest exact-node classifications from the August 10 merge.
test_mps_basic.pyandtest_xpu_basic.pyare permanently excluded;their
50nodes remain struck through for history and are omittedfrom all totals and rates.
Overall Result
22,9625022,912PASSEDoutcomes:19,124(was17,633)2026-08-06) (passed + skipped + xfailed) / included nodes:21,150 / 22,912 = 92.31%2026-08-13) (passed + skipped + xfailed) / included nodes:22,645 / 22,912 = 98.83%267 / 22,912 = 1.17%(was1,762 / 22,912 = 7.69%)Suite Summary
Struck-through suites are excluded from the included total and will not be updated again.
test_cudacodecache.pyandtest_extension_backend.pyremain in the observed totals below but are pruned from follow-up rerun scope as pre-existing MI308X failures.GitHub Markdown does not support row background colors, so the table uses a color-coded priority column and bold suite names.
Precedence is
P1 > P2 > P3: 🔴 P1 = MI308X-clean suite with an MI450/gfx1250 timeout/miss; 🟡 P2 = remaining MI308X-clean suite with an MI450/gfx1250 failure/error; 🟣 P3 = remaining suite with more than10unresolved outcomes.test_alignment.pytest_analysis.pytest_aot_inductor.pytest_aot_inductor_arrayref.pytest_aot_inductor_custom_ops.pytest_aot_inductor_package.pytest_async_compile.pytest_augmented_graph_helper.pytest_auto_chunker.pytest_auto_functionalize.pytest_benchmark_fusion.pytest_benchmarking.pytest_best_config.pytest_binary_folding.pytest_block_analysis.pytest_cache.pytest_caching.pytest_ck_backend.pytest_codecache.pytest_codegen_triton.pytest_collective_autotuning.pytest_combo_kernels.pytest_compile.pytest_compile_subprocess.pytest_compile_worker.pytest_compiled_autograd.pytest_compiled_optimizers.pytest_config.pytest_control_deps.pytest_control_flow.pytest_cooperative_reductions.pytest_coordinate_descent_tuner.pytest_cpp_wrapper_hipify.pytest_cpu_repro.pytest_cuda_repro.pytest_cudacodecache.pytest_cudagraph_trees.pytest_cudagraph_trees_expandable_segments.pytest_custom_lowering.pytest_custom_op_autotune.pytest_custom_partitioner_fn.pytest_custom_post_grad_passes.pytest_cutedsl_grouped_mm.pytest_cutedsl_template.pytest_cutlass_backend.pytest_cutlass_evt.pytest_debug_trace.pytest_decompose_mem_bound_mm.pytest_dependencies.pytest_deterministic.pytest_device_assert.pytest_distributed_patterns.pytest_efficient_conv_bn_eval.pytest_exc_lowering_stack_trace.pytest_extension_backend.pytest_external_callables.pytest_flex_attention.pytest_flex_decoding.pytest_flex_flash.pytest_foreach.pytest_fp8.pytest_fused_attention.pytest_fusion_regions.pytest_fuzzer.pytest_fx_fusion.pytest_fxir_backend.pytest_gpu_cpp_wrapper.pytest_gpu_select_algorithm.pytest_graph_transform_observer.pytest_group_batch_fusion.pytest_halide.pytest_helion_kernels.pytest_indexing.pytest_inductor_annotations.pytest_inductor_freezing.pytest_inductor_scheduler.pytest_inductor_utils.pytest_inplace_padding.pytest_inplacing_pass.pytest_kernel_benchmark.pytest_kernel_optimization.pytest_lookup_table.pytest_loop_ordering.pytest_max_autotune.pytest_max_autotune_blackwell.pytest_mem_estimation.pytest_memory.pytest_memory_planning.pytest_metrics.pytest_minifier.pytest_minifier_isolate.pytest_minifier_utils.pytest_mix_order_reduction.pytest_mkldnn_pattern_matcher.pytest_mmdecomp.pytest_move_constructors_to_gpu.pytest_mps_basic.py46240020020test_multi_kernel.pytest_native_matmul.pytest_needs_exact_strides.pytest_nv_universal_gemm.pytest_online_softmax.pytest_op_completeness.pytest_op_dtype_prop.pytest_ordered_set.pytest_pad_mm.pytest_padding.pytest_pattern_matcher.pytest_perf.pytest_profiler.pytest_provenance_tracing.pytest_quantization.pytest_remote_cache.pytest_scatter_optimization.pytest_segmented_tree.pytest_select_algorithm.pytest_selective_lowering.pytest_smoke.pytest_snode_runtime.pytest_split_cat_fx_aten_passes.pytest_split_cat_fx_passes.pytest_static_triton_launcher.pytest_subgraph_choice.pytest_template_heuristics_registry.pytest_torchbind.pytest_torchinductor.pytest_torchinductor_codegen_config_overrides.pytest_torchinductor_codegen_dynamic_shapes.pytest_torchinductor_dynamic_shapes.pytest_torchinductor_opinfo.pytest_torchinductor_strided_blocks.pytest_triton_extension_backend.pytest_triton_helpers.pytest_triton_heuristics.pytest_triton_kernels.pytest_triton_syntax.pytest_triton_wrapper.pytest_unbacked_symints.pytest_utils.pytest_xpu_basic.py44000000Execution, Improvement, and Provenance Notes
Latest MI308X-clean 20-suite full rerun
4,985 / 4,985nodes across20complete suites in2h 51m 59s.4,315passed,404skipped,196xfailed,58failed,1error,11timed out, and0missed.4,915 / 4,985 = 98.60%.186to70:127previous problems became non-failing, while11previously non-failing nodes became problematic.11 / 20suites now have no failed, error, timed-out, or missed outcomes:test_external_callables.py,test_kernel_benchmark.py,test_loop_ordering.py,test_multi_kernel.py,test_unbacked_symints.py,test_triton_kernels.py,test_gpu_cpp_wrapper.py,test_select_algorithm.py,test_native_matmul.py,test_pad_mm.py,test_pattern_matcher.py.120sper-test timeout, a12hper-file timeout, and no retries. The GPU-health monitor exited cleanly and the post-run GPU smoke test passed.Earlier failed/error fix-validation rerun
These August 8-9 results remain the source for retained rows outside OpInfo and the 20 latest full-suite replacements.
1,739 / 1,739selected exact nodes.1,390test_torchinductor_opinfo.pynodes first, then the other349failed/error nodes.2026-08-08 01:03 CDTthrough2026-08-09 09:06 CDT(about32h 03m).1,281passed,3xfailed,26failed,253timed out, and176missed.120sper-test/file timeout, and no consecutive-outcome stop. Retries were2initially and changed to0after1,410completed nodes.What improved in the earlier fix configuration
Note
These improvements were observed under a combined configuration: LLVM
850a2b1b, Triton PR #10721,HSA_HOTSWAP_ENABLE=1, andTRITON_HIP_USE_EXPERT_SCHEDULING=0. The run was not a single-variableexperiment, so it does not isolate the causal contribution of each change.
1,284exact nodes transitioned from baselineFAILED/ERRORtoPASSED/XFAILED:1,281passed and3xfailed.test_torchinductor_opinfo.py. Its1,390prior errors became1,281passed,3xfailed,6timed out, and100missed.60selected OpInfo FFT nodes passed. This is broad batch confirmation of the COMGR translation path selected byHSA_HOTSWAP_ENABLE=1, beyond the earlier single-node check.1,281passed,3xfailed,87timed out, and114missed.26remained failed,166timed out, and62missed. No prior failure reached a non-failing terminal state in this run.1,284nodes and the pass rate increased from92.23%to97.82%in the August 10 all-suite snapshot.Exact result provenance
3,674OpInfo outcomes and4,985outcomes from the 20-suite follow-up under LLVMb010a18d; the other14,253included outcomes are retained from the August 10 exact-node snapshot.50MPS/XPU nodes remain visible only as struck-through historical rows.14,897original-run outcomes,6,326continuation outcomes, and1,739LLVM-850 fix-validation outcomes.failed_error_llvm850_expert0_hotswap1_20260808.provenance.json.merged_latest_results_20260809_llvm850_expert0_hotswap1.json; the replacement states are inopinfo_full_llvmb010_expert_default_hotswap1_sdma_unset_20260812.state.jsonandmi308x_clean20_full_llvmb010_expert_default_hotswap1_sdma_unset_20260812.state.json.Unresolved outcomes and interpretation
67failed,1error,165timed out, and34missed nodes.MISSEDpredominantly means pytest file collection exceeded the file timeout before the runner could identify the active exact node. It is not an assertion failure.ERRORistest_flex_attention.py::TestLearnableBiasesCUDA::test_flex_attention_with_dynamic_max_autotune_cuda, which aborted during the full rerun.16timeouts and0misses across5suites.148timeouts and34misses, but remain excluded until their23MI308X failing node IDs can be matched exactly.Environment
gfx12502.11.0+rocm7.15.0a20260721, ROCm/HIP7.15.0, Triton3.8.02.11.0+rocm10.1.0a20260803, ROCm/HIP7.15.26305, Triton3.8.02.11.0+rocm10.1.0a20260803(commit2c6687ca5c33), ROCm/HIP7.15.26305, Triton3.8.0(commit475d7bc8e67c, PR #10721), LLVM850a2b1b(PR #10739)HSA_HOTSWAP_ENABLE=1,TRITON_HIP_USE_EXPERT_SCHEDULING=02.11.0+rocm10.1.0a20260803(commit2c6687ca5c33), ROCm/HIP7.15.26305, Triton3.8.0(commite9c329c8cb49), LLVMb010a18dHSA_HOTSWAP_ENABLE=1; SDMA, expert scheduling, and CoExec variables were unset, so both gfx1250 schedulers used their default-enabled behavior.5.15.0.dev0was preinstalled in/opt/venvbefore the 20-suite run.1.787sper node.Failure Clusters and Current Triage
Note
Inclusion rule: triage only current
TIMED OUTorMISSEDnodes from suitesthat had no failed, error, timed-out, or missed outcomes on MI308X.
Current failures/errors are intentionally outside this section.
These timeout/miss nodes are 🔴 Priority 1.
The in-scope comparison covers the
21MI308X-clean suites (8,659nodes). It contains16gfx1250 timeouts and0misses across5suites.Pooling — 8 timeouts
max_pool2d/max_pool3dcases timed out.avg_pool3d_backward2cases timed out.test_torchinductor.py::GPUTests::test_avg_pool3d_backward2_cudatimed out.120slimit; none was problematic on MI308X.Deterministic model execution — 6 timeouts
5.15.0.dev0was already installed, so these are execution timeouts rather than package-setup errors.Stable sort — 1 timeout
test_torchinductor.py::GPUTests::test_sort_stable_cudatimed out at120s.FlexAttention block-mask recompilation — 1 timeout
test_create_block_mask_recompile_persistent_reduction_oob_cudatimed out at120s.Miss triage: there are no current misses in the 21 MI308X-clean suites.
Exact in-scope timeout/miss nodes (16)
TIMED OUT—test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_inference_precision_ampTIMED OUT—test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_ampTIMED OUT—test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_bfloat16TIMED OUT—test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_float16TIMED OUT—test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_BertForMaskedLM_training_or_inference_training_precision_float32TIMED OUT—test/inductor/test_deterministic.py::DeterministicTest::test_run2run_determinism_model_name_DistillGPT2_training_or_inference_training_precision_bfloat16TIMED OUT—test/inductor/test_flex_attention.py::TestFlexAttentionCUDA::test_create_block_mask_recompile_persistent_reduction_oob_cudaTIMED OUT—test/inductor/test_torchinductor.py::GPUTests::test_avg_pool3d_backward2_cudaTIMED OUT—test/inductor/test_torchinductor.py::GPUTests::test_sort_stable_cudaTIMED OUT—test/inductor/test_torchinductor_codegen_dynamic_shapes.py::DynamicShapesCodegenCpuTests::test_avg_pool3d_backward2_dynamic_shapes_cpuTIMED OUT—test/inductor/test_torchinductor_codegen_dynamic_shapes.py::DynamicShapesCodegenGPUTests::test_avg_pool3d_backward2_dynamic_shapes_cudaTIMED OUT—test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float16TIMED OUT—test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float32TIMED OUT—test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool2d_cuda_float64TIMED OUT—test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool3d_cuda_float32TIMED OUT—test/inductor/test_torchinductor_opinfo.py::TestInductorOpInfoCUDA::test_comprehensive_nn_functional_max_pool3d_cuda_float64Priority 2 queue
After Priority 1, the following MI308X-clean suites are 🟡 Priority 2 because MI450/gfx1250 has current failures/errors:
test_aot_inductor_package.py— failed4, error0test_benchmark_fusion.py— failed1, error0test_custom_lowering.py— failed1, error0test_fp8.py— failed45, error0test_triton_heuristics.py— failed1, error0Priority 3 queue
The remaining rows with more than
10total unresolved outcomes are 🟣 Priority 3:test_max_autotune.py— unresolved108test_mix_order_reduction.py— unresolved18test_torchinductor_strided_blocks.py— unresolved24Deferred from triage: the nine suites that contain
23MI308X failures currently have148gfx1250 timeouts and34misses. They are excluded until the MI308X failing node IDs are matched exactly; suite-level counts cannot prove that an individual gfx1250 problem is unique.Recommended next step: rerun these
16exact nodes in fresh processes with a600sper-test timeout. No broad full-suite rerun is needed for this scoped triage.The current OpInfo replacement is documented in this issue comment. The earlier failed/error rerun checkpoint is in this issue comment. Detailed reproduction steps and high-risk test notes are in this earlier comment.
The comments below are intermediate checkpoints and can be ignored.