Daily Cross-Workflow Summary — 2026-06-23
Snapshot: 2026-06-23 21:01 UTC · Only completed runs/jobs counted in trends · Auto-updated every 30 min
TL;DR
🔴 RED · ~16 active clusters · 3 🆕 today (R204, R205, R206 — all aiter-caused) · ~12 carrying over · 3 review-ready fixes (#27757, #27141, #28889) + 1 merged (#28941)
👉 Today's ask: The Jun-22 AMD AITER Scout 27983814629 is now COMPLETE and — unlike yesterday's in-flight "zero aiter-caused" snapshot — its MI35x jobs reveal 3 A/B-confirmed aiter-caused accuracy regressions under candidate aiter a0320631: 🆕 R204 (GLM-5.1-FP8 DSA collapse), 🆕 R205 (DeepSeek-V3.2-MTP gsm8k), 🆕 R206 (gpt-oss-20b). Triage these before the next Thu scout — re-run each MI35x job with default aiter (AITER_COMMIT_OVERRIDE= empty) to confirm, then flag the aiter bump. Also chase the 3 review-ready fixes. ⚠️ Both nightly Jun-23 runs are IN-FLIGHT (only the early 2-GPU accuracy jobs done) — counts partial. Good news: both release-docker workflows are ✅ green today, and pr-test-amd-rocm720 produced a clean ✅ SUCCESS run (28027779773) — the concurrency churn is still present but a full run completed.
Workflow status
| Workflow |
Latest completed run |
✅ |
❌ |
Trend (completed real failures) |
Δ vs yesterday |
| nightly-test-amd |
Jun-22 27975987522; Jun-23 28046906309 (IN-FLIGHT) |
0 |
7 (+2 partial) |
5·11·9·7·(IN-FLIGHT) |
pending (Jun-23 in-flight) |
| nightly-test-amd-rocm720 |
Jun-22 27975850577; Jun-23 28046795462 (IN-FLIGHT) |
0 |
11 (+2 partial) |
12·11·(IN-FLIGHT) |
pending (Jun-23 in-flight) |
| release-docker-amd-nightly |
Jun-23 28028570862 |
✅ |
0 |
0·0·0 |
0 |
| release-docker-amd-rocm720-nightly |
Jun-23 28027960865 |
✅ |
0 |
0·0·0 |
0 |
| amd-aiter-scout |
Jun-22 27983814629 (now COMPLETE) |
0 |
~13 real (3 aiter-caused · ~8 pre-existing · rest infra/cancel) |
— |
scout completed; +3 aiter-caused surfaced |
| pr-test-amd |
Jun-23 28009346928 (schedule) + PR runs |
0 |
~4/run |
~4–6 |
~same |
| pr-test-amd-rocm720 |
Jun-23 ✅ 28027779773; 28046873327 concurrency-cancelled |
✅ |
0 (clean run) |
~10·~10·0 |
much better (clean run landed) |
Notes: (1) Both nightlies' Jun-23 runs (28046906309, 28046795462) are queued/IN-FLIGHT at snapshot — only the early 2-GPU accuracy jobs have reported; the 8-GPU jobs (Qwen3.5/Kimi/MXFP4) have not run yet, so their carry-over clusters are NOT yet confirmed for today. (2) The scout run is the Mon/Thu cron from Jun-22, now completed (conclusion cancelled overall, but real test failures ran first). No new scout today (next ≈ Thu Jun-25). (3) pr-test-amd-rocm720 still loses overlapping scheduled runs to cancel-in-progress, but run 28027779773 completed green — the in-flight CI fix #28987 (draft) targets this churn.
Failure clusters (active, deduplicated across all workflows)
R204 · 🆕 GLM-5.1-FP8 DSA accuracy collapse under candidate aiter — MI35x (scout, aiter-caused)
- Status: NEW today (surfaced now scout completed). Near-zero accuracy, no crash, on candidate aiter
a0320631.
- Top hypothesis:
[MEDIUM] aiter override regresses the GLM-5.1 FP8 sparse-MLA/DSA compute path; disconfirming: sister baseline-job log expired (404) — pass inferred from job success, not directly read.
- In-flight fix: ❌ none (#28922 would move nightly off GLM-5.1→GLM-5.2; #28975 adds an opt-in gfx950 GLM5 kernel — neither a fix).
| Workflow |
Job (shard) |
Test File |
Test Function |
Origin |
Error |
Log |
| amd-aiter-scout |
nightly-8-gpu-mi35x-glm51-rocm720 |
test/registered/amd/accuracy/mi35x/test_glm51_eval_mi35x.py |
test_glm51_accuracy |
aiter-caused |
near-zero accuracy (no crash) |
Log |
R205 · 🆕 DeepSeek-V3.2-MTP gsm8k TP-divergence/accuracy — MI35x (scout, aiter-caused)
- Status: NEW today. Sister (default aiter) passed at 0.965; scout (candidate aiter) hangs/diverges during gsm8k decode (DSA + EAGLE MTP).
- Top hypothesis:
[MEDIUM] aiter override changes a collective/MTP code path → TP-rank scheduling divergence → deadlock; disconfirming: hang vs accuracy drop are different symptoms (R205 vs R204) — may be two aiter regressions.
- In-flight fix: ❌ none (aiter-override-only; sglang fix not expected).
R206 · 🆕 gpt-oss-20b GSM8K below threshold (AITER MXFP4 fused-MoE, gfx950) — MI35x (scout, aiter-caused)
- Status: NEW today. Sister (default aiter) passed at 0.520; scout below threshold.
- Top hypothesis:
[MEDIUM] candidate aiter MXFP4 fused-MoE kernel regresses gpt-oss-20b accuracy on gfx950; disconfirming: gpt-oss thresholds are tight — margin not stated, could be borderline.
- In-flight fix: ❌ none.
| Workflow |
Job (shard) |
Test File |
Test Function |
Origin |
Error |
Log |
| amd-aiter-scout |
nightly-accuracy-8-gpu-mi35x-rocm720 |
test/registered/amd/accuracy/mi35x/test_gpt_oss_eval_mi35x.py |
test_gpt_oss_accuracy |
aiter-caused |
gsm8k below threshold |
Log |
R195 · Mamba extra_buffer needs CUDA/MUSA/NPU (FLA) rejected on ROCm (Qwen3.5) — 4 nightly jobs
- Status: persistent since ~Jun-18/19 on all four Qwen3.5 nightly jobs (Jun-22 completed; Jun-23 8-GPU jobs not yet run). Top hypothesis:
[HIGH] #28151 made auto→extra_buffer, removing the ROCm fallback. In-flight fix: ✅ #27141 (open, ready, blocked) — unblock/land.
R2 · Mistral/Mixtral GSM8K below threshold (chat-eval) — both nightlies + scout (pre-existing)
- Status: never-passed in window (≥Jun-13); Mistral/Mixtral score ~0.17–0.37 vs 0.46–0.69. Scout A/B confirms
pre-existing (sglang) (Mistral-7B fails ~0.36 with default aiter too) — aiter exonerated. Still failing Jun-23 (nightly-accuracy-2-gpu, IN-FLIGHT run). In-flight fix: ✅ #27757 (open, ready, blocked) — unblock.
| Workflow |
Job (shard) |
Test File |
Test Function |
Error |
Log |
| nightly-test-amd |
nightly-accuracy-2-gpu (Jun-23) |
test_gsm8k_eval_amd.py |
test_gsm8k_all_models |
Mistral/Mixtral < thr (+ Qwen2-57B server-start timeout) |
Log |
| nightly-test-amd-rocm720 |
nightly-accuracy-2-gpu-rocm720 |
test_gsm8k_eval_amd.py |
test_gsm8k_all_models |
5 models below thr |
Log |
| amd-aiter-scout |
nightly-accuracy-2-gpu (pre-existing) |
test_gsm8k_eval_amd.py |
test_gsm8k_all_models |
Mistral/Mixtral < thr |
Log |
R1 · VLM MMMU accuracy below threshold (5 models, ~random) — both nightlies + scout (pre-existing)
- Status: never-passed (≥Jun-13); same 5 models (Qwen2-VL, Qwen2.5-VL, Qwen3-VL-30B, MiniCPM-o, deepseek-vl2) ~0.26–0.29, bit-stable. Scout A/B confirms
pre-existing (sglang). Still failing Jun-23. In-flight fix: ❌ none (stale WIP #26178 only touches perf tests).
R196 · VLM DP-encoder CUDA-IPC invalid device pointer crash/hang — nightlies + scout (pre-existing)
- Status: ongoing since ~Jun-16; HIP
invalid device pointer in cuda_ipc_transport_utils.py reconstruct path → exit -9; intermittent. Scout A/B confirms pre-existing (sglang). In-flight fix: ❌ none.
| Workflow |
Job (shard) |
Test File |
Test Function |
Error |
Log |
| nightly-test-amd |
nightly-4-gpu |
test_encoder_dp.py |
test_vlm_mmmu_benchmark |
mem fault → -9 |
Log |
| nightly-test-amd-rocm720 |
nightly-4-gpu-rocm720 |
test_encoder_dp.py |
test_vlm_mmmu_benchmark |
hipErrorInvalidDevicePointer |
Log |
| amd-aiter-scout |
nightly-4-gpu-rocm720 (pre-existing) |
test_encoder_dp.py |
test_vlm_mmmu_benchmark |
invalid device pointer |
Log |
R192 · FLUX.2 modelopt-FP8 torch._scaled_mm HIPBLAS_STATUS_NOT_SUPPORTED — pr-test-amd multimodal-gen (now has fix)
- Status: recurring on every pr-test-amd
multimodal-gen-2gpu run (flux2). In-flight fix: ✅ #28889 (open, ready, blocked) — "dequant→bf16 fallback in apply_fp8_linear", exact-symptom fix; unblock/land. ← upgraded from ❌ yesterday.
R202 · Mooncake RDMA transfer-engine segfault / KV-transfer timeout (PD-disagg) — pr-test-amd MI35x
- Status: recurring on
stage-b-test-large-8-gpu-mi35x-disaggregation-amd (RDMA ENOMEM / mooncake segfault / decode WaitingForInput 300s), cascades to GPU-occupied sibling tests. In-flight fix: ⚠️ #27012 (open, ready) — Revert "[AMD] Fix stage-b-test-large-8-gpu-mi35x-disaggregation-amd", exact job-name match; review before duplicating.
R199 · Meta-Llama-3.1-70B-FP8 HIP OOM during decode CUDA-graph capture — scout (unclear)
- Status: carried from Jun-22. Scout A/B:
unclear — no FP8-70B sister baseline (sister ran only non-FP8 70B, which passed 0.910). One analysis labels it aiter-caused (LOW, likely flaky OOM); treat as unclear pending a default-aiter FP8-70B re-run. In-flight fix: ❌ none.
| Workflow |
Job (shard) |
Test File |
Test Function |
Origin |
Error |
Log |
| amd-aiter-scout |
nightly-accuracy-2-gpu-rocm720 |
test_gsm8k_eval_amd.py |
test_gsm8k_all_models (70B-FP8) |
unclear |
HIP OOM at capture → -9 |
Log |
R19 · Qwen3-235B-MXFP4 HIP stream-capture abort (hipErrorCapturedEvent) — MI35x, nightlies + scout
- Status: never-passed (Jun 10–22);
--tp4 --ep2 mxfp4-specific custom-all-reduce capture abort. Scout A/B pre-existing (sglang). In-flight fix: ⚠️ #23581 (default SGLANG_USE_AITER_AR false) — candidate.
R194 · Kimi-K2.6 MoE w2 narrow IndexError (TP8) — rocm720 nightly
- Status:
narrow(start=shard_size*tp_rank) overruns a 64-wide w2 shard. In-flight fix: ⚠️ #28905 (open, DRAFT) — targets the w2 scale path; verify it also covers the w2 weight narrow.
Known/stable + resolved clusters (carrying over or fixed) · click to expand
| ID |
Cluster |
Where (latest) |
In-flight fix |
| R200 |
DeepSeek-V4 compress-state KV HIP OOM (rocm720 dsv4-flash) |
rocm720 dsv4-flash |
✅ #28941 MERGED (cee1caaf) — needs post-fix run to confirm |
| R197 |
anthropic SDK init — httpx 0.23.3 lacks socket_options (AMD image dep) |
pr-test-amd stage-b-1gpu-small (13) |
❌ none — bump AMD image httpx≥0.25 |
| R193 |
torch._inductor AssertionError in diffusion torch.compile (rocm720) |
quiet today (rocm720 ran green) |
⚠️ #25969 (DRAFT) |
| R155 |
DeepSeek-V3.2 MI35x GSM8K catastrophic/below-thr (scout pre-existing) |
rocm720 dsv32-mi35x |
⚠️ #25559 (WIP) |
| R6 |
Qwen3-30B-A3B unquantized MoE accuracy ~0.38 (MI35x) |
rocm720 accuracy-8gpu-mi35x |
❌ none |
| R198 |
MiniMax-M2.7 GSM8K marginal 0.920 < 0.93 (borderline) |
rocm720 minimax-m27 |
❌ none |
| R187 |
DeepSeek-R1-MXFP4 EAGLE/MTP accept-length low (MI35x) |
pr-test-amd stage-c-mi35x |
⚠️ #28378 |
| R203 |
Qwen2-57B/72B-FP8 server-start / decode capture hang (TP2) |
nightly-accuracy-2-gpu Jun-23 |
❌ none |
| R201 |
Qwen3-Coder-Next basic server hang (aiter JIT, flaky 1/3) |
rocm720 accuracy-8gpu-mi35x |
❌ none |
| R181 |
MORI build invalid-ELF-header → gtest discovery fail (infra) |
pr-test-amd stage-a-1gpu-small |
❌ none (infra) |
| — |
Ideogram-4-fp8 gated-repo 403 at diffusion setup |
pr-test-amd + scout (pre-existing) |
✅ #28225 (draft) |
| — |
RCCL allreduce SIGSEGV 8-GPU smoke test (test_rccl_multi_gpu.py) |
pr-test-amd stage-c-8gpu (0) |
❌ none |
Infrastructure / orchestration noise (not test failures) · click to expand
- pr-test-amd-rocm720 concurrency churn: dozens of
cancel-in-progress rows across runs (27975943807, 28046873327) — overlapping 30 17 daily + 0 */6 crons share github.ref. In-flight fix ⚠️ #28987 (draft). A clean run did complete today (28027779773 ✅).
- Scout gate-cancellation cascade:
call-gate / pr-gate (82820330340) cancelled → extra-a-test jobs cancelled (0 steps). No test ran.
- HF weight-download stalls (xet): grok-1-W4A8KV8, GLM-5.1-MXFP4, Kimi-K2.6 MI35x, DeepSeek-R1 HiCache, DeepSeek-V3-KV-FP8, DeepSeek-V4-pro — across nightly + scout. These are infra/cache, not aiter (scout A/B explicitly exonerates aiter for the download hangs).
- grok2-rocm720 VRAM not cleared (zombie KFD context) — scout 82820327688; runner-ops (node reboot), upstream ROCm/aiter#2061.
- Diffusion 1-GPU infra: runner lost-communication, port-5555 contention, 45/90-min step timeouts on pr-test-amd multimodal-gen shards.
- jit-kernel-unit-test-amd (scout 82820364884): C++ template compile error — already fixed by merged #27947; stale on the scout's older sglang SHA.
Workflow drill-down (per-workflow view)
amd-aiter-scout · Jun-22 [27983814629](https://github.com/sgl-project/sglang/actions/runs/27983814629) (now COMPLETE) · ~13 real · override [`a0320631`](https://github.com/ROCm/aiter/commit/a0320631f5567e8723916a5d3b24ae7d6c5a32c2)
Scout verdict: yesterday's in-flight snapshot said "zero aiter-caused"; the completed run shows 3 A/B-confirmed aiter-caused MI35x accuracy regressions (R204/R205/R206) under candidate aiter a0320631. All other real failures are pre-existing (sglang) (R1/R2/R19/R155/R196/R198) or unclear/infra (R199 + downloads).
nightly-test-amd · Jun-22 [27975987522](https://github.com/sgl-project/sglang/actions/runs/27975987522) (7) · Jun-23 [28046906309](https://github.com/sgl-project/sglang/actions/runs/28046906309) IN-FLIGHT (2 so far)
nightly-test-amd-rocm720 · Jun-22 [27975850577](https://github.com/sgl-project/sglang/actions/runs/27975850577) (11) · Jun-23 [28046795462](https://github.com/sgl-project/sglang/actions/runs/28046795462) IN-FLIGHT (2 so far)
pr-test-amd · Jun-23 [28009346928](https://github.com/sgl-project/sglang/actions/runs/28009346928) (schedule) + [28048998470](https://github.com/sgl-project/sglang/actions/runs/28048998470) IN-FLIGHT
| Job (shard) |
Test File |
Test Function |
Cluster |
Error |
| multimodal-gen-2gpu (1) |
diffusion (flux2) |
N/A |
R192 |
HIPBLAS_STATUS_NOT_SUPPORTED |
| stage-b-mi35x-disaggregation |
PD-disagg decode |
N/A |
R202 |
mooncake RDMA segfault |
| stage-c-8gpu (0) |
test_rccl_multi_gpu.py |
N/A |
RCCL |
SIGSEGV smoke test |
| stage-a-1gpu-small |
mori build |
N/A |
R181 |
invalid ELF header |
| stage-b-1gpu-small (13) / extra-a |
test_anthropic_server.py |
test_in_messages_system_role |
R197 |
httpx socket_options |
| multimodal-gen-1gpu / 2gpu |
diffusion |
N/A |
infra |
runner lost / port-5555 / timeout |
pr-test-amd-rocm720 · Jun-23 ✅ [28027779773](https://github.com/sgl-project/sglang/actions/runs/28027779773) SUCCESS · [28046873327](https://github.com/sgl-project/sglang/actions/runs/28046873327) concurrency-cancelled
A clean scheduled run completed green today. The remaining "failures" in the daily issue are cancel-in-progress concurrency supersessions (overlapping crons share github.ref) — orchestration noise, not test failures. In-flight CI fix ⚠️ #28987 (draft) targets this. The substantive rocm720 cluster when runs do complete is R200 (DSV4 OOM), now ✅ fixed by merged #28941.
How this report is generated
- Only
status == "completed" runs/jobs counted in trends. Both nightly Jun-23 runs are IN-FLIGHT (only 2-GPU accuracy jobs reported) — their rows are partial; 8-GPU carry-over clusters (R195/R194/R19) are NOT yet confirmed for Jun-23. Scout 27983814629 is now COMPLETE. Both release-docker workflows ✅ green.
- 🆕 NEW today: R204 (GLM-5.1-FP8 DSA collapse), R205 (DeepSeek-V3.2-MTP gsm8k), R206 (gpt-oss-20b) — all
aiter-caused, surfaced because the scout completed. R200 resolved (#28941 merged); R192 gained a fix (#28889).
- Scout Origins are A/B-derived: R204/R205/R206
aiter-caused; R1/R2/R19/R155/R196/R198 pre-existing (sglang); R199 + all download hangs unclear/infra. PR states verified live: #27757/#27141/#28889 open+ready+blocked; #28905/#28987 draft; #28941 merged; #27012 open+ready.
- Confidence:
FACT/HIGH/MEDIUM/LOW/SPECULATION. Bot does NOT assign Priority — engineers decide from cluster size + persistence + fix availability.
Generated by amd-bot · last updated 2026-06-23 21:01 UTC
Generated by amd-bot using Claude Code CLI (last updated: 2026-06-23 21:01 UTC)
CI Monitor — 2026-06-23
Repo: sgl-project/sglang
Monitored Workflows:
nightly-test-amd.yml
nightly-test-amd-rocm720.yml
release-docker-amd-nightly.yml
release-docker-amd-rocm720-nightly.yml
amd-aiter-scout.yml
pr-test-amd.yml
pr-test-amd-rocm720.yml
Per-workflow failure reports are appended as comments below; the Daily Cross-Workflow Summary is rendered above this section.
Daily Cross-Workflow Summary — 2026-06-23
Snapshot: 2026-06-23 21:01 UTC · Only completed runs/jobs counted in trends · Auto-updated every 30 min
TL;DR
🔴 RED · ~16 active clusters · 3 🆕 today (R204, R205, R206 — all aiter-caused) · ~12 carrying over · 3 review-ready fixes (#27757, #27141, #28889) + 1 merged (#28941)⚠️ Both nightly Jun-23 runs are IN-FLIGHT (only the early 2-GPU accuracy jobs done) — counts partial. Good news: both release-docker workflows are ✅ green today, and pr-test-amd-rocm720 produced a clean ✅ SUCCESS run (28027779773) — the concurrency churn is still present but a full run completed.
👉 Today's ask: The Jun-22 AMD AITER Scout 27983814629 is now COMPLETE and — unlike yesterday's in-flight "zero aiter-caused" snapshot — its MI35x jobs reveal 3 A/B-confirmed
aiter-causedaccuracy regressions under candidate aitera0320631: 🆕 R204 (GLM-5.1-FP8 DSA collapse), 🆕 R205 (DeepSeek-V3.2-MTP gsm8k), 🆕 R206 (gpt-oss-20b). Triage these before the next Thu scout — re-run each MI35x job with default aiter (AITER_COMMIT_OVERRIDE=empty) to confirm, then flag the aiter bump. Also chase the 3 review-ready fixes.Workflow status
cancelledFailure clusters (active, deduplicated across all workflows)
R204 · 🆕 GLM-5.1-FP8 DSA accuracy collapse under candidate aiter — MI35x (scout, aiter-caused)
a0320631.[MEDIUM]aiter override regresses the GLM-5.1 FP8 sparse-MLA/DSA compute path; disconfirming: sister baseline-job log expired (404) — pass inferred from jobsuccess, not directly read.test/registered/amd/accuracy/mi35x/test_glm51_eval_mi35x.pytest_glm51_accuracyR205 · 🆕 DeepSeek-V3.2-MTP gsm8k TP-divergence/accuracy — MI35x (scout, aiter-caused)
[MEDIUM]aiter override changes a collective/MTP code path → TP-rank scheduling divergence → deadlock; disconfirming: hang vs accuracy drop are different symptoms (R205 vs R204) — may be two aiter regressions.test_deepseek_v32_mtp_eval_mi35x.pyTestDeepseekV32TPMTP.test_a_gsm8kR206 · 🆕 gpt-oss-20b GSM8K below threshold (AITER MXFP4 fused-MoE, gfx950) — MI35x (scout, aiter-caused)
[MEDIUM]candidate aiter MXFP4 fused-MoE kernel regresses gpt-oss-20b accuracy on gfx950; disconfirming: gpt-oss thresholds are tight — margin not stated, could be borderline.test/registered/amd/accuracy/mi35x/test_gpt_oss_eval_mi35x.pytest_gpt_oss_accuracyR195 · Mamba
extra_buffer needs CUDA/MUSA/NPU (FLA)rejected on ROCm (Qwen3.5) — 4 nightly jobs[HIGH]#28151 madeauto→extra_buffer, removing the ROCm fallback. In-flight fix: ✅ #27141 (open, ready,blocked) — unblock/land.test_qwen35_eval_amd.pysetUpClassextra_buffer needs CUDA/MUSA/NPU (FLA)test_qwen35_eval_mi35x.pytest_lm_evaltest_qwen35_eval_amd.pysetUpClasstest_qwen35_eval_mi35x.pytest_lm_evalR2 · Mistral/Mixtral GSM8K below threshold (chat-eval) — both nightlies + scout (pre-existing)
pre-existing (sglang)(Mistral-7B fails ~0.36 with default aiter too) — aiter exonerated. Still failing Jun-23 (nightly-accuracy-2-gpu, IN-FLIGHT run). In-flight fix: ✅ #27757 (open, ready,blocked) — unblock.test_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_gsm8k_eval_amd.pytest_gsm8k_all_modelsR1 · VLM MMMU accuracy below threshold (5 models, ~random) — both nightlies + scout (pre-existing)
pre-existing (sglang). Still failing Jun-23. In-flight fix: ❌ none (stale WIP #26178 only touches perf tests).test_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelsR196 · VLM DP-encoder CUDA-IPC
invalid device pointercrash/hang — nightlies + scout (pre-existing)invalid device pointerincuda_ipc_transport_utils.pyreconstruct path → exit -9; intermittent. Scout A/B confirmspre-existing (sglang). In-flight fix: ❌ none.test_encoder_dp.pytest_vlm_mmmu_benchmarktest_encoder_dp.pytest_vlm_mmmu_benchmarkhipErrorInvalidDevicePointertest_encoder_dp.pytest_vlm_mmmu_benchmarkR192 · FLUX.2 modelopt-FP8
torch._scaled_mmHIPBLAS_STATUS_NOT_SUPPORTED — pr-test-amd multimodal-gen (now has fix)multimodal-gen-2gpurun (flux2). In-flight fix: ✅ #28889 (open, ready,blocked) — "dequant→bf16 fallback inapply_fp8_linear", exact-symptom fix; unblock/land. ← upgraded from ❌ yesterday.R202 · Mooncake RDMA transfer-engine segfault / KV-transfer timeout (PD-disagg) — pr-test-amd MI35x
stage-b-test-large-8-gpu-mi35x-disaggregation-amd(RDMA ENOMEM / mooncake segfault / decodeWaitingForInput300s), cascades to GPU-occupied sibling tests. In-flight fix:R199 · Meta-Llama-3.1-70B-FP8 HIP OOM during decode CUDA-graph capture — scout (unclear)
unclear— no FP8-70B sister baseline (sister ran only non-FP8 70B, which passed 0.910). One analysis labels itaiter-caused (LOW, likely flaky OOM); treat asunclearpending a default-aiter FP8-70B re-run. In-flight fix: ❌ none.test_gsm8k_eval_amd.pytest_gsm8k_all_models(70B-FP8)R19 · Qwen3-235B-MXFP4 HIP stream-capture abort (
hipErrorCapturedEvent) — MI35x, nightlies + scout--tp4 --ep2mxfp4-specific custom-all-reduce capture abort. Scout A/Bpre-existing (sglang). In-flight fix:SGLANG_USE_AITER_ARfalse) — candidate.test_qwen3_instruct_mxfp4.pysetUpClasshipErrorCapturedEvent→ -6test_qwen3_instruct_mxfp4.pysetUpClassR194 · Kimi-K2.6 MoE w2
narrowIndexError (TP8) — rocm720 nightlynarrow(start=shard_size*tp_rank)overruns a 64-wide w2 shard. In-flight fix:test_kimi_k26_eval_amd.pysetUpClassIndexError narrow start out of range→ -9test_kimi_k26_eval_amd.pysetUpClassKnown/stable + resolved clusters (carrying over or fixed) · click to expand
cee1caaf) — needs post-fix run to confirmhttpx 0.23.3lackssocket_options(AMD image dep)httpx≥0.25torch._inductorAssertionError in diffusion torch.compile (rocm720)basicserver hang (aiter JIT, flaky 1/3)test_rccl_multi_gpu.py)Infrastructure / orchestration noise (not test failures) · click to expand
cancel-in-progressrows across runs (27975943807, 28046873327) — overlapping30 17daily +0 */6crons sharegithub.ref. In-flight fixcall-gate / pr-gate(82820330340) cancelled → extra-a-test jobs cancelled (0 steps). No test ran.Workflow drill-down (per-workflow view)
amd-aiter-scout · Jun-22 [27983814629](https://github.com/sgl-project/sglang/actions/runs/27983814629) (now COMPLETE) · ~13 real · override [`a0320631`](https://github.com/ROCm/aiter/commit/a0320631f5567e8723916a5d3b24ae7d6c5a32c2)
test_glm51_eval_mi35x.pytest_glm51_accuracytest_deepseek_v32_mtp_eval_mi35x.pytest_a_gsm8ktest_gpt_oss_eval_mi35x.pytest_gpt_oss_accuracytest_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_encoder_dp.pytest_vlm_mmmu_benchmarktest_qwen3_instruct_mxfp4.pysetUpClasshipErrorCapturedEventtest_deepseek_v32_eval_mi35x.pytest_deepseek_v32_accuracytest_minimax_m27_eval_amd.pysetUpClass/N/AScout verdict: yesterday's in-flight snapshot said "zero aiter-caused"; the completed run shows 3 A/B-confirmed
aiter-causedMI35x accuracy regressions (R204/R205/R206) under candidate aitera0320631. All other real failures arepre-existing (sglang)(R1/R2/R19/R155/R196/R198) orunclear/infra (R199 + downloads).nightly-test-amd · Jun-22 [27975987522](https://github.com/sgl-project/sglang/actions/runs/27975987522) (7) · Jun-23 [28046906309](https://github.com/sgl-project/sglang/actions/runs/28046906309) IN-FLIGHT (2 so far)
test_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_encoder_dp.pytest_vlm_mmmu_benchmarktest_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_qwen35_eval_*.pysetUpClass/test_lm_evaltest_qwen3_instruct_mxfp4.pysetUpClasstest_kimi_k26_eval_amd.pysetUpClasssetUpClassnightly-test-amd-rocm720 · Jun-22 [27975850577](https://github.com/sgl-project/sglang/actions/runs/27975850577) (11) · Jun-23 [28046795462](https://github.com/sgl-project/sglang/actions/runs/28046795462) IN-FLIGHT (2 so far)
test_vlms_mmmu_eval_amd.pytest_mmmu_vlm_modelstest_gsm8k_eval_amd.pytest_gsm8k_all_modelstest_encoder_dp.pytest_vlm_mmmu_benchmarktest_qwen35_eval_*.pysetUpClass/test_lm_evaltest_kimi_k26_eval_amd.pysetUpClasstest_qwen3_moe_eval_mi35x.py/..._coder_next...test_minimax_m27_eval_amd.pytest_deepseek_v32_eval_mi35x.pysetUpClasspr-test-amd · Jun-23 [28009346928](https://github.com/sgl-project/sglang/actions/runs/28009346928) (schedule) + [28048998470](https://github.com/sgl-project/sglang/actions/runs/28048998470) IN-FLIGHT
test_rccl_multi_gpu.pytest_anthropic_server.pytest_in_messages_system_rolepr-test-amd-rocm720 · Jun-23 ✅ [28027779773](https://github.com/sgl-project/sglang/actions/runs/28027779773) SUCCESS · [28046873327](https://github.com/sgl-project/sglang/actions/runs/28046873327) concurrency-cancelled
A clean scheduled run completed green today. The remaining "failures" in the daily issue are⚠️ #28987 (draft) targets this. The substantive rocm720 cluster when runs do complete is R200 (DSV4 OOM), now ✅ fixed by merged #28941.
cancel-in-progressconcurrency supersessions (overlapping crons sharegithub.ref) — orchestration noise, not test failures. In-flight CI fixHow this report is generated
status == "completed"runs/jobs counted in trends. Both nightly Jun-23 runs are IN-FLIGHT (only 2-GPU accuracy jobs reported) — their rows are partial; 8-GPU carry-over clusters (R195/R194/R19) are NOT yet confirmed for Jun-23. Scout 27983814629 is now COMPLETE. Both release-docker workflows ✅ green.aiter-caused, surfaced because the scout completed. R200 resolved (#28941 merged); R192 gained a fix (#28889).aiter-caused; R1/R2/R19/R155/R196/R198pre-existing (sglang); R199 + all download hangsunclear/infra. PR states verified live: #27757/#27141/#28889 open+ready+blocked; #28905/#28987 draft; #28941 merged; #27012 open+ready.FACT/HIGH/MEDIUM/LOW/SPECULATION. Bot does NOT assign Priority — engineers decide from cluster size + persistence + fix availability.Generated by amd-bot · last updated 2026-06-23 21:01 UTC
Generated by amd-bot using Claude Code CLI (last updated: 2026-06-23 21:01 UTC)
CI Monitor — 2026-06-23
Repo: sgl-project/sglang
Monitored Workflows:
nightly-test-amd.ymlnightly-test-amd-rocm720.ymlrelease-docker-amd-nightly.ymlrelease-docker-amd-rocm720-nightly.ymlamd-aiter-scout.ymlpr-test-amd.ymlpr-test-amd-rocm720.ymlPer-workflow failure reports are appended as comments below; the Daily Cross-Workflow Summary is rendered above this section.