You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Snapshot: 2026-07-25 23:45 UTC · Only completed runs/jobs counted in trends · Auto-updated every 30 min
TL;DR
🟡 YELLOW · 2 🆕 today (R290 mi355x accuracy gate, R291 GHA platform error) · carry-over nightly families (R257 17d, R270 10d, R273, R262, R289) · R288 resolved (DCP hang did not recur) · in-flight fixes: R270→#31741 (blocked), R257→#31068/#32003, R290→#31727 (candidate)
👉 Today's ask: rerun the mi355x dsv4flash-fp8 leg 2–3× to separate R290 gate-flake from a real 3-day accuracy slide (0.917→0.914→0.905); rerun pr-test-amd to clear the transient R291 GHA platform error (it skipped all 22 AMD test jobs → no test signal on the latest run). Keep chasing #31741 (R270, 10th day) and #31068/#32003 (R257, 17th day). No fresh code regression.
Caution
pr-test-amd latest completed run 30157771919 (Jul-25 12:19) has NO test signal — a GitHub Actions internal platform error hit check-changes+gate jobs (R291), so all 22 downstream AMD test jobs were skipped, not tested. The real test-family failures come from the prior run 30136388246 (Jul-25 00:29). Rerun/update the branch before trusting a "green".
Caution
Jul-25 nightly-test-amd 30168349328 and nightly-test-amd-rocm720 30168262867 are IN-FLIGHT (queued, head 69a3c54c) → excluded from trends. Latest completed nightlies are the Jul-24 runs (30115195247 / 30115085856). pr-test-amd-rocm720 30168318210 also IN-FLIGHT; latest completed = Jul-24 30115153295.
Note
release-docker-amd-rocm720-nightly 30157811305 = startup_failure today (infra — workflow never provisioned, 0 jobs; was green Jul-18→24). release-docker-amd-nightly 30157965906 ✅ green. amd-aiter-scout: no run (cron Mon/Thu; today Sat).
Status: first RED for this workflow since Jul-18; config passed the prior 6 runs (Jul-19→24). This exact gate signature is intermittent: 3/13 completed runs (Jul-12, Jul-18, Jul-25). Marginal — 0.905 vs 0.91 gate.
Top hypothesis: [HIGH] the 0.91 GSM8K gate is too tight for this DSV4-Flash-FP8 DP8/EP8 config (real accuracy 0.905–0.914, ±0.005 sampling variance); runs straddle the bar. Disconfirming: the 3-day trend is monotonic (0.917→0.914→0.905), weakly more consistent with a small real drift than pure noise — a rerun is needed to decide.
In-flight fix: ❌ no threshold-retune PR; ⚠️ candidate #31727 (DSV4-Flash FP8 fused-RMS scale metadata on gfx950) could nudge numerics. Suggested triage: rerun the leg 2–3× at HEAD ebcb74ab; if it oscillates ≥0.91 it's gate flake, else bisect 99b29bf1..ebcb74ab (48 commits, FP8/MoE/DSV4 only).
Status: isolated to the single run30157771919 (Jul-25 12:22 UTC, 44 s window); prior run 12h earlier passed these jobs normally. Not a code regression.
Top hypothesis: [FACT] GitHub Actions platform internal error failed check-changes+gate/finish jobs before any step ran (0 steps, no runner, identical annotation on all 4 jobs). Because check-changes never emitted its change-filter outputs, all 22 downstream AMD test jobs were skipped → this run has zero test signal. Disconfirming: (none — directly observable).
In-flight fix: ❌ none (nothing to fix; platform-side). Suggested triage: simply rerun the run / update the branch; if it recurs, escalate to GitHub Actions infra.
Status: carry-over, 17th day (first Jul-09/10), never-passed on MI35x. Largest cluster: ~9 job-legs across both nightlies on the Jul-24 completed runs. Model-/workflow-independent; alternates on identical SHAs → environment/HW over a bisectable code regression.
Top hypothesis: [MEDIUM] a TP/MoE collective (aiter custom_all_reduce / IPC-meta gather) is captured into the decode HIP graph; the RCCL ProcessGroupNCCL watchdog's isCompleted()→hipEventQuery on that event is illegal mid-capture → hipErrorCapturedEvent → uncaught → SIGABRT (exit -6); compounded by MI35x node/driver hangs & runner-loss. Disconfirming: alternates pass/fail on identical SHAs; Jul-13 showed a distinct GPU Hang exit 134 → not a clean code regression.
In-flight fix: ⚠️#31068 (hung-GPU pre-flight recovery), #32003 (route scheduled jobs to MI300). Suggested triage: make MoE/TP all-reduce capture-safe on ROCm; chase #31068/#32003; validate MI35x pool health.
Status: carry-over, 10th day (first Jul-16), never-passed on AMD. 12 NameError subtests per kernel leg; identical on both Jul-24 completed runs.
Top hypothesis: [FACT]transfer_kv_mamba_* imports gated behind if _is_cuda: while the call site is reachable on HIP → NameError at runtime. Disconfirming: (none — directly observable).
In-flight fix: ✅ #31741 (open since Jul-20, mergeable_state: blocked). Suggested triage: rebase/merge #31741, then verify the JIT kernel actually compiles under ROCm, not just that the import resolves.
R289 · DeepSeek-V3.2 MTP GPU hang → scheduler watchdog timeout (MI35x, DSA backend) — nightly-amd · ⚠️ intermittent, no fix
Status: carry-over from Jul-24 (now 2nd data point), confirmed intermittent — passes ~4 of 6 recent nights; the identical signature already occurred Jul-22 on an earlier SHA, so no Jul23→24 commit introduced it. Not a fresh regression.
Top hypothesis: [LOW] intermittent MI35x (gfx950) GPU hang in the MTP draft-extend decode path — CPU parked on a HIP D2H sync (dsa_backend.py:793.item()); classic GPU-side stall. Disconfirming: no window commit touches dsa_backend.py/eagle_worker; the .item() is where CPU waits on the GPU, not the root cause.
In-flight fix: ❌ none (#32209 is a different model/config). Suggested triage: rerun 2–3× same SHA to confirm flakiness before any bisect; sync with the MI35x debug-branch owners.
Only status == "completed" runs/jobs counted in trends. Jul-25 nightly-amd/rocm720 and pr-rocm720 runs are queued (IN-FLIGHT) → excluded; latest completed nightlies are the Jul-24 runs. mi355x-disagg Jul-25 and pr-test-amd Jul-25 12:19 completed and are used directly.
NEW today: R290 mi355x GSM8K gate miss (workflow flipped green→red, marginal/flaky 3/13, candidate #31727); R291 GitHub Actions platform error (pr-test-amd, isolated single run, 22 test jobs skipped).
Resolved/not-reproduced: R288 DCP all-reduce hang — qwen35-triton-dcp-rocm720 passed in the Jul-24 completed run.
Daily Cross-Workflow Summary — 2026-07-25
Snapshot: 2026-07-25 23:45 UTC · Only completed runs/jobs counted in trends · Auto-updated every 30 min
TL;DR
🟡 YELLOW · 2 🆕 today (R290 mi355x accuracy gate, R291 GHA platform error) · carry-over nightly families (R257 17d, R270 10d, R273, R262, R289) · R288 resolved (DCP hang did not recur) · in-flight fixes: R270→#31741 (blocked), R257→#31068/#32003, R290→#31727 (candidate)
👉 Today's ask: rerun the mi355x
dsv4flash-fp8leg 2–3× to separate R290 gate-flake from a real 3-day accuracy slide (0.917→0.914→0.905); rerun pr-test-amd to clear the transient R291 GHA platform error (it skipped all 22 AMD test jobs → no test signal on the latest run). Keep chasing #31741 (R270, 10th day) and #31068/#32003 (R257, 17th day). No fresh code regression.Caution
pr-test-amd latest completed run 30157771919 (Jul-25 12:19) has NO test signal — a GitHub Actions internal platform error hit
check-changes+gate jobs (R291), so all 22 downstream AMD test jobs were skipped, not tested. The real test-family failures come from the prior run 30136388246 (Jul-25 00:29). Rerun/update the branch before trusting a "green".Caution
Jul-25 nightly-test-amd 30168349328 and nightly-test-amd-rocm720 30168262867 are IN-FLIGHT (
queued, head69a3c54c) → excluded from trends. Latest completed nightlies are the Jul-24 runs (30115195247 / 30115085856). pr-test-amd-rocm720 30168318210 also IN-FLIGHT; latest completed = Jul-24 30115153295.Note
release-docker-amd-rocm720-nightly 30157811305 =
startup_failuretoday (infra — workflow never provisioned, 0 jobs; was green Jul-18→24). release-docker-amd-nightly 30157965906 ✅ green. amd-aiter-scout: no run (cron Mon/Thu; today Sat).Workflow status
❌·❌·⊘·❌·❌·❌·❌❌·❌·⊘·❌·❌·❌·❌✅·✅·✅·✅·✅·✅·❌✅·✅·✅·✅·⚠️·✅·✅✅·✅·✅·✅·✅·✅·⚠️❌·❌·❌·❌·❌·❌·❌❌·❌·❌·❌·❌·❌·❌Failure clusters (deduplicated across all workflows)
R290 · 🆕 · GSM8K accuracy gate miss — DeepSeek-V4-Flash-FP8 1P1D (MI355X disagg) — mi355x-disagg ·⚠️ candidate #31727
[HIGH]the 0.91 GSM8K gate is too tight for this DSV4-Flash-FP8 DP8/EP8 config (real accuracy0.905–0.914, ±0.005 sampling variance); runs straddle the bar. Disconfirming: the 3-day trend is monotonic (0.917→0.914→0.905), weakly more consistent with a small real drift than pure noise — a rerun is needed to decide.ebcb74ab; if it oscillates ≥0.91 it's gate flake, else bisect99b29bf1..ebcb74ab(48 commits, FP8/MoE/DSV4 only).few_shot_gsm8k.py, 1319q)accuracy=0.905 threshold=0.91 → failing before sweepcancelled~7s (runnermia1-p01-g20infra) — separate, benchmark never ranR291 · 🆕 · GitHub Actions internal platform error (gate jobs, 0 steps) — pr-test-amd · ❌ transient infra
[FACT]GitHub Actions platform internal error failedcheck-changes+gate/finish jobs before any step ran (0 steps, no runner, identical annotation on all 4 jobs). Becausecheck-changesnever emitted its change-filter outputs, all 22 downstream AMD test jobs were skipped → this run has zero test signal. Disconfirming: (none — directly observable).GitHub Actions has encountered an internal error…R257 · MI35x (gfx950) GPU-hang / HW-exception family —⚠️ #31068/#32003
hipErrorCapturedEvent· GPU mem-fault · runner-loss — nightly-amd + nightly-rocm720 ·[MEDIUM]a TP/MoE collective (aitercustom_all_reduce/ IPC-meta gather) is captured into the decode HIP graph; the RCCLProcessGroupNCCLwatchdog'sisCompleted()→hipEventQueryon that event is illegal mid-capture →hipErrorCapturedEvent→ uncaught → SIGABRT (exit -6); compounded by MI35x node/driver hangs & runner-loss. Disconfirming: alternates pass/fail on identical SHAs; Jul-13 showed a distinctGPU Hang exit 134→ not a clean code regression.test_qwen3_instruct_mxfp4.pyTestQwen3Instruct2507MXFP4.setUpClasshipErrorCapturedEventduring decode graph capture → -6test_qwen3_instruct_mxfp4.pysetUpClasshipErrorCapturedEvent→ -6test_qwen3_instruct_mxfp4.pysetUpClasshipErrorCapturedEvent(aitercustom_all_reduce) → -6test_qwen3_moe_eval_mi35x.pytest_qwen3_moe_accuracytest_glm51_mxfp4_tp2_gsm8k_mi35x.pysetUpClassR270 · Mamba JIT transfer kernel
NameErroron ROCm (CUDA-only import gate) — nightly-amd + nightly-rocm720 kernel · ✅ #31741 (blocked)NameErrorsubtests per kernel leg; identical on both Jul-24 completed runs.[FACT]transfer_kv_mamba_*imports gated behindif _is_cuda:while the call site is reachable on HIP →NameErrorat runtime. Disconfirming: (none — directly observable).mergeable_state: blocked). Suggested triage: rebase/merge #31741, then verify the JIT kernel actually compiles under ROCm, not just that the import resolves.test/registered/kernels/ops/mamba/test_transfer_mamba.pyNameError: transfer_kv_mamba_lf_pf→ exit 255test_transfer_mamba.pyR289 · DeepSeek-V3.2 MTP GPU hang → scheduler watchdog timeout (MI35x, DSA backend) — nightly-amd ·⚠️ intermittent, no fix
[LOW]intermittent MI35x (gfx950) GPU hang in the MTP draft-extend decode path — CPU parked on a HIP D2H sync (dsa_backend.py:793.item()); classic GPU-side stall. Disconfirming: no window commit touchesdsa_backend.py/eagle_worker; the.item()is where CPU waits on the GPU, not the root cause.test/registered/amd/accuracy/mi35x/test_deepseek_v32_mtp_eval_mi35x.pyTestDeepseekV32TPMTP.test_a_gsm8kKeyError:'answer'→ exit 255Known stable / recurring clusters (carrying over, no new action today) · click to expand
test_gsm8k_eval_amd.py/test_*mxfp4*.py/test_vlms_mmmu_eval_amd.pyCLIPImageProcessor has no .tokenizer— server crashtest_vlms_mmmu_eval_amd.py(InternVL2_5-2B)wo_adequant)test_deepseek_v4_pro_*_mtp.pyHIPBLAS_STATUS_NOT_SUPPORTED,update_weights_from_disk400 shape-mismatchtest_server_1_gpu.py/test_server_2_gpu.pyKeyError:'answer'), Kimi-K2.5-MXFP4 loadmonitored_barrierSIGKILLtest_disaggregation_pp.py/test_kimi_k25_mxfp4_bcg_mi35x.pyModuleNotFoundError: msgpackon ROCm 7.2 image (diffusion extra not installed)test_msgpack_roundtrip_*test_*_mxfp4_*.pytest_bench_serving_*OUT_OF_RESOURCESWorkflow drill-down · click to expand
nightly-test-amd · latest completed 30115195247 (Jul-24) · 8 failures: R257 (qwen3-235b-mxfp4 ×2 legs, hipErrorCapturedEvent) · R270 (kernel) · R289 (dsv32-mtp watchdog hang) · R273 (acc-2gpu Mixtral+Qwen3-4B config, acc-2gpu-vlm MMMU, minimax-m25, nightly-4-gpu GLM-4.1V MMMU timeout). Jul-25 30168349328 IN-FLIGHT.
nightly-test-amd-rocm720 · latest completed 30115085856 (Jul-24) · 10 failures: R270 (kernel) · R257 (qwen3-235b-mxfp4, acc-8gpu mem-fault, glm51 runner-loss) · R262 (InternVL2_5) · R273 (dsv32, dsr1-mxfp4 MTP, minimax-m25, Mixtral/Qwen3-4B) · dsv4-pro-mtp launch timeout · dsv4-flash. R288 (DCP hang) did NOT recur —
qwen35-triton-dcpabsent from today's failures. Jul-25 30168262867 IN-FLIGHT.nightly-amd-mi355x-disagg · 30143908399 (Jul-25) · 3 failures = R290 🆕 (dsv4flash-fp8 gate 0.905<0.91) + collect-results rollup + 1 unrelated leg cancellation (runner
mia1-p01-g20infra). Was green Jul-19→24.pr-test-amd · latest completed 30157771919 (Jul-25 12:19) · R291 🆕 GHA platform error → 4 gate/finish jobs failed, 22 test jobs skipped (no signal). Prior run 30136388246 (Jul-25 00:29): diffusion + disagg families only.
pr-test-amd-rocm720 · latest completed 30115153295 (Jul-24) · 8 failures: diffusion family (FLUX.2 FP8/timeout/runner-loss) · msgpack ModuleNotFound · triton throughput gate · MXFP4 borderline · stage-c-8gpu HSA-OOM (DeepSeek-V3.2). Jul-25 30168318210 IN-FLIGHT.
How this report is generated
status == "completed"runs/jobs counted in trends. Jul-25 nightly-amd/rocm720 and pr-rocm720 runs arequeued(IN-FLIGHT) → excluded; latest completed nightlies are the Jul-24 runs. mi355x-disagg Jul-25 and pr-test-amd Jul-25 12:19 completed and are used directly.qwen35-triton-dcp-rocm720passed in the Jul-24 completed run.NameError(10th day, #31741 blocked), R273 accuracy near-misses, R262 InternVL, R289 DSV3.2-MTP (confirmed intermittent).startup_failure(Jul-25) — platform-side, 0 jobs, no cluster.FACT/HIGH/MEDIUM/LOW/SPECULATION. Bot does NOT assign Priority — engineers decide from cluster size + persistence + fix availability.Generated by amd-bot · last updated 2026-07-25 23:45 UTC
Generated by amd-bot using Claude Code CLI (last updated: 2026-07-25 23:45 UTC)
CI Monitor — 2026-07-25
Repo: sgl-project/sglang
Monitored Workflows:
nightly-test-amd.ymlnightly-test-amd-rocm720.ymlrelease-docker-amd-nightly.ymlrelease-docker-amd-rocm720-nightly.ymlnightly-amd-mi355x-disagg.ymlamd-aiter-scout.ymlpr-test-amd.ymlpr-test-amd-rocm720.ymlPer-workflow failure reports are appended as comments below; the Daily Cross-Workflow Summary is rendered above this section.