Skip to content

[CI Monitor] Daily Report - 2026-06-23 #113

Description

@amd-bot

Daily Cross-Workflow Summary — 2026-06-23

Snapshot: 2026-06-23 21:01 UTC · Only completed runs/jobs counted in trends · Auto-updated every 30 min

TL;DR

🔴 RED · ~16 active clusters · 3 🆕 today (R204, R205, R206 — all aiter-caused) · ~12 carrying over · 3 review-ready fixes (#27757, #27141, #28889) + 1 merged (#28941)
👉 Today's ask: The Jun-22 AMD AITER Scout 27983814629 is now COMPLETE and — unlike yesterday's in-flight "zero aiter-caused" snapshot — its MI35x jobs reveal 3 A/B-confirmed aiter-caused accuracy regressions under candidate aiter a0320631: 🆕 R204 (GLM-5.1-FP8 DSA collapse), 🆕 R205 (DeepSeek-V3.2-MTP gsm8k), 🆕 R206 (gpt-oss-20b). Triage these before the next Thu scout — re-run each MI35x job with default aiter (AITER_COMMIT_OVERRIDE= empty) to confirm, then flag the aiter bump. Also chase the 3 review-ready fixes. ⚠️ Both nightly Jun-23 runs are IN-FLIGHT (only the early 2-GPU accuracy jobs done) — counts partial. Good news: both release-docker workflows are ✅ green today, and pr-test-amd-rocm720 produced a clean ✅ SUCCESS run (28027779773) — the concurrency churn is still present but a full run completed.

Workflow status

Workflow Latest completed run Trend (completed real failures) Δ vs yesterday
nightly-test-amd Jun-22 27975987522; Jun-23 28046906309 (IN-FLIGHT) 0 7 (+2 partial) 5·11·9·7·(IN-FLIGHT) pending (Jun-23 in-flight)
nightly-test-amd-rocm720 Jun-22 27975850577; Jun-23 28046795462 (IN-FLIGHT) 0 11 (+2 partial) 12·11·(IN-FLIGHT) pending (Jun-23 in-flight)
release-docker-amd-nightly Jun-23 28028570862 0 0·0·0 0
release-docker-amd-rocm720-nightly Jun-23 28027960865 0 0·0·0 0
amd-aiter-scout Jun-22 27983814629 (now COMPLETE) 0 ~13 real (3 aiter-caused · ~8 pre-existing · rest infra/cancel) scout completed; +3 aiter-caused surfaced
pr-test-amd Jun-23 28009346928 (schedule) + PR runs 0 ~4/run ~4–6 ~same
pr-test-amd-rocm720 Jun-23 ✅ 28027779773; 28046873327 concurrency-cancelled 0 (clean run) ~10·~10·0 much better (clean run landed)

Notes: (1) Both nightlies' Jun-23 runs (28046906309, 28046795462) are queued/IN-FLIGHT at snapshot — only the early 2-GPU accuracy jobs have reported; the 8-GPU jobs (Qwen3.5/Kimi/MXFP4) have not run yet, so their carry-over clusters are NOT yet confirmed for today. (2) The scout run is the Mon/Thu cron from Jun-22, now completed (conclusion cancelled overall, but real test failures ran first). No new scout today (next ≈ Thu Jun-25). (3) pr-test-amd-rocm720 still loses overlapping scheduled runs to cancel-in-progress, but run 28027779773 completed green — the in-flight CI fix #28987 (draft) targets this churn.

Failure clusters (active, deduplicated across all workflows)

R204 · 🆕 GLM-5.1-FP8 DSA accuracy collapse under candidate aiter — MI35x (scout, aiter-caused)

  • Status: NEW today (surfaced now scout completed). Near-zero accuracy, no crash, on candidate aiter a0320631.
  • Top hypothesis: [MEDIUM] aiter override regresses the GLM-5.1 FP8 sparse-MLA/DSA compute path; disconfirming: sister baseline-job log expired (404) — pass inferred from job success, not directly read.
  • In-flight fix: ❌ none (#28922 would move nightly off GLM-5.1→GLM-5.2; #28975 adds an opt-in gfx950 GLM5 kernel — neither a fix).
Workflow Job (shard) Test File Test Function Origin Error Log
amd-aiter-scout nightly-8-gpu-mi35x-glm51-rocm720 test/registered/amd/accuracy/mi35x/test_glm51_eval_mi35x.py test_glm51_accuracy aiter-caused near-zero accuracy (no crash) Log

R205 · 🆕 DeepSeek-V3.2-MTP gsm8k TP-divergence/accuracy — MI35x (scout, aiter-caused)

  • Status: NEW today. Sister (default aiter) passed at 0.965; scout (candidate aiter) hangs/diverges during gsm8k decode (DSA + EAGLE MTP).
  • Top hypothesis: [MEDIUM] aiter override changes a collective/MTP code path → TP-rank scheduling divergence → deadlock; disconfirming: hang vs accuracy drop are different symptoms (R205 vs R204) — may be two aiter regressions.
  • In-flight fix: ❌ none (aiter-override-only; sglang fix not expected).
Workflow Job (shard) Test File Test Function Origin Error Log
amd-aiter-scout nightly-accuracy-8-gpu-mi35x-deepseek-v32-mtp-rocm720 test_deepseek_v32_mtp_eval_mi35x.py TestDeepseekV32TPMTP.test_a_gsm8k aiter-caused collective deadlock during gsm8k decode Log

R206 · 🆕 gpt-oss-20b GSM8K below threshold (AITER MXFP4 fused-MoE, gfx950) — MI35x (scout, aiter-caused)

  • Status: NEW today. Sister (default aiter) passed at 0.520; scout below threshold.
  • Top hypothesis: [MEDIUM] candidate aiter MXFP4 fused-MoE kernel regresses gpt-oss-20b accuracy on gfx950; disconfirming: gpt-oss thresholds are tight — margin not stated, could be borderline.
  • In-flight fix: ❌ none.
Workflow Job (shard) Test File Test Function Origin Error Log
amd-aiter-scout nightly-accuracy-8-gpu-mi35x-rocm720 test/registered/amd/accuracy/mi35x/test_gpt_oss_eval_mi35x.py test_gpt_oss_accuracy aiter-caused gsm8k below threshold Log

R195 · Mamba extra_buffer needs CUDA/MUSA/NPU (FLA) rejected on ROCm (Qwen3.5) — 4 nightly jobs

  • Status: persistent since ~Jun-18/19 on all four Qwen3.5 nightly jobs (Jun-22 completed; Jun-23 8-GPU jobs not yet run). Top hypothesis: [HIGH] #28151 made autoextra_buffer, removing the ROCm fallback. In-flight fix: ✅ #27141 (open, ready, blocked) — unblock/land.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-8-gpu-qwen35 test_qwen35_eval_amd.py setUpClass extra_buffer needs CUDA/MUSA/NPU (FLA) Log
nightly-test-amd nightly-8-gpu-mi35x-qwen35 test_qwen35_eval_mi35x.py test_lm_eval same assert Log
nightly-test-amd-rocm720 nightly-8-gpu-qwen35-rocm720 test_qwen35_eval_amd.py setUpClass same assert Log
nightly-test-amd-rocm720 nightly-8-gpu-mi35x-qwen35-rocm720 test_qwen35_eval_mi35x.py test_lm_eval same assert Log

R2 · Mistral/Mixtral GSM8K below threshold (chat-eval) — both nightlies + scout (pre-existing)

  • Status: never-passed in window (≥Jun-13); Mistral/Mixtral score ~0.17–0.37 vs 0.46–0.69. Scout A/B confirms pre-existing (sglang) (Mistral-7B fails ~0.36 with default aiter too) — aiter exonerated. Still failing Jun-23 (nightly-accuracy-2-gpu, IN-FLIGHT run). In-flight fix: ✅ #27757 (open, ready, blocked) — unblock.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-accuracy-2-gpu (Jun-23) test_gsm8k_eval_amd.py test_gsm8k_all_models Mistral/Mixtral < thr (+ Qwen2-57B server-start timeout) Log
nightly-test-amd-rocm720 nightly-accuracy-2-gpu-rocm720 test_gsm8k_eval_amd.py test_gsm8k_all_models 5 models below thr Log
amd-aiter-scout nightly-accuracy-2-gpu (pre-existing) test_gsm8k_eval_amd.py test_gsm8k_all_models Mistral/Mixtral < thr Log

R1 · VLM MMMU accuracy below threshold (5 models, ~random) — both nightlies + scout (pre-existing)

  • Status: never-passed (≥Jun-13); same 5 models (Qwen2-VL, Qwen2.5-VL, Qwen3-VL-30B, MiniCPM-o, deepseek-vl2) ~0.26–0.29, bit-stable. Scout A/B confirms pre-existing (sglang). Still failing Jun-23. In-flight fix: ❌ none (stale WIP #26178 only touches perf tests).
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-accuracy-2-gpu-vlm (Jun-23) test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models 4 models < thr (+deepseek-vl2 download hang) Log
nightly-test-amd-rocm720 nightly-accuracy-2-gpu-vlm-rocm720 (Jun-23) test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models 5 models < thr Log
amd-aiter-scout nightly-accuracy-2-gpu-vlm (pre-existing) test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models 5 models < thr Log

R196 · VLM DP-encoder CUDA-IPC invalid device pointer crash/hang — nightlies + scout (pre-existing)

  • Status: ongoing since ~Jun-16; HIP invalid device pointer in cuda_ipc_transport_utils.py reconstruct path → exit -9; intermittent. Scout A/B confirms pre-existing (sglang). In-flight fix: ❌ none.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-4-gpu test_encoder_dp.py test_vlm_mmmu_benchmark mem fault → -9 Log
nightly-test-amd-rocm720 nightly-4-gpu-rocm720 test_encoder_dp.py test_vlm_mmmu_benchmark hipErrorInvalidDevicePointer Log
amd-aiter-scout nightly-4-gpu-rocm720 (pre-existing) test_encoder_dp.py test_vlm_mmmu_benchmark invalid device pointer Log

R192 · FLUX.2 modelopt-FP8 torch._scaled_mm HIPBLAS_STATUS_NOT_SUPPORTED — pr-test-amd multimodal-gen (now has fix)

  • Status: recurring on every pr-test-amd multimodal-gen-2gpu run (flux2). In-flight fix: ✅ #28889 (open, ready, blocked) — "dequant→bf16 fallback in apply_fp8_linear", exact-symptom fix; unblock/land. ← upgraded from ❌ yesterday.
Workflow Job (shard) Test File Test Function Error Log
pr-test-amd multimodal-gen-test-2-gpu-amd (1) diffusion (flux2) N/A HIPBLAS_STATUS_NOT_SUPPORTED Log

R202 · Mooncake RDMA transfer-engine segfault / KV-transfer timeout (PD-disagg) — pr-test-amd MI35x

  • Status: recurring on stage-b-test-large-8-gpu-mi35x-disaggregation-amd (RDMA ENOMEM / mooncake segfault / decode WaitingForInput 300s), cascades to GPU-occupied sibling tests. In-flight fix: ⚠️ #27012 (open, ready)Revert "[AMD] Fix stage-b-test-large-8-gpu-mi35x-disaggregation-amd", exact job-name match; review before duplicating.
Workflow Job (shard) Test File Test Function Error Log
pr-test-amd stage-b-test-large-8-gpu-mi35x-disaggregation-amd PD-disagg decode N/A mooncake RDMA segfault Log

R199 · Meta-Llama-3.1-70B-FP8 HIP OOM during decode CUDA-graph capture — scout (unclear)

  • Status: carried from Jun-22. Scout A/B: unclear — no FP8-70B sister baseline (sister ran only non-FP8 70B, which passed 0.910). One analysis labels it aiter-caused (LOW, likely flaky OOM); treat as unclear pending a default-aiter FP8-70B re-run. In-flight fix: ❌ none.
Workflow Job (shard) Test File Test Function Origin Error Log
amd-aiter-scout nightly-accuracy-2-gpu-rocm720 test_gsm8k_eval_amd.py test_gsm8k_all_models (70B-FP8) unclear HIP OOM at capture → -9 Log

R19 · Qwen3-235B-MXFP4 HIP stream-capture abort (hipErrorCapturedEvent) — MI35x, nightlies + scout

  • Status: never-passed (Jun 10–22); --tp4 --ep2 mxfp4-specific custom-all-reduce capture abort. Scout A/B pre-existing (sglang). In-flight fix: ⚠️ #23581 (default SGLANG_USE_AITER_AR false) — candidate.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-8-gpu-mi35x-qwen3-235b-mxfp4 test_qwen3_instruct_mxfp4.py setUpClass hipErrorCapturedEvent → -6 Log
amd-aiter-scout nightly-8-gpu-mi35x-qwen3-235b-mxfp4-rocm720 (pre-existing) test_qwen3_instruct_mxfp4.py setUpClass same Log

R194 · Kimi-K2.6 MoE w2 narrow IndexError (TP8) — rocm720 nightly

  • Status: narrow(start=shard_size*tp_rank) overruns a 64-wide w2 shard. In-flight fix: ⚠️ #28905 (open, DRAFT) — targets the w2 scale path; verify it also covers the w2 weight narrow.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd-rocm720 nightly-8-gpu-kimi-k26-rocm720 test_kimi_k26_eval_amd.py setUpClass IndexError narrow start out of range → -9 Log
nightly-test-amd nightly-8-gpu-kimi-k26 test_kimi_k26_eval_amd.py setUpClass same Log
Known/stable + resolved clusters (carrying over or fixed) · click to expand
ID Cluster Where (latest) In-flight fix
R200 DeepSeek-V4 compress-state KV HIP OOM (rocm720 dsv4-flash) rocm720 dsv4-flash #28941 MERGED (cee1caaf) — needs post-fix run to confirm
R197 anthropic SDK init — httpx 0.23.3 lacks socket_options (AMD image dep) pr-test-amd stage-b-1gpu-small (13) ❌ none — bump AMD image httpx≥0.25
R193 torch._inductor AssertionError in diffusion torch.compile (rocm720) quiet today (rocm720 ran green) ⚠️ #25969 (DRAFT)
R155 DeepSeek-V3.2 MI35x GSM8K catastrophic/below-thr (scout pre-existing) rocm720 dsv32-mi35x ⚠️ #25559 (WIP)
R6 Qwen3-30B-A3B unquantized MoE accuracy ~0.38 (MI35x) rocm720 accuracy-8gpu-mi35x ❌ none
R198 MiniMax-M2.7 GSM8K marginal 0.920 < 0.93 (borderline) rocm720 minimax-m27 ❌ none
R187 DeepSeek-R1-MXFP4 EAGLE/MTP accept-length low (MI35x) pr-test-amd stage-c-mi35x ⚠️ #28378
R203 Qwen2-57B/72B-FP8 server-start / decode capture hang (TP2) nightly-accuracy-2-gpu Jun-23 ❌ none
R201 Qwen3-Coder-Next basic server hang (aiter JIT, flaky 1/3) rocm720 accuracy-8gpu-mi35x ❌ none
R181 MORI build invalid-ELF-header → gtest discovery fail (infra) pr-test-amd stage-a-1gpu-small ❌ none (infra)
Ideogram-4-fp8 gated-repo 403 at diffusion setup pr-test-amd + scout (pre-existing) #28225 (draft)
RCCL allreduce SIGSEGV 8-GPU smoke test (test_rccl_multi_gpu.py) pr-test-amd stage-c-8gpu (0) ❌ none
Infrastructure / orchestration noise (not test failures) · click to expand
  • pr-test-amd-rocm720 concurrency churn: dozens of cancel-in-progress rows across runs (27975943807, 28046873327) — overlapping 30 17 daily + 0 */6 crons share github.ref. In-flight fix ⚠️ #28987 (draft). A clean run did complete today (28027779773 ✅).
  • Scout gate-cancellation cascade: call-gate / pr-gate (82820330340) cancelled → extra-a-test jobs cancelled (0 steps). No test ran.
  • HF weight-download stalls (xet): grok-1-W4A8KV8, GLM-5.1-MXFP4, Kimi-K2.6 MI35x, DeepSeek-R1 HiCache, DeepSeek-V3-KV-FP8, DeepSeek-V4-pro — across nightly + scout. These are infra/cache, not aiter (scout A/B explicitly exonerates aiter for the download hangs).
  • grok2-rocm720 VRAM not cleared (zombie KFD context) — scout 82820327688; runner-ops (node reboot), upstream ROCm/aiter#2061.
  • Diffusion 1-GPU infra: runner lost-communication, port-5555 contention, 45/90-min step timeouts on pr-test-amd multimodal-gen shards.
  • jit-kernel-unit-test-amd (scout 82820364884): C++ template compile error — already fixed by merged #27947; stale on the scout's older sglang SHA.

Workflow drill-down (per-workflow view)

amd-aiter-scout · Jun-22 [27983814629](https://github.com/sgl-project/sglang/actions/runs/27983814629) (now COMPLETE) · ~13 real · override [`a0320631`](https://github.com/ROCm/aiter/commit/a0320631f5567e8723916a5d3b24ae7d6c5a32c2)
Job (shard) Test File Test Function Cluster Origin Error
nightly-8-gpu-mi35x-glm51-rocm720 test_glm51_eval_mi35x.py test_glm51_accuracy R204🆕 aiter-caused accuracy collapse
nightly-accuracy-8-gpu-mi35x-deepseek-v32-mtp-rocm720 test_deepseek_v32_mtp_eval_mi35x.py test_a_gsm8k R205🆕 aiter-caused collective deadlock
nightly-accuracy-8-gpu-mi35x-rocm720 test_gpt_oss_eval_mi35x.py test_gpt_oss_accuracy R206🆕 aiter-caused gsm8k < thr
nightly-accuracy-2-gpu test_gsm8k_eval_amd.py test_gsm8k_all_models R2 / R199 pre-existing / unclear Mistral < thr + 70B-FP8 OOM
nightly-accuracy-2-gpu-rocm720 test_gsm8k_eval_amd.py test_gsm8k_all_models R2 / R199 pre-existing / unclear Mistral < thr + 70B-FP8 OOM
nightly-accuracy-2-gpu-vlm test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models R1 pre-existing 5 < thr
nightly-accuracy-2-gpu-vlm-rocm720 test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models R1 pre-existing 5 < thr
nightly-4-gpu-rocm720 test_encoder_dp.py test_vlm_mmmu_benchmark R196 pre-existing invalid device pointer
nightly-8-gpu-mi35x-qwen3-235b-mxfp4-rocm720 test_qwen3_instruct_mxfp4.py setUpClass R19 pre-existing hipErrorCapturedEvent
nightly-accuracy-8-gpu-mi35x-deepseek-v32-rocm720 test_deepseek_v32_eval_mi35x.py test_deepseek_v32_accuracy R155 pre-existing catastrophic ~0.01
2× minimax-m27 (amd/rocm720) test_minimax_m27_eval_amd.py accuracy R198 pre-existing 0.920 < 0.93
6+ download/gate/grok2 jobs various setUpClass/N/A infra unclear/infra HF xet stall / cancel / VRAM

Scout verdict: yesterday's in-flight snapshot said "zero aiter-caused"; the completed run shows 3 A/B-confirmed aiter-caused MI35x accuracy regressions (R204/R205/R206) under candidate aiter a0320631. All other real failures are pre-existing (sglang) (R1/R2/R19/R155/R196/R198) or unclear/infra (R199 + downloads).

nightly-test-amd · Jun-22 [27975987522](https://github.com/sgl-project/sglang/actions/runs/27975987522) (7) · Jun-23 [28046906309](https://github.com/sgl-project/sglang/actions/runs/28046906309) IN-FLIGHT (2 so far)
Job (shard) Test File Test Function Cluster Error
nightly-accuracy-2-gpu-vlm (Jun-23) test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models R1 4 < thr + deepseek-vl2 download
nightly-accuracy-2-gpu (Jun-23) test_gsm8k_eval_amd.py test_gsm8k_all_models R2 / R203 Mistral < thr + Qwen2-57B server-start
nightly-4-gpu test_encoder_dp.py test_vlm_mmmu_benchmark R196 IPC mem fault → -9
nightly-accuracy-2-gpu-vlm test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models R1 5 < thr
nightly-accuracy-2-gpu test_gsm8k_eval_amd.py test_gsm8k_all_models R2 Mistral/Mixtral+gemma
nightly-8-gpu-qwen35 + mi35x-qwen35 test_qwen35_eval_*.py setUpClass/test_lm_eval R195 extra_buffer assert
nightly-8-gpu-mi35x-qwen3-235b-mxfp4 test_qwen3_instruct_mxfp4.py setUpClass R19 HIP capture abort -6
nightly-8-gpu-kimi-k26 test_kimi_k26_eval_amd.py setUpClass R194 narrow IndexError
grok1-int4 / mi35x-kimi-k26 / mi35x-dsr1-hicache download setUpClass infra HF xet stall
nightly-test-amd-rocm720 · Jun-22 [27975850577](https://github.com/sgl-project/sglang/actions/runs/27975850577) (11) · Jun-23 [28046795462](https://github.com/sgl-project/sglang/actions/runs/28046795462) IN-FLIGHT (2 so far)
Job (shard) Test File Test Function Cluster Error
nightly-accuracy-2-gpu-vlm-rocm720 (Jun-23) test_vlms_mmmu_eval_amd.py test_mmmu_vlm_models R1 near-chance
nightly-accuracy-2-gpu-rocm720 (Jun-23) test_gsm8k_eval_amd.py test_gsm8k_all_models R2 Mistral + watchdog storms
nightly-4-gpu-rocm720 test_encoder_dp.py test_vlm_mmmu_benchmark R196 invalid device pointer
nightly-8-gpu-qwen35 + mi35x-qwen35 test_qwen35_eval_*.py setUpClass/test_lm_eval R195 extra_buffer assert
nightly-8-gpu-kimi-k26-rocm720 test_kimi_k26_eval_amd.py setUpClass R194 narrow IndexError
accuracy-8-gpu-mi35x-rocm720 test_qwen3_moe_eval_mi35x.py / ..._coder_next... accuracy R6 / R201 0.380 / basic hang
minimax-m27-rocm720 test_minimax_m27_eval_amd.py accuracy R198 0.920 < 0.93
dsv32-mi35x-rocm720 test_deepseek_v32_eval_mi35x.py accuracy R155 below thr / timeout
glm5-mxfp4 / glm51 / mi35x-kimi-k26 / dsr1-hicache / dsv4-pro download setUpClass infra HF xet stall
pr-test-amd · Jun-23 [28009346928](https://github.com/sgl-project/sglang/actions/runs/28009346928) (schedule) + [28048998470](https://github.com/sgl-project/sglang/actions/runs/28048998470) IN-FLIGHT
Job (shard) Test File Test Function Cluster Error
multimodal-gen-2gpu (1) diffusion (flux2) N/A R192 HIPBLAS_STATUS_NOT_SUPPORTED
stage-b-mi35x-disaggregation PD-disagg decode N/A R202 mooncake RDMA segfault
stage-c-8gpu (0) test_rccl_multi_gpu.py N/A RCCL SIGSEGV smoke test
stage-a-1gpu-small mori build N/A R181 invalid ELF header
stage-b-1gpu-small (13) / extra-a test_anthropic_server.py test_in_messages_system_role R197 httpx socket_options
multimodal-gen-1gpu / 2gpu diffusion N/A infra runner lost / port-5555 / timeout
pr-test-amd-rocm720 · Jun-23 ✅ [28027779773](https://github.com/sgl-project/sglang/actions/runs/28027779773) SUCCESS · [28046873327](https://github.com/sgl-project/sglang/actions/runs/28046873327) concurrency-cancelled

A clean scheduled run completed green today. The remaining "failures" in the daily issue are cancel-in-progress concurrency supersessions (overlapping crons share github.ref) — orchestration noise, not test failures. In-flight CI fix ⚠️ #28987 (draft) targets this. The substantive rocm720 cluster when runs do complete is R200 (DSV4 OOM), now ✅ fixed by merged #28941.

How this report is generated

  • Only status == "completed" runs/jobs counted in trends. Both nightly Jun-23 runs are IN-FLIGHT (only 2-GPU accuracy jobs reported) — their rows are partial; 8-GPU carry-over clusters (R195/R194/R19) are NOT yet confirmed for Jun-23. Scout 27983814629 is now COMPLETE. Both release-docker workflows ✅ green.
  • 🆕 NEW today: R204 (GLM-5.1-FP8 DSA collapse), R205 (DeepSeek-V3.2-MTP gsm8k), R206 (gpt-oss-20b) — all aiter-caused, surfaced because the scout completed. R200 resolved (#28941 merged); R192 gained a fix (#28889).
  • Scout Origins are A/B-derived: R204/R205/R206 aiter-caused; R1/R2/R19/R155/R196/R198 pre-existing (sglang); R199 + all download hangs unclear/infra. PR states verified live: #27757/#27141/#28889 open+ready+blocked; #28905/#28987 draft; #28941 merged; #27012 open+ready.
  • Confidence: FACT/HIGH/MEDIUM/LOW/SPECULATION. Bot does NOT assign Priority — engineers decide from cluster size + persistence + fix availability.

Generated by amd-bot · last updated 2026-06-23 21:01 UTC


Generated by amd-bot using Claude Code CLI (last updated: 2026-06-23 21:01 UTC)


CI Monitor — 2026-06-23

Repo: sgl-project/sglang

Monitored Workflows:

  • nightly-test-amd.yml
  • nightly-test-amd-rocm720.yml
  • release-docker-amd-nightly.yml
  • release-docker-amd-rocm720-nightly.yml
  • amd-aiter-scout.yml
  • pr-test-amd.yml
  • pr-test-amd-rocm720.yml

Per-workflow failure reports are appended as comments below; the Daily Cross-Workflow Summary is rendered above this section.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions