Repository navigation
Conversation
|
this makes sense. the current design isn't my favorite approach because it feels like it's not extensible. but I can't see a better way of doing right now |
|
actually on secodn thought, this isn't specific to just mori it will work for any |
|
@YukioZzz can you run |
临时携带 NVIDIA/srt-slurm#508 在 dc1e825d 的完整补丁,为显式 MoRI/MultiConnector 注入运行时 discovery 地址。依赖 pin 包含该 PR 后删除本提交;不更换依赖仓库,不重复携带已合入的 #504。
临时携带 NVIDIA/srt-slurm#508 在 dc1e825d 的完整补丁,为显式 MoRI/MultiConnector 注入运行时 discovery 地址。依赖 pin 包含该 PR 后删除本提交;不更换依赖仓库,不重复携带已合入的 #504。
临时携带 NVIDIA/srt-slurm#508 在 dc1e825d 的完整补丁,为显式 MoRI/MultiConnector 注入运行时 discovery 地址。依赖 pin 包含该 PR 后删除本提交;不更换依赖仓库,不重复携带已合入的 #504。
ok, working on it now~ |
dc1e825 to
949a892
Compare
|
Done, thanks! Rebased onto main and ran I also clarified the template-loading code and documentation, added regression coverage, and updated the example inventory. |
|
@YukioZzz can you remove |
I see, got it. It it not necessarily needed. |
949a892 to
51cee89
Compare
|
Done, removed The PR now contains only the discovery-template fix, its 16 regression tests, and documentation pointing to the existing MoRIIO discovery recipe. The implementation is unchanged from the previous revision.
Additional real-weight GPU integration at unchanged HEAD This integration evidence validates discovery binding and the routed request path only; it does not qualify shutdown behavior, accuracy, performance, long-duration stability or CPU-offload eviction/reload. The PR description includes the additional mock and main-compatibility validation results. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #508 +/- ##
=======================================
Coverage ? 83.40%
=======================================
Files ? 154
Lines ? 22158
Branches ? 0
=======================================
Hits ? 18481
Misses ? 3677
Partials ? 0 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
There was a problem hiding this comment.
Reviewed against REVIEW.md and the Design Rules. The fix goes through the existing resolver (kv_transfer_config), reads the table row (row.discovery, row.kv_connector) rather than a connector name, gets its ports from the allocator, and cites upstream at 387bcd39714c. Tests cover the new behavior well, and since no example changed, the launch snapshots show no diff.
0 blocking, 4 nits (inline).
Checks run locally at 51cee89: make lint passes. tests/test_discovery_connector_templates.py, test_vllm_router_frontend.py, test_vllm_connectors.py, test_vllm_port_allocation.py and test_launch_snapshots.py pass (88/88).
| ) | ||
| payload = row.transfer_config(mode) | ||
| payload["kv_connector_extra_config"] = self._discovery_extra_config(process, runtime) | ||
| template = self.get_config_for_mode(mode).get("kv-transfer-config") |
There was a problem hiding this comment.
Nit: This only reads the hyphenated key. _config_to_cli_args maps _ to -, so roles.prefill.args.kv_transfer_config: {...} skips the template and renders a second, unbound --kv-transfer-config. The flags are sorted, so the unbound one comes last and argparse keeps it, which is the bug this PR fixes, still reachable through the other spelling. I reproduced this with the test's _command helper: the output has two --kv-transfer-config flags, and the second has no host_ip/ports. The Dynamo path at L1750 already treats both spellings as explicit. Suggest popping either spelling here (and in build_worker_command), or rejecting the underscore one. No recipe uses that spelling today, hence a nit.
There was a problem hiding this comment.
Fixed. The discovery resolver now accepts either argument spelling, and the command builder removes the underscore alias before emitting one bound --kv-transfer-config. Supplying both spellings is rejected rather than relying on flag ordering. Regression coverage exercises both spellings with object/string templates and checks that only one flag is emitted, sibling connectors are preserved, and the input is unchanged.
| for child in children: | ||
| visit(child) | ||
|
|
||
| visit(payload) |
There was a problem hiding this comment.
Nit: A template with no or duplicate MoRIIOConnector, or the wrong kv_role, is a static recipe error. Right now it only raises in build_worker_command, after the allocation is granted, and srtctl dry-run doesn't catch it. You could split the structural half (parse, find the unique target, check the role) into a helper and call it from VLLMRouterFrontend.validate, so dry-run rejects it before submit. The topology-conflict check can stay here. The PR description already calls this out, so this could be a follow-up.
There was a problem hiding this comment.
Agreed that earlier structural validation would be useful. I am leaving that for a follow-up as suggested, keeping parsing, target matching, and role checks in the existing backend resolver in this PR. The documentation now explicitly says validation happens when the worker command is built; it does not claim that dry-run rejects malformed templates before submission.
| target["kv_role"] = expected_role | ||
| extra = target.setdefault("kv_connector_extra_config", {}) | ||
| for key, value in self._discovery_extra_config(process, runtime).items(): | ||
| if key in extra and extra[key] != value: |
There was a problem hiding this comment.
Nit: The comparison is type-strict. _discovery_extra_config emits ports as strings, so a YAML object template that sets http_port: 6101 (an int, matching the allocation) is rejected as a conflict. The docs discourage hard-coding these, so this is minor. Comparing str(extra[key]) != str(value) for the port keys, or noting in the error that values must be strings, would avoid a confusing failure.
There was a problem hiding this comment.
Fixed. Matching integer port values are normalized for comparison with the allocator's strings; the rendered config retains the allocator's string values. This is limited to port keys: mismatched ports, floats, and booleans remain rejected. Tests cover matching string/integer ports and conflicting values without mutating the input template.
| ``` | ||
|
|
||
| The pinned vLLM image must support the supplied connectors and their fields. | ||
| This changes command rendering only; it does not change Router or engine code. |
There was a problem hiding this comment.
Nit: "This changes command rendering only; it does not change Router or engine code." describes the PR, not the current behavior (docs/AGENTS.md: docs describe the implementation). Suggest dropping it and keeping the image-support sentence. Also, "rejected before worker launch" (L192) could say "when the worker command is built" to match where the check actually runs.
There was a problem hiding this comment.
Updated as suggested: removed the PR-scope sentence, retained the image-support requirement, and changed the validation timing to "when the worker command is built." The section also documents the accepted argument spellings and port value types.
Resolve explicit discovery templates through the existing backend resolver. Bind the unique matching connector to the allocated worker topology while preserving sibling connectors, transfer options and non-discovery precedence. Add command-rendering regression tests and document template usage with the existing discovery recipe. Signed-off-by: Yichao Zhu <Yichao.Zhu@amd.com>
51cee89 to
683c7f9
Compare
|
@cquil11 Addressed the alias handling, equivalent integer ports, and documentation comments in
The updated CI and copyright check require workflow approval. Could you approve the runs and take another look when convenient? Once checks pass, this is ready for merge from my side. |
Problem
With a discovery connector selected, an explicit
kv-transfer-configcurrently overrides the generated discovery configuration. For example, aMultiConnectorcontaining MoRIIO and CPU offload reaches vLLM without the MoRIIO child's worker-specific addresses and allocated ports.Changes
kv-transfer-configorkv_transfer_config, reject both together, and emit one bound CLI flag. Copy the template, locate exactly one connector matching the selected discovery table row, and bind its role and runtime topology.Before: the explicit MoRIIO child contains only recipe-supplied options such as
backend: rdma. After: the same child also receives the router address, worker address/HTTP port, allocated handshake/notify ports and existing READ-mode setting; the CPU-offload child stays unchanged.No new schema fields, router behavior, vLLM engine changes, or transport tuning are introduced. Connector selection still follows the existing table; topology binding reuses the existing allocator and discovery mapping.
The connector contract was checked against vLLM
387bcd39714c:moriio_common.pyconsumes the discovery fields, andmulti_connector.pyconstructs children fromkv_connector_extra_config.connectors. The documentation snippet requires an image that supports the supplied connectors and fields; this PR does not add support for other discovery protocols.Validation
Rebased onto main at
c2d437c10bcato include the launch-snapshot tooling.make snapshotsremoved the deleted example's snapshot;make snapshots-checkpassed. There is no net example, snapshot, or example-inventory diff against main.make checkat683c7f9c2951, including the alias/port review fixes: 3449 passed, 2 skipped, 6 deselected, including Ruff, blocking type checks and schema-documentation checks. These fixes were validated with CPU-only checks; no new cluster job was submitted.srtctl dry-run -f examples/vllm/vllm-router-moriio-disagg.yamlpassed. Schema documentation is unchanged and its freshness check passed.All 24 dedicated template regressions pass; the combined template/router/connector/allocator suite passes all 67 tests. Coverage includes both argument spellings, duplicate-alias rejection, JSON/object inputs, input immutability, nested templates, sibling preservation, matching string/integer ports, conflicting bindings, malformed chains, and unchanged non-discovery precedence. Binding and conflict checks run when each worker command is built, not as comprehensive pre-submit template validation.
Earlier documentation-to-launch checks at
51cee8904a0band its conflict-free merge with main at12089b15ebbfpassed for JSON object/string inputs with colocated 1P1D and 2P2D. The native mock orchestrator produced independent per-worker listener ports and preserved sibling connectors and transport options. These checks predate the alias/port review fixes.Earlier real-weight GPU integration at
51cee8904a0b: two-node 1P1D, TP8/DCP8 per role, prefillMultiConnectorwith MoRIIO and SimpleCPUOffload, and direct MoRIIO on decode. The unmodified native orchestrator supplied discovery bindings absent from the input template; the sibling connector and transport options were preserved. This GPU run predates the alias/port review fixes and was not repeated for this revision.Routed requests: 394/394 succeeded, with nonempty HTTP 200 completions, covering one cold request, four serial requests and a 120-second concurrency-8 window. Decode external-prefix-hit counters increased by 1,014,535 tokens. No serving-time RDMA flush, memory-registration failure or engine fatal error was observed.
The GPU evidence validates discovery binding and the routed request path only. It is not a qualification of shutdown behavior, accuracy, throughput, long-duration stability, CPU-offload eviction/reload, or other discovery protocols.
中文
修复显式
kv-transfer-config遮蔽自动 discovery 配置的问题:后端在模板副本中定位唯一匹配的 connector,注入当前 worker 的角色、地址和已分配端口,保留 CPU offload 等兄弟 connector 及原有传输选项,并拒绝歧义或冲突配置。支持连字符和下划线两种参数写法,只输出一个完成绑定的 CLI 参数,同时设置两种写法会报错;匹配的整数或字符串端口均可接受,输出仍使用 allocator 的值。按 review 要求移除独立 CPU-offload 示例及其生成快照,保留 YAML 文档片段、已有 discovery 示例引用及 24 项模板回归测试。未增加 schema 字段,未修改 router、vLLM 引擎或传输调优逻辑。绑定与冲突检查发生在 worker 命令生成时,不是完整的提交前模板验证。
本轮别名和端口修复后的
make check为 3449 项通过、2 项跳过、6 项 integration 测试未选中;定向测试 67 项通过。静态检查、类型检查、schema 文档检查、已有 discovery 示例 dry-run 和快照检查均通过。本轮仅进行 CPU 验证,没有提交新的集群任务。相对 PR 基线,示例、快照和示例清单均无变化。此前在
51cee8904a0b及其与 main12089b15ebbf的无冲突合并版本上,验证了文档 YAML 到原生 mock 启动命令的完整路径:JSON 对象/字符串两种输入,各覆盖同节点 1P1D、2P2D,确认 worker 监听端口独立且兄弟 connector 不变。该项补充验证早于本轮别名和端口修复。此前在
51cee8904a0b上完成双节点真实权重 1P1D 集成验证:P/D 均为 TP8/DCP8,P 使用 MoRIIO 与 SimpleCPUOffload 组成的 MultiConnector,D 使用直接 MoRIIO。原生编排正确注入 discovery 配置并保留兄弟 connector;冷请求、串行请求及 120 秒并发 8 共 394/394 成功,decode 外部命中计数增加 1,014,535 token,服务期间未观察到 RDMA flush、内存注册失败或引擎致命错误。本轮别名和端口修复未重复 GPU 实验。GPU 证据仅验证 discovery 绑定和路由请求路径,不构成退出行为、精度、吞吐、长时间稳定性、CPU offload 驱逐回载或其他 discovery 协议的验证。