Skip to content

cuda.core: narrow the memcpy update release note and test the host-operand boundary - #3035

Open
Andy-Jost wants to merge 3 commits into
NVIDIA:mainfrom
Andy-Jost:ajost/memcpy-update-followup
Open

Andy-Jost wants to merge 3 commits into
NVIDIA:mainfrom
Andy-Jost:ajost/memcpy-update-followup

Conversation

@Andy-Jost

@Andy-Jost Andy-Jost commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Follow-up to #3029, which fixed #2649. The review there found that the release note overstated what Graph.update() accepts once a captured memcpy node's operand is replaced, and asked for a test of that boundary and a comment in the executable update path. This PR makes those three changes and nothing else.

Changes

  • Release note: a replaced operand stays unified, and a device-to-device replacement can also be applied with Graph.update(). A replacement with host memory always takes effect in a new instantiation; whether an executable update accepts it depends on the driver, and Linux drivers reject it.
  • test_memcpy_update_captured_node_host_operand replaces a captured node's source with a pinned host buffer and checks that a fresh instantiation copies from it. For Graph.update() and the executable view it accepts either outcome: a CUDAError (with PARAMETERS_CHANGED for Graph.update()), or a successful update whose launch copies from the host buffer. The first CI run showed why: Linux and Windows WDDM/MCDM rows reject the update, Windows TCC rows accept it.
  • ExecutableMemcpyNode.update() records in a comment that it reads the memory types from the definition node, which is correct only while MemcpyNode.update() keeps unified operands unified.
  • The three tests added by cuda.core: accept stream-captured memcpy nodes in MemcpyNode.update #3029 say why they keep the unused builder.

Related Work

#3029, #2649, and the review comment that requested these changes: #3029 (comment)

🤖 Generated with Claude Code

…erand boundary

Follow-up to NVIDIA#3029. The release note said Graph.update() applies any
replaced operand of a captured memcpy node; the driver accepts only a
device-to-device replacement in an executable update, and a replacement
with host memory takes effect in a new instantiation. A test pins that
boundary, and ExecutableMemcpyNode.update() records why it may read the
memory types from the definition node.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost Andy-Jost added this to the cuda.core 1.3.0 milestone Oct 6, 2026
@Andy-Jost Andy-Jost added documentation Improvements or additions to documentation P1 Medium priority - Should do cuda.core Everything related to the cuda.core module labels Oct 6, 2026
@Andy-Jost Andy-Jost self-assigned this Oct 6, 2026
@coderabbitai

coderabbitai Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuda-python/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: 42c2ef72-7445-44a2-926a-578a7164535c
📥 Commits

Reviewing files that changed from the base of the PR and between 1f96478 and 3f35fc5.

📒 Files selected for processing (2)
  • cuda_core/docs/source/release/1.3.0-notes.rst
  • cuda_core/tests/graph/test_graph_node_update.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • cuda_core/tests/graph/test_graph_node_update.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Documentation
    • Clarified that captured memcpy nodes preserve unified-address operand types during updates. Replacing an operand with host memory applies to a new graph instantiation; executable graph updates may be accepted or rejected depending on allocation type and driver support.
  • Tests
    • Expanded coverage for captured memcpy updates to verify operand and size changes, preservation of unchanged operands, and copied data after fresh instantiation. Added checks for driver-dependent handling of executable updates involving pinned host memory.

Walkthrough

The changes document executable memcpy update handling, clarify captured operand behavior in the release notes, and add a test for replacing a captured device operand with pinned host memory.

Changes

Captured memcpy updates

Layer / File(s) Summary
Document and test captured operand updates
cuda_core/cuda/core/graph/_subclasses.pyx, cuda_core/tests/graph/test_graph_node_update.py, cuda_core/docs/source/release/1.3.0-notes.rst
Comments describe executable update handling of memory types. The test checks a fresh instantiation with a pinned host operand and allows an existing executable update to succeed or fail according to driver support. The release note distinguishes device and host operand replacements.

Assessment against linked issues

Objective Addressed Explanation
Allow MemcpyNode.update() on stream-captured memcpy nodes to replace the copy size [#2649] ❌ The provided changes do not modify or test copy-size updates.

Suggested reviewers: juenglin

Priority: ➖ Normal

Change: Other

Merge Risk: 🔵 Low · up to 3f35f

The new test leaves a focused gap in coverage for direct executable-node updates from device to host memory. The PR is otherwise documentation, comments, and tests, so it is mergeable with that limitation tracked.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

CI showed that Windows TCC drivers accept a host buffer as the replacement
operand of a captured memcpy node in an executable update, while Linux
drivers reject it. The test now accepts either outcome and checks the copy
when the update is accepted, and the release note says the driver decides.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
cuda_core/tests/graph/test_graph_node_update.py (1)

931-932: 🎯 Functional Correctness | 🔵 Trivial | ⚡ Quick win

suggestion: Instantiate other before updating the definition node.

_capture_device_memcpy captures a device-to-device copy. Since other is created after node.update(src=host_src), its node-level update supplies the host source it already uses. The earlier executable only exercises instantiated_before.update(graph_def), the whole-graph update path.

Suggested fix
@@ -899,6 +899,7 @@
     # The builder is unused; it keeps the captured graph alive.
     builder, graph_def, node, dst, src = _capture_device_memcpy(init_cuda, stream)
     instantiated_before = graph_def.instantiate()
+    other = graph_def.instantiate()
     host_src = LegacyPinnedMemoryResource().allocate(64)
     ctypes.memset(int(host_src.handle), 0xA5, 64)

@@ -928,7 +929,6 @@
     else:
         assert "PARAMETERS_CHANGED" in str(rejected)

-    other = graph_def.instantiate()
     rejected = _cuda_error_from(lambda: other[node].update(dst=dst, src=host_src, size=64))

ℹ️ Review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuda-python/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: f917a584-4241-4e7c-965a-8e62ad5e5600
📥 Commits

Reviewing files that changed from the base of the PR and between 6bb6d63 and 1f96478.

📒 Files selected for processing (2)
  • cuda_core/docs/source/release/1.3.0-notes.rst
  • cuda_core/tests/graph/test_graph_node_update.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.

@juenglin juenglin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agree with CodeRabbit,

One new test does not cover the node-level executable update path it intends to check; adjusting the instantiation order would fix that.

A diagnostic run across Linux, Windows MCDM, and Windows TCC rows showed
that the executable-update outcome follows the allocation kind, not the
platform: the driver compares the memory class of pool-backed and
virtual-memory operands and rejects a replacement by host memory or by a
cuMemAlloc buffer, while plain cuMemAlloc operands accept it. The TCC
runners report no memory-pool support, so the default memory resource is
not pool-backed there, which is why the update was accepted on them. The
release note and the test docstring now say so.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda.core Everything related to the cuda.core module documentation Improvements or additions to documentation P1 Medium priority - Should do

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG]: MemcpyNode.update() rejects stream-captured memcpy nodes

2 participants