Skip to content

cuda.core: share VirtualMemoryResource buffers between processes - #3028

Draft
Andy-Jost wants to merge 7 commits into
NVIDIA:mainfrom
Andy-Jost:ajost/vmm-ipc-fabric
Draft

Andy-Jost wants to merge 7 commits into
NVIDIA:mainfrom
Andy-Jost:ajost/vmm-ipc-fabric

Conversation

@Andy-Jost

@Andy-Jost Andy-Jost commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

VirtualMemoryResource buffers can now be shared with other processes the way pool-backed buffers are. Buffer.ipc_descriptor on a VirtualMemoryBuffer exports every physical allocation that backs the buffer as a POSIX file descriptor, Buffer.from_ipc_descriptor imports the descriptor with a VirtualMemoryResource of the receiving process for the device that owns the memory, and a buffer or its descriptor can be sent through multiprocessing directly. VirtualMemoryResource.is_ipc_enabled is True when the resource's handle_type is "posix_fd" on Linux; handle_type is the switch and is_ipc_enabled the query. A grown buffer imports as one contiguous range with the same byte layout.

Changes

  • _rt: import_mem_allocation_handle() wraps cuMemImportFromShareableHandle in the same box and cuMemRelease deleter as create_mem_allocation_handle(), with the importer's access descriptors; the box records whether the allocation was imported (mem_allocation_is_imported()), and the mark travels with the allocation into every range that maps it. cuMemImportFromShareableHandle joins the driver function table.
  • _ipc: VirtualMemoryIPCBufferDescriptor, an IPCBufferDescriptor subclass that carries one IPCAllocationHandle and one size per physical allocation, in address order. It exposes handle_type, sizes, handles, and fds, and from_fds() rebuilds a descriptor in the receiving process from file descriptors received through a transport of your own (integers are duplicated, so the caller keeps ownership of its own; IPCAllocationHandle objects are shared). The descriptor's own pickling stays multiprocessing-only, which duplicates the file descriptors into the receiver; plain pickle raises. Buffer.from_ipc_descriptor dispatches on the descriptor's class and rejects a mismatched resource with TypeError.
  • VirtualMemoryResource: is_ipc_enabled; the import path, which imports every allocation with the resource's handle type, checks that each one lives where the resource allocates (the same device, or the same host location type), reserves the total, maps the allocations in order with the resource's access options for its device and peers, and records the deallocation stream as allocate() does; __reduce__, so a buffer pickles as (resource, descriptor). VirtualMemoryBuffer.ipc_descriptor exports on every access and returns a new descriptor; nothing is cached on the buffer. A modify_allocation() config must keep the resource's handle_type.
  • Error reporting: a descriptor is data from another process. Malformed fields (handle type, chunk count, chunk sizes that are not positive multiples of the granularity, a size beyond the exported allocations, a size with no allocations, a closed handle) raise ValueError before any driver call. Driver failures during export, import, or map carry a note that names the allocation and, for the map, says that a size which does not match the exported allocation fails there.
  • Docs: a "Sharing across processes" section in VMM_DESIGN.md, docstrings, VirtualMemoryIPCBufferDescriptor in the private API reference, and a 1.3.0 release note.

Rules

  • A descriptor pins the physical memory while it exists, in every process that holds one, and nothing else does. Each exported allocation costs one file descriptor while the descriptor lives. A descriptor parked in a Queue or sent to a Pool pins the memory in the sender until the receiver has unpickled it. Pickling a buffer builds a transient descriptor, so a sender holds no file descriptors between sends, and the importer does not keep the descriptor it imported from. A buffer passed as a Process argument is the one exception: a spawned child receives the file descriptors when it is created, after pickling, so that descriptor lives on the Popen until the Process object is released.
  • The importing resource must be for the device that owns the memory; other devices go in its peers option. A resource for another device gets a ValueError that names both devices.
  • Memory imported from another process cannot be exported again. ipc_descriptor and pickling on an imported buffer, or on any buffer modify_allocation() derived from one (alias, in-place grow, or move), raise a RuntimeError before any driver call that says so and points to the two alternatives: forward the descriptor it was imported from, or copy into a buffer you own. A Queue feeder thread reports the same error and drops the item; the queue stays usable. This matches the contract PyTorch applies to received CUDA tensors.

Tests

The existing memory_ipc suite runs against VirtualMemoryResource through a third ipc_memory_resource parameter (VirtualMR), with the scenarios that need the pool half of the protocol (allocation handles, the registry, uuid, plain pickle of a buffer, re-export of an imported buffer) guarded or skipped for it. test_vmm_ipc.py covers what is specific to virtual memory: a three-chunk range placed by a relocating grow, import on a second device with and without peer access, the descriptor validation and rollback paths, the re-export errors for imported and derived buffers (direct, Pipe, Process arguments, and the Queue feeder thread), file descriptor counts returning to baseline on both sides, export under a lowered RLIMIT_NOFILE, raw-fd round trips through os.dup and a Unix socket with SCM_RIGHTS, empty buffers, a DLPack consumer, same-process aliasing, and file descriptor ownership.

Fabric handles

The path for handle_type="fabric" would differ only in the payload: a 64-byte CUmemFabricHandle per allocation instead of a file descriptor. It is left out of this PR because no available test system reports CU_DEVICE_ATTRIBUTE_HANDLE_TYPE_FABRIC_SUPPORTED or exposes an IMEX channel, so it could not be verified. is_ipc_enabled reports False for "fabric" and the docs say so.

Related Work

Part of #2980. Builds on the _rt handle layer that #2917 gave VirtualMemoryResource.

🤖 Generated with Claude Code

Buffers of a VirtualMemoryResource whose handle_type is "posix_fd" can now
be exported and imported like pool-backed buffers. Buffer.ipc_descriptor
exports every physical allocation that backs a VirtualMemoryBuffer as a
file descriptor, Buffer.from_ipc_descriptor imports the descriptor with a
VirtualMemoryResource of the receiving process, and a buffer or its
descriptor can be sent through multiprocessing directly. A grown buffer
imports as one contiguous range with the same byte layout.

The import reuses MemAllocationHandle: import_mem_allocation_handle wraps
cuMemImportFromShareableHandle in the same box and cuMemRelease deleter as
create_mem_allocation_handle, with the importer's access descriptors.
VirtualMemoryIPCBufferDescriptor carries one IPCAllocationHandle and one
size per allocation and owns the file descriptors; multiprocessing
duplicates them into the receiving process. VirtualMemoryResource pickles
as (device, options), which lets a buffer pickle as (resource, descriptor).

Fabric handles are not shared yet: no available test system reports fabric
support or exposes an IMEX channel, so that path could not be verified.

Part of NVIDIA#2980.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost Andy-Jost added this to the cuda.core 1.3.0 milestone Oct 6, 2026
@Andy-Jost Andy-Jost added P0 High priority - Must do! feature New feature or request cuda.core Everything related to the cuda.core module labels Oct 6, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@Andy-Jost Andy-Jost self-assigned this Oct 6, 2026
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test 2231df7

The CI wheel build compiles with -Werror and -Wsign-compare, and the VMM
import path compared a size_t count with the Py_ssize_t that len() returns.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test 682db0f

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Under pytest-run-parallel the same test body runs in several threads. The
alias test checks mapping state by address, and a freed address can be
reused by another thread's allocation; the descriptor test checks a closed
file descriptor number, which another thread can reuse.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test 4ed2c03

Address the review of the VirtualMemoryResource IPC path.

Import: check every imported allocation's location against the
resource's before mapping (the same device, or the same host location
type) and raise a ValueError that names both; reject a descriptor with
a size but no allocations; attach notes to driver failures that say
which chunk failed and why a peer-controlled size fails at the map.
The imported buffer no longer keeps the descriptor: the driver copies
the shared allocation into a handle of its own and keeps no reference
to the file descriptor, so the descriptors live only as long as the
caller holds the descriptor object.

Export: an imported buffer, and a buffer grown from one, cannot be
exported again because the driver exports only allocations created
with the requested handle type; say so instead of surfacing
INVALID_VALUE. modify_allocation() rejects a config whose handle_type
differs from the resource's, so every chunk of a buffer is exportable
the same way. __reduce__ preserves subclasses.

Tests: parametrize the memory_ipc suite over VirtualMemoryResource
(VirtualMR), with the pool-only scenarios guarded or skipped;
PatternGen works on the first `size` bytes of a larger buffer; the
fd-leak harness asserts on the success path; new VMM-only tests for
many-chunk ranges, a second device with and without peers, descriptor
validation and rollback, re-export, empty buffers, and a DLPack
consumer.

Docs: file descriptor and memory lifetime, multiprocessing-only
transport, trust boundary, device rule, re-export limit;
VirtualMemoryIPCBufferDescriptor in api_private; the
from_ipc_descriptor stub has its mr type again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test ef9b712

…M IPC

Apply the decisions on the VirtualMemoryResource IPC design.

No exporter-side cache. VirtualMemoryBuffer.ipc_descriptor exports on
every access and returns a new descriptor that owns its file
descriptors; they close when the descriptor is released. Pickling a
buffer for multiprocessing builds a transient descriptor, so a sender
holds no file descriptors between sends. A spawned Process is the one
exception: it receives the file descriptors when it is created, after
pickling, so the transient descriptor stays on the spawning Popen until
the Process object is released. A descriptor pins the physical memory
while it exists, and nothing else does.

No re-export of imported memory. Every physical allocation records
whether cuMemImportFromShareableHandle produced it
(mem_allocation_is_imported), and modify_allocation reuses the
allocation handles of its input, so the mark reaches every alias,
in-place grow, and move. Exporting such a buffer, directly or through
pickling, raises a RuntimeError before any driver call that says the
buffer contains memory imported from another process, that it cannot be
exported again, and that the caller should forward the descriptor it
imported from or copy into a buffer it owns. A Queue feeder thread
reports the same error and drops the item.

Raw-fd transport. VirtualMemoryIPCBufferDescriptor exposes handle_type,
sizes, handles, and fds, and from_fds() rebuilds a descriptor in the
receiving process from file descriptors received another way (the
integers are duplicated; IPCAllocationHandle objects are shared). The
descriptor's own pickling stays multiprocessing-only.

No ipc_enabled option: handle_type is the switch and is_ipc_enabled the
query, and the docs say so.

Tests: file descriptor counts return to baseline on both sides, export
under a lowered RLIMIT_NOFILE, the re-export errors for imported and
derived buffers (direct, Pipe, Process arguments, Queue feeder thread),
and raw-fd round trips through os.dup and a Unix socket with
SCM_RIGHTS.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test b804e18

…warding

Write down in VMM_DESIGN.md the opt-in design that would let an imported
VirtualMemoryResource buffer be forwarded (keep the received handles with
the imported buffer and duplicate them on export), why the option would
belong on the resource, and its cost, so that it does not have to be
re-derived. Not implemented; nothing needs it yet.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The module-level import of the POSIX-only resource module made pytest
fail to collect tests/memory_ipc/test_vmm_ipc.py on Windows, so the
Windows rows ran no cuda.core tests at all. Every test in the module is
skipped on Windows by its fixture; only the import has to move.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda.core Everything related to the cuda.core module feature New feature or request P0 High priority - Must do!

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant