Repository navigation
Conversation
PyTorch >= 2.14 exposes torch_tensor_from_pyobject / aoti_torch_delete_tensor_object as part of the stable C ABI (landed via pytorch/pytorch#183323). These let us obtain an AtenTensorHandle from a Python tensor object without peering at THPVariable's internal layout. Dispatch is driven by *symbol presence* via dlsym / GetProcAddress on first call to view_as_torch_tensor: - resolved -> primary path (owned handle, deleted after metadata reads) - NULL -> fall back to the existing pyobj_to_aten_handle pointer-arithmetic trick, which stays correct for torch 2.3-2.13 (the only versions where this matters). The outer version gate in _memoryview.pyx (``2.3 <= torch <= 2.14``) is unchanged -- it still bounds the fallback. Once we drop support for torch < 2.14 we can remove the upper bound entirely because the shim itself is part of torch's stable ABI contract. Diagnostics kept module-private: - _get_pyobj_path_counts() / _reset_pyobj_path_counts() - _get_shim_available() - _set_use_shim(bool) to force the fallback for A/B benchmarking.
Contributor
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Experimental: adopt the stable
PyObject -> TensorC ABI shim that landed in PyTorch 2.14 (pytorch/pytorch#183323, merged asb69f7838a7) as the primary path in the tensor bridge, keeping the existingpyobj_to_aten_handlepointer-arithmetic trick as the fallback for torch 2.3-2.13.Changes
cuda_core/cuda/core/_memoryview.pyx: bump the outer version gate from(2, 12)to(2, 14). Resolves theBufferError: only CUDA Array Interface v3 or above is supportedsymptom I hit in Bump nightly PyTorch highest version to 2.14.1 #3031 against torch 2.14.1.cuda_core/cuda/core/_tensor_bridge.pyx: resolvetorch_tensor_from_pyobject+aoti_torch_delete_tensor_objectviadlsym/GetProcAddresson first call. Dispatch key is symbol presence, not version string:try/finally)_memoryview.pyx)The outer version gate (
2.3 <= torch <= 2.14) is unchanged. Once we drop torch < 2.14 support we can remove the upper bound entirely because the shim itself is a stable ABI contract.Module-private diagnostics (used for the benchmark below):
_get_pyobj_path_counts()/_reset_pyobj_path_counts()_get_shim_available()_set_use_shim(bool)for A/B benchmarking against the same torch versionVerification (local)
Env: fresh conda, python 3.12, CUDA 13.4 toolkit, RTX 6000 Ada; test target
cuda_core/tests/test_utils.py -k torch(63 tests:TestViewGPU[torch-*],TestViewCudaArrayInterfaceGPU[torch-*],test_torch_tensor_bridge_dtypes[*],test_ml_dtypes_bfloat16_torch_dlpack).The ptr/itemsize assertions in
test_torch_tensor_bridge_dtypescover 13 dtypes end-to-end — if the shim's owned handle were disagreeing with torch's view of the storage, these would fail. Path counters confirm each call is routed as expected.Benchmark
Same env, same tensor (32×32 float32 CUDA), 30 outer × 10 000 inner iterations, 2 000 warm-up, toggle via
_set_use_shim:The ~190 ns overhead is the cost of the shim's
at::Tensorcopy-construction (bumps the storage refcount) plus the matchingaoti_torch_delete_tensor_object. The pointer-arithmetic fallback does neither — the Pythonobjalready keeps the storage alive viabuf.exporting_obj = obj.Open question — do we need the refcount bump?
The shim returns a new reference (
// returns new referenceinshim.h). Forview_as_torch_tensorwe only read metadata off the handle (get_data_ptr,get_sizes,get_strides,get_dtype,get_device_type,get_device_index) and the Pythonobjoutlives the handle, so a borrowed variant that skipped the refcount bump would be functionally equivalent and would close the perf gap.Worth asking upstream for a
torch_tensor_borrow_from_pyobject(or equivalent flag on the existing shim)? Logging here for discussion.-- Leo's bot