Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

37 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

torch-infini

Experimental PyTorch plugin for the Infini stack.

The plugin keeps upstream PyTorch unchanged and registers the infini device through PyTorch's PrivateUse1 backend slot:

import torch
import torch_infini

src = torch.arange(16, dtype=torch.float32).reshape(4, 4)
x = torch.empty(src.shape, dtype=src.dtype, device="infini:0")
x.copy_(src)
y = torch.add(x, x)

out = torch.empty_like(src)
out.copy_(y)
torch.testing.assert_close(out, src + src)

This first-step bridge is intentionally narrow. It wires PyTorch device and stream management, device and pinned-host allocation, synchronization, contiguous tensor copies, shared ATen tensor metadata adapters, and aten::view, aten::add.Tensor, aten::mul.Tensor, aten::mm, aten::bmm, aten::addmm, aten::relu, aten::rms_norm, and aten::silu to the Infini stack. General ATen operator coverage is left to later integration work.

The implementation follows PyTorch's documented out-of-tree backend path: PrivateUse1 is renamed to infini, C++ kernels are registered through the dispatcher, and the extension is built with torch.utils.cpp_extension.

Build

Build and install InfiniRT first. Then build InfiniOps without its PyTorch backend and with the generated C++ operator call surface needed by downstream consumers. A focused CPU build can use a small operator allowlist:

cmake -S /path/to/InfiniOps -B /tmp/infini-ops-build \
  -DCMAKE_INSTALL_PREFIX=/path/to/infini-ops-prefix \
  -DINFINI_RT_ROOT=/path/to/infini-rt-prefix \
  -DWITH_CPU=ON \
  -DWITH_TORCH=OFF \
  -DGENERATE_OPERATOR_CALL_INSTANTIATIONS=ON \
  -DINFINI_OPS_OPS=add,gemm,mul,relu,rms_norm,silu
cmake --build /tmp/infini-ops-build --target infiniops -j
cmake --install /tmp/infini-ops-build

Enable the InfiniOps backend options that match the InfiniRT build, such as WITH_NVIDIA=ON, when targeting an accelerator. Point this package at both installed prefixes:

export INFINI_RT_PREFIX=/path/to/infini-rt-prefix
export INFINI_OPS_PREFIX=/path/to/infini-ops-prefix
pip install --no-build-isolation --no-deps .

INFINI_OPS_INCLUDE_DIRS, INFINI_OPS_LIBRARY_DIRS, INFINI_RT_INCLUDE_DIRS, and INFINI_RT_LIBRARY_DIRS can be used when headers or libraries are not under their respective install prefixes.

The wheel links to libinfiniops.so before libinfinirt.so but does not bundle either library or store their absolute build paths. Before importing torch_infini, make both library directories available to the dynamic loader. Add the directories to the system loader configuration and run ldconfig, or expose them for the current shell:

export LD_LIBRARY_PATH="/path/to/infini-ops-prefix/lib:/path/to/infini-rt-prefix/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"

Use lib64 instead of lib when that is the library directory in either install prefix.

Some InfiniRT installations currently expose CUDA headers through their public headers. For those installations, set CUDA_INCLUDE_DIRS when CUDA headers are outside the standard toolkit paths. torch-infini does not otherwise depend on CUDA; InfiniRT should eventually export any required transitive include paths.

Compatibility

Source-build compatibility is tested for every combination of Python 3.10, 3.11, and 3.12 with PyTorch 2.12 and 2.13. The tested native dependencies are InfiniRT commit 95c70080f9551e61241110497d163dfcdf9dc7e7 and InfiniOps commit 6c66469dcc5cb229a02dcfa348f8c32ef8155b96.

Binary wheel builds, editable builds, and in-place builds record the full PyTorch version, normalized major.minor version, and CXX11 ABI mode used to compile the extension. At import time, torch-infini requires the runtime PyTorch major.minor version and CXX11 ABI mode to match before it registers the infini backend or loads the native extension. Patch versions and local, development, alpha, beta, and release-candidate suffixes may differ when the leading major.minor version matches. Rebuild or reinstall torch-infini after changing the PyTorch minor version or CXX11 ABI mode.

Runtime backend

torch-infini automatically uses an accelerator backend compiled into InfiniRT when that backend reports at least one available device. When no compiled accelerator backend has an available device, it falls back to CPU when CPU support is included in the InfiniRT build. torch-infini requires this CPU support so the fallback is always available. InfiniRT stores its selection per thread, so torch-infini binds the selected backend whenever a thread enters a runtime operation.

Conformance testing

The required conformance suite checks the supported torch.infini API and runs backend-neutral device, stream, and event contracts:

python -m pytest -q tests/conformance

These tests gate the capabilities classified as required in tests/conformance/api_profile.json. CUDA cases run when CUDA is available and skip otherwise. Infini cases always run through the automatically selected InfiniRT backend, including the CPU fallback.

The full torch.cuda API surface is tracked by an advisory report:

python tools/report_api_gaps.py \
  --profile tests/conformance/api_profile.json \
  --markdown build/conformance/torch-cuda-gaps.md \
  --json build/conformance/torch-cuda-gaps.json

The command records required, planned, excluded, and unclassified symbols in both Markdown and JSON. Compatibility gaps do not make the command fail, but an invalid profile or a report-generation failure does.

CUDA is a behavioral reference rather than the specification, so the contracts compare stable state transitions instead of hardware-specific values. Future operator conformance work will use CPU execution as the numerical oracle and CUDA as an additional accelerator reference.

Scope

The initial implementation supports:

  • device="infini:0"
  • torch.infini.is_available()
  • torch.infini.device_count()
  • torch.infini.current_device()
  • torch.infini.set_device(index)
  • torch.infini.synchronize()
  • torch.infini.Stream()
  • torch.infini.current_stream()
  • torch.infini.default_stream()
  • torch.infini.set_stream(stream)
  • torch.infini.stream(stream)
  • torch.infini.Event()
  • event record, query, synchronize, elapsed-time, and stream-wait operations
  • torch.empty(..., device="infini")
  • torch.empty_strided(..., device="infini")
  • storage-sharing torch.as_strided, Tensor.view(shape), and matrix-transpose metadata views, including view-compatible reshape and nn.Flatten
  • Tensor.pin_memory("infini") and Storage.pin_memory("infini")
  • contiguous copy_ between CPU and Infini tensors, with asynchronous return for pinned CPU memory when non_blocking=True and the selected InfiniRT backend advertises the required capabilities
  • internal ATen-to-InfiniRT TensorView and InfiniOps execution-context adapters
  • same-dtype, same-device torch.add(tensor, tensor) through native InfiniOps implementation index 0, including broadcasted and strided inputs
  • same-dtype, same-device torch.mul(tensor, tensor) through native InfiniOps implementation index 0 on CPU, NVIDIA, Iluvatar, MetaX, Moore, and Ascend, including broadcasted and strided inputs
  • float32 torch.mm, torch.bmm, and torch.addmm inference through native InfiniOps implementation index 0, plus nn.Linear inference for matrix inputs and contiguous [batch, sequence, hidden] inputs, including transposed dense weights and broadcast bias
  • out-of-place torch.relu and nn.ReLU inference through native InfiniOps implementation index 0 on CPU, NVIDIA, Iluvatar, MetaX, and Moore, including supported floating-point and integer dtypes and PyTorch-compatible output layouts
  • float32 torch.nn.functional.rms_norm and nn.RMSNorm inference through native InfiniOps implementation index 0 for contiguous two-dimensional and three-dimensional inputs with one normalized dimension and an affine weight
  • out-of-place torch.nn.functional.silu and nn.SiLU inference through native InfiniOps implementation index 0 on CPU, NVIDIA, Iluvatar, MetaX, and Moore, including supported floating-point dtypes and PyTorch-compatible output layouts
  • end-to-end float32 inference for two-layer nn.Sequential MLPs composed of nn.Linear, nn.ReLU, and nn.Linear, with matrix or contiguous [batch, sequence, hidden] inputs, on the ReLU-enabled runtime backends listed above
  • end-to-end float32 inference for gated MLP blocks composed of parallel, bias-free nn.Linear gate and up projections, torch.nn.functional.silu, tensor multiplication, and a bias-free nn.Linear down projection, with contiguous [batch, sequence, hidden] inputs, validated on the InfiniRT CPU and NVIDIA backends

The torch.infini module follows torch.cuda naming and semantics for the device and stream-management operations it implements. Stream priorities, random-number generation, and other general ATen operators are not exposed yet. reshape operations that require materializing a contiguous copy are not supported yet. For CPU-to-Infini and Infini-to-CPU copies, non_blocking=True returns before completion only when the CPU tensor uses torch-infini pinned memory and the selected backend supports both pinned-host allocation and asynchronous memcpy. Copies involving ordinary host memory, an unsupported backend, or device storage whose lifetime cannot be tracked complete synchronously. This guarantee is limited to contiguous CPU-to-Infini and Infini-to-CPU copies and does not cover the lifetime of storage used by other asynchronous operators. InfiniRT does not currently expose the capabilities needed for blocking or interprocess events, so those event constructor options raise NotImplementedError. Event operations are validated with the InfiniRT CPU and NVIDIA backends; other backends require corresponding InfiniRT event support. The initial tensor Add path requires alpha == 1 and does not perform dtype promotion. The initial tensor Mul path does not provide scalar, out, or in-place overloads, dtype promotion, or backward support. The initial matrix multiplication paths are limited to float32 non-overlapping dense inputs: two-dimensional inputs for mm and three-dimensional inputs with matching batch dimensions for bmm. PyTorch's composite nn.Linear path uses these matrix kernels and metadata views to support contiguous three-dimensional inputs with or without bias. The initial addmm path composes InfiniOps Gemm and Add and requires alpha == 1 and beta == 1. The path does not provide mm.out, bmm.out, addmm.out, training, or backward support. On NVIDIA, it follows the InfiniOps TF32 path and matches PyTorch's default high float32 matrix-multiplication precision; selecting highest precision does not yet change InfiniOps execution. The initial ReLU path does not provide relu_, an out overload, or backward support. InfiniOps does not yet provide native ReLU implementation 0 for Cambricon or Ascend, so torch-infini reports those runtime backends as unsupported for this operator. The initial RMSNorm path does not support noncontiguous inputs, more than one normalized dimension, non-float32 dtypes, missing affine weights, an out overload, or backward. Its CPU and NVIDIA paths are validated here; other InfiniRT backends with InfiniOps RmsNorm implementation 0 still require independent validation. Unsupported operations should fail clearly instead of silently falling back through CPU. The initial SiLU path does not provide silu_, an out overload, or backward support. InfiniOps does not yet provide native SiLU implementation 0 for Cambricon, Ascend, or Hygon, so torch-infini reports those runtime backends as unsupported for this operator.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages