Skip to content

Repository files navigation

UCL-MPComm

UCL-MPComm

Unified Communication Library — Memory Pool Communication

RDMA NIC CUDA Python

UCL-MPComm (Unified Communication Library — Memory Pool Communication) is a high-performance, RDMA-native data-plane library for large-scale heterogeneous memory pools, developed by the Tencent Astral Network Team.

Large volumes of time-sensitive data—such as reinforcement learning datasets, R3 "expert IDs," and inference KV caches—are rapidly generated, transmitted, and consumed across diverse storage media like DRAM and HBM, giving rise to large-scale memory pooling architectures. In these pooling scenarios, traditional inter-memory communication schemes encounter bottlenecks such as inefficient bandwidth utilization and rigid communication patterns. Designed specifically for large-scale, heterogeneous memory pooling environments, UCL-MPComm addresses these limitations through extreme optimization, achieving performance levels that approach the physical limits of the hardware.

Key Features | Getting Started | Performance | Open Source Plan


Key Features

  • Native ibverbs one-sided ops — RDMA WRITE / READ, bypassing the kernel and extra copies to reach remote memory directly.
  • Multi-NIC aggregation (multi-rail) — a single MPComm instance stripes across multiple CX-7 NICs in parallel; bandwidth scales near-linearly with NIC count.
  • Multi-QP concurrency — configurable QPs per NIC per connection (MPCOMM_QPS_PER_CONNECTION) for deeper pipelining.
  • NUMA affinity — auto-detects NIC ↔ NUMA topology; allocates / publishes / matches buffers per NUMA node.
  • Two data pathszero-copy (pre-registered user buffer, direct one-sided access) and non-zero-copy (unregistered buffer routed through a multi-threaded staging pool).
  • HBM ↔ DRAM movement — Hopper TMA (cp.async.bulk) and SM int4-vectorized H2D kernels (tmaGather / tmaScatter).

Getting Started

Dependencies

  • Linux + Mellanox/NVIDIA RDMA NIC (RoCE / IB), with rdma-core / OFED installed
  • libnuma (NUMA allocation)
  • CUDA (optional; enables GPU direct and TMA kernels, requires sm_90+ / H100·H800·H20)
  • CMake ≥ 3.x, C++17, Python ≥ 3.8

Build & Install

Option A: CMake build (C++ + Python extension)

cd ucl-mpcomm
sh build.sh
# Artifacts:
#   mpcomm-install/lib/libmpcomm.so
#   mpcomm-install/lib/python/mpcomm   (Python module)

To use it from Python without configuring PYTHONPATH, add it to your environment:

export PYTHONPATH=/path/to/ucl-mpcomm/mpcomm-install/lib/python:$PYTHONPATH

Option B: install as a Python wheel

pip install .          # or: sh build_wheel.sh
python -c "import mpcomm; print(mpcomm.__file__)"

Performance

Test environment: Tencent cluster, 2 nodes, single NUMA (1 CPU) × 4 CX-7 NICs aggregated over a dual-plane network (each NIC a 2×200G bond), batch=4, block size swept 512KB → 1GB. Results are reported on two platforms: AMD Turin / RTX PRO 5000 and AMD Genoa / NVIDIA H20.

UCL-MPComm 4-NIC batch=4 bandwidth: zero-copy vs Mooncake TENT, non-zero-copy vs UCX

  • Zero-copy path — on a single-NUMA, 4-NIC dual-plane network, UCL-MPComm delivers roughly 30% higher point-to-point memory transfer bandwidth than Mooncake TENT.
  • Non-zero-copy path — benefiting from the Two-Stage pipeline and multi-threaded concurrency, throughput improves by up to 5×.
Measured peak bandwidth (1GB block, batch=4, 4 NICs)
Path Platform UCL-MPComm PUT / GET Baseline PUT / GET Gain
Zero-copy AMD Turin / RTX PRO 5000 195.7 / 195.7 GB/s Mooncake TENT 142.2 / 142.3 GB/s ~1.38×
Zero-copy AMD Genoa / NVIDIA H20 165.7 / 152.3 GB/s Mooncake TENT 122.0 / 111.7 GB/s ~1.36×
Non-zero-copy AMD Turin / RTX PRO 5000 91.9 / 84.5 GB/s UCX 20.9 / 16.7 GB/s up to ~5×
Non-zero-copy AMD Genoa / NVIDIA H20 77.9 / 73.3 GB/s UCX 15.0 / 10.3 GB/s up to ~5×

Zero-copy pre-registers the user buffer and accesses it directly with one-sided verbs; non-zero-copy routes unregistered buffers through a multi-threaded staging pool, contrasted with UCX's copy-in mechanism.


Open Source Plan

  • MPComm main framework
  • NVLink forwarding
  • H2D/D2H accelerating
  • Non-zero-copy convenience API
  • Centralized deployment mode
  • Intra-/internode unified QoS
  • Scatter/gather acceleration

License

Licensed under the Apache License, Version 2.0.

Copyright (C) 2026 Tencent. All rights reserved.

About

A high-performance, RDMA-native data-plane library for large-scale heterogeneous memory pools, developed by the Tencent Astral Network Team.

Resources

Stars

31 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages