Skip to content

Repository files navigation

CiruStrixLink

Direct GPU tensor transport and a two-host operator console for AMD Strix Halo over USB4.

CiruStrixLink prepares and manages the USB4 connection behind the GLM5.3 Flash CIRU STRIX IU4 deployment. Its direct NHI path uses GPU-owned DMA-BUF pools and USB4 hardware DMA rings to exchange tensor partials without the TCP/IP payload path. The accompanying HIP adapter validates the received message and performs the reduction on the GPU. The same connection retains USB4NET for control and other traffic.

The project brings that transport together with endpoint discovery, correctness checks, runtime configuration and an embedded browser console. The measured exact 64-KiB BF16 all-reduce completed in about 110 µs, versus about 342 µs through the matched RCCL socket path. The complete setup requires the StrixLink binary, the patched stream driver and a compatible runtime adapter; their roles and sources are documented below.

Get started · GPU transport architecture · Measurements and methodology · Download v0.3.3 · GLM deployment

What makes the transport useful

  • GPU-owned transfer pools. Dedicated uncached HIP allocations are exported as DMA-BUFs and imported by USB4STREAM. GPU kernels pack TX and consume RX; the CPU manages submission and completion rather than copying tensor payloads through socket buffers.
  • Exact two-rank reductions. The published adapter exchanges contiguous BF16 [8,4096] partials in both directions and sums on the GPU. Unmatched collectives retain the runtime's existing backend.
  • Explicit message ownership. A 32-byte footer carries rank, epoch, sequence and payload length, with a GPU-written commit word. Circular ring addressing, completion accounting and RX reposting govern safe reuse.
  • A qualified two-host lifecycle. StrixLink checks both peers, preserves USB4NET's HopID allocation, discovers the stream devices, and requires matching endpoints and an available lease before launching NHI.
  • An operator console. Inspect transport state, run bounded TCP link tests, watch model PP/TG and draft acceptance, and optionally configure/load/unload the packaged GLM pair with coordinated readiness and rollback.

The console is a single static Go binary. Its link tests need no Python, iperf3, or Go installation. Other runtimes can use the portable IP link; using the direct GPU path requires an adapter for their tensor operations.

Measured GPU-operation latency

Recorded on two Ryzen AI Max+ 395 / Radeon 8060S systems with the qualified Linux 7.2.2 kernel and ROCm 10. These are completed-operation timings in microseconds, not cable ping figures.

Operation Control Direct NHI result Evidence
64-KiB BF16 [8,4096] all-reduce, median RCCL 340.904–342.745 µs 109.897–110.807 µs Exact sums; 67.5–67.9% less latency
Single-token GPU activation + consumer credit — 45.505 µs BF16 pipeline mechanism gate
T128 activation + credit, median BF16 1,642.803 µs I8G32 1,043.402 µs 36.49% shorter cycle; experimental codec
6-MiB BF16 [768,4096] all-reduce, median RCCL 4,943–4,961 µs 4,516–4,521 µs Exact sums; separate PP768 experiment

The M8 gate used 16 warmups and 64 measured calls per rank. All 80 NHI exchanges per rank passed with zero timeouts, protocol errors, driver failures or dropped events. Raw records and timing boundaries are published with the results.

The ~3.1× M8 operation-rate improvement is a component result. Model impact depends on how much communication is exposed and how completion is scheduled. A later deferred-retirement development comparison improved verification cadence 4.05% under fixed synthetic acceptance, and a natural structured request improved 21.23 → 23.09 tok/s with the same accepted proposal count. Those experiments are distinct from the synchronous bridge in the published GLM source archive. The performance report identifies each variant, including the compressed path's scope and rejected M8 INT8 reduction.

Install and connect

Run the following on both Linux x86-64 hosts. Download tools are curl, tar and sha256sum; StrixLink reports any missing network dependencies.

VERSION=0.3.3
RELEASE_DIR=$(mktemp -d)
cd "$RELEASE_DIR"
RELEASE_URL="https://github.com/ciru-ai/CiruStrixLink/releases/download/v${VERSION}"
curl -fLO "$RELEASE_URL/ciru-strixlink-${VERSION}-linux-amd64.tar.gz"
curl -fLO "$RELEASE_URL/SHA256SUMS"
sha256sum -c SHA256SUMS
mkdir unpacked
tar -xzf "ciru-strixlink-${VERSION}-linux-amd64.tar.gz" -C unpacked
sudo install -m 0755 unpacked/ciru-strixlink /usr/local/bin/ciru-strixlink
ciru-strixlink version
ciru-strixlink prerequisites

Connect the USB4 cable, load thunderbolt-net on both hosts, and configure complementary addresses. Setup previews its changes until --apply is given.

# Host A: 10.77.77.1/30
sudo modprobe thunderbolt-net
ciru-strixlink setup --role a
sudo ciru-strixlink setup --role a --apply

# Host B: 10.77.77.2/30
sudo modprobe thunderbolt-net
ciru-strixlink setup --role b
sudo ciru-strixlink setup --role b --apply

After both sides are configured, run doctor on each against the other address. Start the temporary test server on B, then test from A:

# B: leave running until the test finishes
ciru-strixlink doctor --peer 10.77.77.1
ciru-strixlink serve

# A: in a separate host terminal
ciru-strixlink doctor --peer 10.77.77.2
ciru-strixlink test --peer 10.77.77.2 --duration 7s --streams 4 \
  --output "$HOME/ciru-strixlink-report.json"

Stop the test listener with Ctrl-C. The test uses TCP 55321 over the private USB4 interface and checks IP transport, integrity and reconnects. It does not run the GPU/NHI benchmarks above. The complete guide covers dependencies, existing profiles, tokens, firewall ports, jumbo frames, runtime environment files and the direct-NHI setup.

Browser console

Create the same private token file on both hosts as described in step 4 of the guide. Then run:

# Host B
ciru-strixlink agent --token-file "$HOME/.config/ciru-strixlink/peer.token"

# Host A
ciru-strixlink ui --peer 10.77.77.2 \
  --token-file "$HOME/.config/ciru-strixlink/peer.token"

Open http://127.0.0.1:7749 on A. From a separate desktop, use ssh -N -L 7749:127.0.0.1:7749 USER_A@HOST_A and open the same localhost URL. The agent binds USB4 TCP 7748; the console stays on loopback. Peer tokens protect agent access, not the browser listener.

Add --model-url http://127.0.0.1:8083 to show an existing model frontend's identity and metrics. The optional Launch tab requires the installed GLM runtime, fixed complementary ranks, shared token and scoped model helpers. Step 7 gives the commands for each host. Model control is disabled by default.

Enable the direct GPU path

Layer What to install or configure
Kernel 16-patch Linux 7.2.2 series, including DMA-BUF import and the stale-MSI-X reliability fix
Link Verified USB4NET first; then prepare both NHI endpoints and reconcile HopID 9/9, ring 4095, throttling 8192 ns and lease availability
Runtime Published GLM HIP bridge and custom vLLM, or an explicitly integrated adapter for another runtime
Privilege Scoped system service granting CAP_SYS_RAWIO to the importing runtime with root-owned executable paths

Use the ordered two-host NHI instructions. --runtime vllm generates configuration; it does not retrofit stock vLLM with the adapter. Architecture explains GPU allocations, packet layout, completion scheduling and the source/release boundary.

Documentation and source

Build and credits

git clone https://github.com/ciru-ai/CiruStrixLink.git
cd CiruStrixLink
make test
make vet
make linux-amd64

The binary is written to dist/ciru-strixlink-linux-amd64; building requires Go and Make. CiruStrixLink's Go utility is Apache-2.0. The bundled Linux patches retain GPL-2.0 and their original authors; the published HIP bridge retains MIT. The kernel foundation includes Linux USB4STREAM and jyatesdotdev/strix-rdma, with Ciru's Linux 7.2.2 rebases, reliability fix, GPU/runtime integration and operator tooling. Source identities and license boundaries are documented explicitly.

About

Fail-closed USB4 link qualification and paired Strix Halo model operations

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages