NTH

OpenURMA: A Clean-Room Open Implementation of the Unified Bus Protocol

AuthorsBojie Li

July 17, 2026 3 min read
Watch on YouTube
The one-line take

OpenURMA is the first open implementation of Huawei’s new Unified Bus protocol, showing that a load/store-style design can cut remote-memory latency and boost throughput versus traditional RDMA.

Key results

500
Remote-fetch latency

Nanoseconds for a 64-byte UB load/store remote fetch.

4.37
Latency reduction

Fold reduction versus the 2,186-nanosecond matched OpenRoCE baseline.

4855
Connection-state reduction

Fold reduction at 1,024 local and 1,024 remote endpoints.

322
FPGA timing

Megahertz post-route clock target reached by the OpenURMA RTL.

14.1%
LUT utilization

Alveo U50 LUT budget used by OpenURMA.

9.2
YCSB-A throughput gain

Fold higher application throughput than RoCE DMA in the YCSB-A workload.

What the paper found

OpenURMA, from Bojie Li at Pine AI, is the first clean-room open implementation of Huawei’s Unified Bus protocol, developed to address RDMA’s Queue-Pair-over-PCIe bottleneck. Unlike RoCEv2, which maintains per-application, per-remote-endpoint state and requires repeated PCIe crossings, UB separates application endpoints called Jetties from per-host TP Channels, reducing state growth from O(N·M) to O(N+M), places the controller on the CPU’s on-chip bus, and exposes native load/store access with four opt-in ordering axes. The implementation uses a shared OpenClickNP description across synthesizable RTL on an Alveo U50, a cycle-accurate two-node SystemC simulator, and a gem5 full-system scaffold, each compared with matched OpenRoCE. For a 64-byte remote fetch, UB’s load/store path reaches 500 ns end to end versus 2,186 ns for RoCE, a 4.37-fold reduction, while sustaining 2.80-fold higher work-request throughput. At 1,024 local and 1,024 remote endpoints, OpenURMA uses 110.6 KB of connection state versus 536.9 MB for RoCE, a 4,855-fold reduction, and the RTL closes at 322 MHz using 14.1% of the U50’s LUT budget. The study also reports 9.2-fold higher YCSB-A throughput, selective acknowledgments that outperform Go-Back-N under loss, and congestion control reaching 96.0% utilization versus 62.5% for DCQCN. These are modeled and FPGA post-route results rather than measurements on physical UB silicon; Huawei’s Ascend 950 is closed, while NVIDIA’s ConnectX-7 data calibrates the RoCE baseline. The paper states that its code and manuscript were produced with Pine Copilot and Claude Code.

Original abstract

Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of megabytes at 1024-application fanout - and pays a four-traversal PCIe round trip on a 64-byte operation, inflating latency an order of magnitude beyond the wire. Both follow from the Queue Pair over PCIe abstraction RDMA inherits from InfiniBand. Huawei's Unified Bus (UB), a public 2025 specification, changes the abstraction: it decouples per-application endpoint state from per-host transport state so connection context grows additively, exposes ordering as opt-in, and reaches remote memory through native CPU load/store to an on-chip-bus controller. UB ships in Huawei's closed Ascend 950 silicon. OpenURMA is the first clean-room open implementation of UB's transport and transaction layers, realised at three tiers - synthesisable RTL on Alveo U50, a cycle-level two-node SystemC simulator, and a gem5 full-system scaffold - each with a matched OpenRoCE (RoCEv2 RC) baseline. The contribution is the implementation, harness, and controlled comparison closed silicon does not admit. On the canonical 64-byte remote fetch - LOAD on UB-spec Sec.8.3, READ on RoCEv2 RC - UB's load/store path delivers ~500 ns end-to-end, 4.37x below the matched baseline (2186 ns), sustains 2.80x higher throughput, and fits in ~14% of a U50's LUTs.

Read the original paper

More in AI Hardware

Browse all 34 papers →