Skip to content

Receiver reorder memory for multipath transport

Tracking memory each design needed, relative to our receiver (ten blocks), by geometric mean over the simulated settings

Our receiver

Our simulator’s model of STrack, a published design: 1.965x

A second fixed-size design, modelled in our simulator: 1.443x

Simulated settings in which the other design needed less than ours

Our STrack model: 9 of 64

The second design: 31 of 64

Counts only: the lit squares are not the settings’ positions. A simulation, not silicon.

The result

In simulation, our receiver tracked out-of-order data in about half the memory of our simulator's model of STrack, a published design, on average; a second design, not cited here, came closer and used less in nearly half the settings

What it does not claim

This is a simulation result, not a measurement on hardware. It does not claim a new kind of bound: designs that keep their tracking state small already exist, and the comparison with them is the one to judge. It does not claim the mechanism beats every such design in every setting, and it does not say what the mechanism costs in throughput.

The limits, in the lab’s words:

Simulation, seeds [1,2,3], generated_utc 2026-06-27.

Excerpts, word for word from the published file, where the full text can be read.

When a network spreads one transfer across many paths, data arrives out of order and the receiver must track what is missing. In simulation, our simulator's model of STrack, a published design cited below that also keeps its tracking memory small, needed by geometric mean over the cells 1.965x as much tracking memory as our mechanism, and held less in 9 of 64 cells. Against a second bounded-state design in the lab's simulator, not cited here, the advantage is smaller, 1.443x on average, and that design held less in 31 of 64 cells. A full buffer, which grows with the data in flight, needs far more, but careful designs do not use one.

What it shows

In this model: a closed-loop network simulation, not hardware, of a receiver, run over a grid of simulated transports and settings; each setting is one cell of the grid.

Transports proposed for clusters that run machine learning (ML) workloads, like those cited below, move data between machines with remote direct memory access (RDMA) over many network paths at once. Spreading data over paths uses the network well, but the pieces then arrive out of order. The receiving network interface card (NIC) must remember which pieces have arrived and which are still missing, so it can ask for the right ones again.

The comparison that matters

Careful designs already keep that record small. STrack, a published multipath transport cited below, is one. On the same simulation grid of 64 cells, our simulator's model of STrack held on average 1.965x as much reorder state as our mechanism. A second bounded-state design in the lab's simulator, not cited here, came closer: it held 1.443x as much on average. These two ratios measure what the mechanism adds.

Neither design is beaten everywhere. Our simulator's model of STrack held less state than our mechanism in 9 of the 64 simulated settings, and the second design held less in 31, nearly half of them.

Against a full buffer

The simple way to track the data is a full buffer, a record for every slot in the window of data that may be in flight. That memory grows with the window, and memory on the card is small: Lu et al., cited below, describe on-chip memory as only a few megabytes. A full buffer needed on average 3,886x as much as our mechanism. That ratio is large because a full buffer is wasteful, so it says little about the mechanism against good designs.

The mechanism

Our mechanism keeps a small, fixed-size record of what is still missing. The details of that record are not published yet. The statement below gives the scaling the simulation showed.

Why it matters, and to whom

The result is for teams that design transport engines on network cards and data-processing units (network cards with processors of their own), and for architects who size on-card memory for multipath transport. Less tracking memory per connection leaves room on the card for more connections or larger windows. The published results do not measure what, if anything, the mechanism costs in throughput or retransmissions.

Why now

Transports for AI clusters, such as the Ultra Ethernet design cited below, spread each transfer over many paths, so out-of-order arrival is becoming the normal case. Window sizes grow with link speed and distance, while the memory on a network card is small. That makes the memory a receiver needs for out-of-order data a design limit now, for the teams building these cards.

How it was checked

The numbers come from a closed-loop simulation, one in which the simulated traffic responds to the network rather than replaying a fixed script. The seeds, the starting values for the simulation's random choices, and the date the results were generated are quoted in the limits below.

The average it reports

Each ratio is a geometric mean over the cells of the grid. A geometric mean is pulled less than an ordinary average by a few very large cells. The ratio against a full buffer averages only the designs that would need one, and leaves out the few cells where that buffer is empty. The lab reports several related measurements separately and does not add them together.

How to reproduce

The published results file names the command and the commit it was read at. The command recomputes the ratio against the full buffer from the grid's saved results and then searches for competitor designs that beat the mechanism; it does not re-run the whole simulation. Anyone with access to the lab's code at that commit can run it. The lab's code is not public yet. A visitor without that access can still check each file named on this page against the fingerprint listed for it in the list of published files.

Observed scaling, in symbols

In words: W is the reorder window, the span of data that can be in flight at once. Across the window sizes the simulation tried, our tracking state stayed the same size while a full buffer grew in proportion to W. This is an observed scaling in simulation, not a proof.

The result, in the lab's exact words

Read it with this condition: in this quote, the faithful buffering field is the full buffer; the faithful bounded-state prior art is our simulator's model of STrack, which has not been checked against STrack's own implementation or results; and the robust geomean is the geometric mean described above. The scaling wording describes the window sizes the simulation tried; it is a simulation result, not a proof. Judge the gain by the ratios against STrack and the uncited second design, not by the ratio against a full buffer.

Bounded descriptor state scales O(1) where the faithful buffering field scales Theta(W), giving a 3,886x robust geomean over the closed-loop grid — and 1.965x against faithful bounded-state prior art on the identical grid.

Quoted word for word from the lab's published claim and limits.

Prior art

Evidence

A name in these files: Meta CTS in the lab's files is the lab's simulator model of the clear-to-send admission Meta describes (Gangidi et al., SIGCOMM 2024, section 5.2.2); no public Meta source states its reorder memory.

Related results

The same result on VerifyCore Labs

All results from AxiomLimit · The AxiomLimit home page

How the numbers on this page are checked

Every number on this page links to the file it comes from. All published files.

  • Loadingshown only once its file has loaded in your browser
  • Not checkeda question we have not checked would say so, with no number
  • Checked, nothing founda search that found nothing would say so, with no number
  • No valuea file that holds no value for the question would say so
  • File missinga number whose file is missing or altered would be hidden
  • Unclear subjecttwo files that disagree about what a number describes would both be shown
  • Small samplea number from a small sample would carry its sample size
  • Conflicting filestwo files giving different values would both be shown
  • Not publishablea file we may not publish would be named by its fingerprint only
  • Out of datea measurement older than a week would carry its age
  • Does not applya question that does not apply to this page would say so
  • Run faileda measurement whose program failed would say so, with no number