Why data arrives out of order
Data arrives out of order when a network spreads one transfer over many paths, because the pieces travel by routes of different speed.
Spreading a transfer over many paths uses the network well. Clusters that train and serve large machine learning (ML) models move huge amounts of data between machines, and a single path can leave capacity idle. Transports designed for these clusters, such as a multipath design for remote direct memory access and STrack, send the pieces of one transfer over many paths at once. The price is order. A later piece can arrive before an earlier one.
These transports often run on remote direct memory access (RDMA), which lets a network card place data straight into the memory of another machine. The card, not the processor, handles each arrival. That makes the card’s own memory the place where the cost of disorder lands.
What the receiver has to remember
The receiver has to remember which pieces have arrived and which are still missing, so that it can ask the sender again for the missing ones.
The network interface card (NIC) does this work in hardware. The window is the span of data that can be in flight at once, and a faster or longer link needs a bigger window to stay busy. The faithful way to track it is to keep a record for every slot in the window. That record grows with the window.
Getting the record wrong has a price either way:
- A receiver that forgets a gap may stall, waiting for data that never comes.
- A receiver that cannot tell what arrived may ask for pieces it already holds, and waste the bandwidth the multipath transport was meant to use.
Memory on the card is scarce. The authors of a well-known multipath RDMA design describe on-chip memory as very limited, and they note that swapping state out to host memory has a cost. Extra tracking state for each connection also cuts the number of connections a card can serve from its own memory. This is a real design limit, not a detail.
- arrived in order
- arrived early
- still missing
- not sent yet
One transfer drawn as slots. The receiver must remember which slots are filled, which arrived early and which are still missing, so that it can ask again for the gaps.
How a bounded-state receiver differs
A bounded-state receiver keeps a fixed amount of tracking state however large the window grows, instead of a record for every slot.
Across the window sizes the simulation tried, the lab’s mechanism did this. It keeps a small, fixed-size record of what is still missing, which its claim calls bounded descriptor state. The lab does not publish the details of that record yet. What it does publish is the scaling. Across the window sizes the simulation tried, the state stayed the same size as the window grew, while a faithful buffer grew in proportion to the window. The research page sets out the formal statement and the lab’s own words.
- a faithful buffer of the whole window
- bounded state
Shape only. This illustrates the two scaling laws from the formal statement. It shows no measured values.
Some designs avoid the buffer in another way. They place each arriving piece straight into its final spot in memory, which is the idea of the direct data placement standard linked below. The lab’s large ratio is computed over designs that do not do this, the ones that pay for a faithful buffer. Designs that do place data directly are a different comparison.
What a simulation says about hardware
A simulation says how a design behaves in a model of the network, and it does not say what the design costs in silicon.
The lab ran a closed-loop simulation over a grid of simulated transports and settings. Closed loop means the simulated traffic responds to the network instead of replaying a fixed script. The results are not measurements on hardware.
The claim gives two ratios over the same grid:
- The first compares the mechanism with a faithful buffer, and it is very large.
- The second compares it with our simulator’s model of STrack, prior art that also bounds its state, and it is 1.965x.
Judge the gain over what exists by the ratios against existing designs, not by the large one. Against a second bounded-state design in the lab’s simulator, its model of the clear-to-send admission Meta describes (Gangidi et al., SIGCOMM), the gain is smaller still: that design held on average 1.443x as much reorder state as the mechanism. The large ratio is against a baseline that bounded-state designs do not use.
STrack, a published multipath design, already bounds its state. In the lab’s simulation, our model of STrack still holds less reorder state than the mechanism in 9 of the 64 simulated settings of the grid, and the second design, the lab’s model of Meta’s clear-to-send admission, holds less in 31, nearly half of them. The lab’s command also searches for competitor designs that escape the mechanism, and the lab reports those cells instead of leaving them out.
What this means for you
The question of reorder memory sits inside the network hardware, so it reaches buyers indirectly, and each kind of reader can still ask for the evidence.
- Network hardware teams. If you design transport engines on network cards, compare any bounded-state design with STrack and with other designs that bound their state, not only with a faithful buffer.
- GPU clouds. Window sizes grow with link speed and distance, and the memory on a network card is small. Ask what limits the number of connections your cards can serve.
- Enterprise AI platforms. You buy this through your cloud’s network. Ask how reordering behaves under load, and where the answer comes from.
- Auditors. Ask for the grid, the seeds and the comparison with prior art, not only the headline ratio.
What this does not show
This is a simulation result, and it does not claim a new kind of bound.
- It is not a measurement on hardware. It does not show what the mechanism costs in throughput or retransmissions.
- Bounded-state designs already exist. The comparison with them is the second ratio.
- It does not claim that the mechanism beats every bounded-state design in every setting.
- This post does not describe the details of the mechanism, because the lab does not publish them yet.
The lab’s own limits for this result, word for word
Simulation, seeds [1,2,3], generated_utc 2026-06-27.