I had a hypothesis. Different exchanges push updates at wildly different rates, so any cross-venue snapshot is comparing quotes of different ages — and that measurement artifact, not real economics, is what makes venue prices look different. It's the kind of claim that sounds obviously right.
So I recorded it. Nine BTC order books, sampled every two seconds for 56 minutes: 1,666 simultaneous cross-sections, 14,992 observations. Price, top-of-book depth, and time since that venue last sent a message.
The hypothesis was wrong. Clock misalignment accounts for about 3% of the observed dispersion. What accounts for the rest is more interesting, and the thing that actually breaks cross-venue analysis isn't price at all.
Finding 1Most of the "dispersion" isn't dispersion
Median spread across all nine venues was 2.78 bps — about $22 on a $78,000 asset. But that number is close to meaningless until you decompose it.
Four of the nine quote BTC against USD. The other five quote against USDT. Those are not the same instrument, and the market knows it: USDT-quoted venues traded above USD-quoted venues in 93.3% of instants, with a median basis of 0.91 bps.
That is not exchanges disagreeing about the price of Bitcoin. It's the price of Tether.
Then there's Bitfinex, which was the cheapest venue in the sample 81% of the time, sitting a median 1.30 bps below the cross-section. Persistent, one-sided, and large enough to dominate the USD group on its own.
| Comparison | Median spread |
|---|---|
| All nine venues | 2.78 bps |
| USD venues only (4) | 2.15 bps |
| USD venues, excluding Bitfinex (3) | 1.32 bps |
| USDT venues only (5) | 1.13 bps |
| USDT/USD basis component | 0.91 bps |
Compare like with like and the dispersion roughly halves. Most of what a naïve nine-venue spread measures is instrument mismatch and one persistent outlier — not disagreement about value.
Finding 2The clock hypothesis fails
The timing differences are real and they're larger than I expected. Median time since last message ranged from 11 ms on Gemini to 496 ms on Gate.io — a 45× spread. Within a single sampling instant, the gap between the freshest and stalest quote had a median of 524 ms and a 95th percentile of 968 ms.
| Venue | Age p50 | p95 | p99 |
|---|---|---|---|
| Gemini | 11 ms | 254 ms | 1,074 ms |
| Bybit | 18 ms | 105 ms | 205 ms |
| Coinbase | 27 ms | 59 ms | 105 ms |
| Bitstamp | 45 ms | 141 ms | 232 ms |
| Binance | 57 ms | 100 ms | 180 ms |
| OKX | 65 ms | 367 ms | 699 ms |
| Bitfinex | 119 ms | 361 ms | 759 ms |
| Crypto.com | 244 ms | 485 ms | 506 ms |
| Gate.io | 496 ms | 960 ms | 1,001 ms |
So the misalignment is undeniable. The question is whether it matters, and there's a direct way to test it: split the instants into the best-aligned quartile and the worst-aligned quartile, then compare their price spreads.
best-aligned quartile (age gap ≤ 362 ms) spread 2.77 bps
worst-aligned quartile (age gap ≥ 769 ms) spread 2.84 bps
─────────────────
difference +0.07 bps (Mann-Whitney p = 0.015)
Statistically significant, economically irrelevant. Doubling the timing misalignment moves the spread by 2.5% of its median value. The correlation between age dispersion and observed spread across all instants is 0.074 — indistinguishable from nothing.
The arithmetic explains why. At the realised volatility in this window, the mid moved a median of 0.16 bps per two-second sample, or roughly 0.08 bps per second. Over 524 ms of misalignment, that's 0.04 bps of price drift — against a 2.78 bps spread. The clock simply isn't fast enough to matter at this horizon.
Which means: if your holding period is measured in seconds or minutes, stop worrying about cross-venue clock alignment and start worrying about which pair you're quoting against. If you're operating at microsecond horizons, none of this applies to you and you weren't reading a browser-based dashboard anyway.
Finding 3The book doesn't agree with itself
This is the one that changed how I think about the problem.
Order book imbalance — (bid − ask) / (bid + ask) over the top levels — is a standard signal, and it was the signal behind the strategy I killed earlier. The natural assumption is that it measures something about the market. It measures something about a venue.
Across eight comparable venues, the sign of the imbalance was unanimous in 7.4% of instants. Median pairwise correlation: 0.195.
The clearest way to see it: take Binance's imbalance as your signal, and ask how often each other venue's book points the same way.
Sign agreement with Binance's book imbalance. The vertical line is 50% — a coin flip.
Gemini agrees with Binance on the direction of book pressure 54% of the time. That is a coin flip with a rounding error. The only strong pair in the whole set is Binance–Bybit at 0.711 correlation, and those two venues share a quote currency, a market-making population, and a great deal of cross-venue arbitrage flow.
If you build an imbalance signal on one venue and describe it as "order flow", be precise about what you have: order flow on that venue. Eight books, eight different answers, most of the time.
Finding 4The books cluster, and not by quote currency
Pairwise correlation of book imbalance across the eight comparable venues. Read it as: how much does knowing one venue's book pressure tell you about another's.
| Binance | Bybit | Crypto.com | Gate.io | Coinbase | Bitstamp | Gemini | Bitfinex | |
|---|---|---|---|---|---|---|---|---|
| Binance | — | 0.71 | 0.55 | 0.35 | 0.27 | 0.35 | 0.16 | 0.24 |
| Bybit | 0.71 | — | 0.53 | 0.30 | 0.27 | 0.28 | 0.14 | 0.13 |
| Crypto.com | 0.55 | 0.53 | — | 0.22 | 0.25 | 0.26 | 0.07 | 0.07 |
| Gate.io | 0.35 | 0.30 | 0.22 | — | 0.15 | 0.17 | 0.13 | 0.17 |
| Coinbase | 0.27 | 0.27 | 0.25 | 0.15 | — | 0.13 | 0.22 | 0.10 |
| Bitstamp | 0.35 | 0.28 | 0.26 | 0.17 | 0.13 | — | 0.04 | 0.13 |
| Gemini | 0.16 | 0.14 | 0.07 | 0.13 | 0.22 | 0.04 | — | 0.14 |
| Bitfinex | 0.24 | 0.13 | 0.07 | 0.17 | 0.10 | 0.13 | 0.14 | — |
Pearson correlation of (bid−ask)/(bid+ask), 1,666 paired observations. With n=1,666, |r| > 0.048 is significant at 5% — so nearly every cell here is distinguishable from zero, and nearly every cell is also small enough not to matter. OKX excluded: my collector reads five levels there and ten elsewhere.
There is a block in the top-left corner and essentially nothing anywhere else.
| Block | Mean pairwise r |
|---|---|
| Within USDT venues | 0.443 |
| Across the two groups | 0.194 |
| Within USD venues | 0.127 |
My first guess was that venues would cluster by quote currency. That is not what happened. The USD venues correlate less with each other (0.127) than they do with the USDT group (0.194) — they are, for practical purposes, eight independent books wearing four different names.
The cluster is specific: Binance, Bybit and Crypto.com, all mutually above 0.53, with Binance–Bybit at 0.71. Gate.io quotes in USDT too and sits outside it, correlating 0.35 with Binance and less with everyone else.
I can't prove the mechanism from quote data alone, but the shape is consistent with shared liquidity provision — the same market-making inventory quoted across several venues, so the books move together because they are, in part, the same book. Whatever the cause, the practical consequence is direct: Binance and Bybit are close to one observation, not two. Averaging their imbalance to build a "market-wide" signal double-counts one venue's flow while treating Gemini, which is nearly orthogonal to everything, as equally weighted noise.
Is one hour enough to say this?
Fair question, so I split the window in half and recomputed the matrix on each 28-minute segment independently. The two matrices correlate at 0.835 across the 28 unique pairs, with a median absolute difference of 0.081. Binance–Bybit was 0.722 in the first half and 0.699 in the second.
So the structure reproduces within the sample. That is not the same as reproducing next Tuesday, or in a selloff, and I would not build anything load-bearing on a single afternoon. But it's stable enough that the block isn't an artifact of one noisy stretch.
CaveatSome of this was my own fault
Depth is where the comparison gets treacherous, and I want to be exact about how much of that is the venues and how much is me.
Median top-of-book depth ranged from 0.88 BTC on OKX to 4.67 BTC on Binance. Within a single instant, the ratio between the deepest and shallowest venue had a median of 9.5× and reached 79×.
But that number is contaminated, because my own collector reads five levels from OKX and ten from everywhere else — a leftover from how I wrote the OKX handler months ago. I found it while auditing this dataset, not before publishing it. So OKX's depth is not comparable to the rest, and I've excluded it from the imbalance comparison above for that reason.
Which is the practical lesson, and it applies well beyond my code. Venues don't agree on what a book snapshot is:
- Different level counts. Some feeds give you 5, some 20, some the full book. Sum "the top of book" across them and you're summing different windows.
- Different update semantics. Some venues push full snapshots, some push incremental diffs you must apply yourself. Get the sequencing wrong on an incremental feed and your book silently drifts from reality.
- Different aggregation. Price-level aggregation versus per-order granularity changes what a "level" even means.
None of this raises an error. You get numbers back and they look like depth. The only defence is to normalise on something the venues actually agree on — depth within a fixed distance of the mid, say, rather than a fixed number of levels — and to state your window explicitly whenever you publish a depth figure.
MethodWhat this is and isn't
Honest limits, because a single-window measurement is easy to over-read:
- One 56-minute window on one afternoon, in calm conditions. The magnitudes will differ in stress; the structure — basis, persistent outliers, book disagreement — is likely more stable than the levels.
- Staleness is measured in the browser, so it's a lower bound. It excludes network latency and exchange-side delay. The real ages are worse, which strengthens the timing result rather than weakening it — the effect was negligible even measured generously.
- Kraken failed to connect during the run, so nine venues rather than ten.
- No trade-through analysis. I measured quotes, not executions. Whether the dispersion is capturable after fees is a different question, and the fee arithmetic suggests the answer is usually no.
What I'd tell someone about to build multi-venue tooling, in the order I'd tell it:
Check the quote currency before anything else. A third of the dispersion I set out to explain was USDT, not disagreement.
Check your own level windows. I had a 5-versus-10 mismatch in my own code and didn't find it until I looked hard at the output.
Don't call single-venue imbalance "order flow." At an hour of data, eight books agreed on direction 7% of the time.
And the meta-lesson, which cost me the original thesis of this post: I was confident enough in the clock hypothesis to plan an article around it. Fifty-six minutes of recording was enough to show it explains 3% of the thing I thought it explained. That measurement took an afternoon. The strategy built on the unmeasured assumption took weeks.