Benjamin Moyer

Two hundred packets changed my mind

Ten pings said Buffalo. Two hundred said the tail was 523 ms and the averages were tied. A short argument for larger samples.

I was choosing between two VPS locations and did what everyone does: ten pings to each, compare the averages, pick the lower one.

minavgmaxmdev
Buffalo18.520.325.21.9
Chicago24.326.139.14.3

Clear enough. Buffalo by 5.8 ms, and Chicago has a 39 ms outlier besides. Except ten samples is nothing, and the outlier was the only hint that the distributions might not be what the means suggested.

Two hundred packets each, same wired connection, minutes apart:

minavgmaxmdev
Buffalo18.324.2523.835.7
Chicago24.025.065.53.2

The averages are tied. Buffalo’s 5.8 ms advantage was an artifact of a sample too small to contain its own tail — and that tail is a half-second stall on a wired university connection, not a cellular hiccup. Chicago’s mean sits one millisecond above its floor, which is what a clean path looks like.

The mechanism turned up in mtr: Buffalo’s route hands off at an interconnect whose worst-case was 92 ms against a best of 18.5, with the highest standard deviation of any hop in either trace. Chicago’s equivalent handoff was clean.

For the thing I was actually building — a VPN endpoint — this matters more than the means do. Mean latency does not determine how a tunnel feels; tail latency does. A path that is 18 ms most of the time but occasionally stalls for half a second produces visible hitches in an interactive session, and TCP inside the tunnel reacts to every one of them.

Ten samples would have bought me the wrong box.

No comments here, by design. If you have something to say — a correction, a different edition of the same source, results of your own — I would like to read it: hello@bmoyer.net.