Shedletsky's Bits: A... Blog?

A random walk through manifold space
Clipping

Latency Numbers Every Programmer Should Know

The classic latency table, with human-scale comparisons.

Jeff Dean's rough latency figures date to 2001. Peter Norvig published an early table to show how long memory, disk, and network operations take. This later version adds SSD and datacenter examples; its values are historical estimates, not current benchmarks.

The classic comparison

For the last column, imagine one nanosecond of computer time as one second of human time. The clock has simply been stretched by a billion; the ratios stay the same.

Operation Approximate computer time At human scale
L1 cache reference 0.5 ns 0.5 seconds
Branch misprediction 5 ns 5 seconds
L2 cache reference 7 ns 7 seconds
Mutex lock and unlock 25 ns 25 seconds
Main memory reference 100 ns 1 minute 40 seconds
Compress 1 KB with Zippy 3 µs 50 minutes
Transfer 1 KB at 1 Gbit/s 10 µs about 2 hours 47 minutes
Read 4 KB randomly from an SSD 150 µs about 1.7 days
Read 1 MB sequentially from memory 250 µs about 2.9 days
Round trip within one datacenter 500 µs about 5.8 days
Read 1 MB sequentially from an SSD 1 ms about 11.6 days
Seek to a new location on a hard disk 10 ms about 116 days
Read 1 MB sequentially from a hard disk 20 ms about 231 days
Packet round trip: California → Netherlands → California 150 ms about 4.8 years

Here ns means nanosecond (one billionth of a second), µs means microsecond (one millionth), and ms means millisecond (one thousandth). A thousand nanoseconds make a microsecond; a thousand microseconds make a millisecond.

The table does not describe fourteen interchangeable operations. A cache reference retrieves a small piece of data; a sequential read moves a megabyte. The 1 Gbit/s row estimates the time to put bytes onto a link at that rate, not the time for a network request to travel to a server and return. Even the apparent precision of “0.5 ns” should be read as a scale, not a promise.

Why keep the numbers in your head?

The expensive boundary is often the one you cross repeatedly. At the listed datacenter estimate, one round trip is about 0.5 ms; one hundred sequential round trips would already cost about 50 ms before the work at either end. An extra memory lookup and an extra remote call may look equally small in source code, but they inhabit different parts of this table.

The same distinction explains why access patterns matter. A sequential read can keep a device busy moving data; scattered reads may repeatedly pay a fixed access cost. Batching, caching, and avoiding unnecessary trips across a process or network boundary can matter more than shaving a few instructions off a loop. Whether a particular change helps depends on the real workload, so measure it.

Read it as a historical scale chart

Colin Scott's interactive latency chart shows how the estimates change over time. SSDs, CPUs, networks, cloud architecture, and storage interfaces have all changed. A number from this chart should not be pasted into a capacity plan as though it were a current device specification.

What survives is the habit: identify whether you are waiting on a cache, memory, storage, or another machine; distinguish latency (time to get a result) from bandwidth (data moved per unit of time); count how many waits lie on the critical path; then benchmark the actual system. The table is a memorable starting point for asking the right question.

Further reading