Latency Numbers Every Programmer Should Know
The classic latency table, with human-scale comparisons.
Jeff Dean's rough latency figures date to 2001. Peter Norvig published an early table to show how long memory, disk, and network operations take. This later version adds SSD and datacenter examples; its values are historical estimates, not current benchmarks.
The classic comparison
For the last column, imagine one nanosecond of computer time as one second of human time. The clock has simply been stretched by a billion; the ratios stay the same.
| Operation | Approximate computer time | At human scale |
|---|---|---|
| L1 cache reference | 0.5 ns | 0.5 seconds |
| Branch misprediction | 5 ns | 5 seconds |
| L2 cache reference | 7 ns | 7 seconds |
| Mutex lock and unlock | 25 ns | 25 seconds |
| Main memory reference | 100 ns | 1 minute 40 seconds |
| Compress 1 KB with Zippy | 3 µs | 50 minutes |
| Transfer 1 KB at 1 Gbit/s | 10 µs | about 2 hours 47 minutes |
| Read 4 KB randomly from an SSD | 150 µs | about 1.7 days |
| Read 1 MB sequentially from memory | 250 µs | about 2.9 days |
| Round trip within one datacenter | 500 µs | about 5.8 days |
| Read 1 MB sequentially from an SSD | 1 ms | about 11.6 days |
| Seek to a new location on a hard disk | 10 ms | about 116 days |
| Read 1 MB sequentially from a hard disk | 20 ms | about 231 days |
| Packet round trip: California → Netherlands → California | 150 ms | about 4.8 years |
Here ns means nanosecond (one billionth of a second), µs means microsecond (one millionth), and ms means millisecond (one thousandth). A thousand nanoseconds make a microsecond; a thousand microseconds make a millisecond.
The table does not describe fourteen interchangeable operations. A cache reference retrieves a small piece of data; a sequential read moves a megabyte. The 1 Gbit/s row estimates the time to put bytes onto a link at that rate, not the time for a network request to travel to a server and return. Even the apparent precision of “0.5 ns” should be read as a scale, not a promise.
Why keep the numbers in your head?
The expensive boundary is often the one you cross repeatedly. At the listed datacenter estimate, one round trip is about 0.5 ms; one hundred sequential round trips would already cost about 50 ms before the work at either end. An extra memory lookup and an extra remote call may look equally small in source code, but they inhabit different parts of this table.
The same distinction explains why access patterns matter. A sequential read can keep a device busy moving data; scattered reads may repeatedly pay a fixed access cost. Batching, caching, and avoiding unnecessary trips across a process or network boundary can matter more than shaving a few instructions off a loop. Whether a particular change helps depends on the real workload, so measure it.
Read it as a historical scale chart
Colin Scott's interactive latency chart shows how the estimates change over time. SSDs, CPUs, networks, cloud architecture, and storage interfaces have all changed. A number from this chart should not be pasted into a capacity plan as though it were a current device specification.
What survives is the habit: identify whether you are waiting on a cache, memory, storage, or another machine; distinguish latency (time to get a result) from bandwidth (data moved per unit of time); count how many waits lie on the critical path; then benchmark the actual system. The table is a memorable starting point for asking the right question.
Further reading
- Peter Norvig’s original timing exercise and answers
- Jonas Bonér’s “Latency Numbers Every Programmer Should Know” gist, the source of this table’s approximate values
- Colin Scott on latency trends and his interactive chart
- Google SRE’s latency rules of thumb, another way to keep the relative scales close at hand