The bug that taught me this took two days to find. A reply, in the logs, timestamped eleven milliseconds before the message it was replying to. Both services were healthy. Both were synced to the same NTP pool. The code was fine.
The clocks just disagreed, and nothing in the system had ever promised they wouldn’t.
Two clocks, and only one of them measures anything
This is the part worth internalising even if you read nothing else.
Your machine exposes two different notions of time, and they are not variations on a theme — they answer different questions and have different guarantees.
The wall clock answers “what time is it?” It’s the one that maps to a calendar, the one you put in a log line, the one a human can compare against their watch. It is also adjustable. NTP corrects it. An administrator corrects it. A VM resuming from suspend corrects it. It can move backwards, and on a long enough timeline it will.
The monotonic clock answers “how much time has passed?” It counts upward from some arbitrary origin, usually boot. It has no idea what day it is and cannot tell you. In exchange, it never goes backwards.
So this is a latent bug:
import time
start = time.time() # wall clock
do_the_work()
elapsed = time.time() - start # can be negative. genuinely.
And this is not:
import time
start = time.monotonic()
do_the_work()
elapsed = time.monotonic() - start # always >= 0
If NTP steps the clock backwards between those two calls, the first version reports a negative duration. Most code doesn’t crash on that — it quietly records a nonsense number, or the elapsed > timeout check reads false forever, and you find out months later when a timeout stops firing.
The rule is short enough to remember: wall clock for communicating a moment, monotonic for measuring an interval. Nearly every language has both. Go has time.Now() carrying a monotonic reading that Sub uses automatically. Java has System.nanoTime() next to currentTimeMillis(). Rust has Instant and SystemTime, and makes you handle the error case on the second one, which is a nice bit of API design.
What NTP actually promises
Less than people assume.
NTP keeps your machines loosely synchronised. On a well-behaved network with a nearby stratum-2 server you’ll typically sit within a few milliseconds of true time. That is usually fine. The trouble is the words “typically” and “usually,” because the failure cases are not rare:
Network paths are asymmetric. NTP estimates offset by assuming the trip out takes as long as the trip back, so a congested return path quietly biases the correction. A loaded host delays the timestamping itself. A VM that was suspended wakes up badly wrong and gets stepped to correct. And an upstream server that goes unreachable means you drift on whatever the local crystal does, which on commodity hardware is tens of parts per million — seconds per day.
The deeper issue is that synchronising is itself a source of discontinuity. NTP corrects either by slewing (speeding up or slowing down the clock until it catches up, smooth but slow) or by stepping (jumping it, instant but discontinuous). Both are visible to your code. Slewing means a measured second isn’t a real second. Stepping means the jump I described above.
So: NTP makes log timestamps from different machines roughly comparable. It does not let you order two events that happened within a few milliseconds of each other on different machines. Nothing you can deploy does, short of the GPS-and-atomic-clock arrangement Google built for Spanner — and notably, even then they didn’t claim exact time. They exposed the uncertainty as an interval and made the database wait it out.
That’s the honest framing. You can’t eliminate skew. You can only know its bound, or stop depending on it.
Ordering things without asking what time it is
Here’s the reframe that makes the problem tractable. Most of the time you don’t actually care what time an event happened. You care which of two events came first — and that’s a different question with a better answer.
If event A caused event B, there’s a chain of messages connecting them. You can track the chain directly instead of inferring it from timestamps.
A Lamport clock is the minimal version: one integer per process.
on local event: counter += 1
on send: counter += 1; attach counter to message
on receive(msg): counter = max(counter, msg.counter) + 1
That’s the whole algorithm. It gives you a useful guarantee: if A happened-before B, then A.counter < B.counter. Events that genuinely influenced each other come out in the right order, with no clocks involved at all.
What it can’t do is the converse. If you see two events with counters 7 and 12, you cannot tell whether the first caused the second or whether they happened on opposite sides of the system with no connection. Vector clocks fix that by keeping a counter per process rather than one total, so you can compare two vectors and get a real answer: A before B, B before A, or concurrent — meaning neither saw the other, and your application has to decide what that means.
I’ll be straight about the practical reality here, because textbooks are coy about it. Vector clocks are correct and they are a nuisance. The vector grows with the number of processes, it rides along on every message, and pruning it safely is fiddly. Plenty of systems have adopted them and later backed out. If you’re reaching for one, it’s usually a signal that the real fix is making the operation commutative or idempotent so ordering stops mattering — which is the approach I’d try first, and it’s why idempotency keys do so much work in practice.
Where they genuinely earn their keep is conflict detection in replicated stores — determining that two writes were concurrent so you can surface both versions rather than silently picking one by timestamp. Picking by timestamp is last-write-wins, and last-write-wins with skewed clocks means the winner is whichever machine’s clock happened to be running fast. That’s not a conflict resolution policy. That’s a coin flip you’ve chosen not to look at.
Where wall-clock time is load-bearing anyway
You can’t banish it entirely. Several things genuinely depend on civil time, and they’re mostly in security:
TLS certificates have validity windows. JWTs carry exp and nbf. Kerberos tickets expire. OAuth tokens expire. Signed URLs expire. Every one of those is a wall-clock comparison across two machines that do not agree, which is exactly why clock skew shows up as an authentication bug far more often than as a data bug. A machine whose clock is five minutes fast will reject tokens that are perfectly valid, and the error it reports will be about the signature, not the time. This is also why most token validators allow a small configurable skew tolerance — typically a minute or two. That tolerance is not sloppiness. It’s an admission.
Then there’s the one that actually corrupts data: leases and locks.
A lease says “you hold this for 30 seconds.” The holder checks the clock, sees it has 20 seconds left, and proceeds to write. The trouble is that a lease expiring and a process noticing the lease expired are different events, and the gap between them can be arbitrarily long. A stop-the-world GC pause, a hypervisor descheduling the VM, a slow disk — any of these can suspend a process past its expiry. It wakes up, finishes what it was doing, and writes. The lock was reassigned eleven seconds ago. Both writes land.
No timeout fixes this, because the paused process cannot observe its own pause. The fix is a fencing token: a number that increments on every lease grant. The holder sends it with every write, and the resource refuses anything below the highest it has seen.
lease granted to A -> token 33
A pauses (GC)
lease expires, granted to B -> token 34
B writes with token 34 -> accepted, resource now at 34
A wakes, writes with token 33 -> REJECTED, 33 < 34
The stale write is rejected on arrival, and no clock was consulted to do it. That’s the pattern: when correctness is on the line, replace the time comparison with a monotonically increasing number that the resource checks. It’s the same instinct behind write-ahead logging — make the durable thing the arbiter, not the thing that might be lying.
What I actually do
Short list, no ceremony:
Use monotonic clocks for every duration, timeout, and retry interval. There’s no case where the wall clock is the better choice for that and several where it’s a bug.
Log in UTC with an explicit offset, always. Local time in logs is a tax you pay later, at 3am, during an incident, while mentally applying DST rules.
Never order events from two machines by comparing their timestamps. If ordering matters, carry a sequence number — or restructure so it doesn’t matter.
Treat last-write-wins as a decision, not a default. If you can’t say out loud which clock you’re trusting and how far off it might be, you haven’t chosen a policy.
Alert on skew directly. Most NTP clients expose their current offset; scrape it, and page when it exceeds what your token validators tolerate. This moves a confusing authentication outage into a boring, named alert, and it’s maybe twenty minutes of work.
And put a fencing token on anything where two writers would be a genuine problem. Leases alone don’t give you mutual exclusion. They give you mutual exclusion most of the time, which is a thing that works until it’s the subject of a postmortem.
Related: Idempotency Keys: The Thing That Makes a Retry Safe · Raft: How a Cluster of Machines Agrees on a Single Truth · Timeouts and Deadline Propagation
Comments