farblog

 by Malcolm Rowe

HOWTO: set initcwnd on Google Compute Engine Debian hosts

In a footnote to a previous post, I mentioned that I’d recently done some TCP tuning to increase initcwnd on this website, which runs on Google Compute Engine (GCE from here on out). Configuring this turned out to be surprisingly tricky, so I thought it would be useful to write up how it all fits together and how to do this on GCE.

Note that I’m specifically going to be talking about the Debian 13 (Trixie) cloud images, which configure networking via systemd-networkd managed by Netplan. Netplan isn’t critical here, but if you’re not using systemd-networkd, this post probably won’t be that useful.

First, what was I trying to achieve?

initcwnd is the TCP initial congestion window, and controls how much data a server can transmit at the start of a connection before it needs to wait for an acknowledgement from the client.

Linux 2.6.39+ defaults to an initcwnd of 10 segments1 (up to 10 × MSS bytes). On GCE, where the default virtual network MTU is 1460 bytes, this allows transmitting almost 14 KiB in a single go. For a webserver, this means that if an HTTP response (including headers) is larger than that, the server must pause partway to wait for the client to send a TCP ACK, and only then continue with the remainder of the response.

Increasing initcwnd to 20 segments is therefore a quick way to reduce latency for small (approximately, 14–28 KiB) responses, since it effectively cuts their transfer time in half. For most webservers, there’s quite a lot that fits in that size range: for example, while this post by itself is comfortably smaller than 14 KiB (assuming compression), my front page (compressed) will typically be just a little larger than that threshold.

The hard limit for the initial TCP window is 65535 bytes, which in practice would be about 44 segments2. However, setting initcwnd higher than about 20 is usually not recommended: sending too many packets can overflow intermediate routers’ receive buffers (leading to packet loss — and packet loss during the initial TCP connection can stall the connection for seconds), and can also contribute to bufferbloat, potentially causing latency spikes for other traffic. 20 is a ‘reasonable’ figure that noticeably improves latency without triggering any of those side effects.

Note that this does assume that the ultimate receiver can actually handle this much data: some very old mobile clients and modern IoT clients (and e.g. Windows XP) will advertise a smaller receive window (their initrwnd), which will cap the size of the initial window.

Anyway, that’s why I wanted to configure initcwnd on this webserver (also, because it’s fun). The rest is how to do so.

As is common, on the GCE version of Debian, the internal IPv4 address is set by DHCP. As is slightly less common, DHCP also provides the server with default and gateway routes (per RFC 3442).

In modern versions of systemd, systemd-networkd can be told how to set initcwnd via an InitialCongestionWindow setting in a .network file.

We can see which .network file systemd-networkd is using with networkctl status (which accepts interface names or shell-style wildcards):

$ networkctl status 'en*'
● 2: ens4
                   Link File: /usr/lib/systemd/network/99-default.link
                Network File: /run/systemd/network/10-netplan-all-en.network
                       State: routable (configured)
...

(or networkctl with no arguments will list all the network links that systemd-networkd is aware of.)

Network File: is the important part here: it shows the name of the .network file that we’d need to change.

However, here we have a dynamic file (in /run/) generated by Netplan, so we can’t modify it directly. Fortunately, systemd-networkd also supports the idea of “drop-in files” that can be used to add additional configuration to an existing network; see man systemd.network.

For example, for the network link above, we could put extra configuration into /etc/systemd/network/10-netplan-all-en.network.d/*.conf, where the name of the .d directory has to match the name of the network file.

So here we’d just need to add a [DHCPv4] section with InitialCongestionWindow=20 to achieve what we want to do.

This approach does rely on the Netplan-generated filename staying static, but that should be the case: it seems to use a hardcoded 10-netplan- prefix, plus a string from the /etc/netplan/ configuration, which should be unlikely to change. Of course, if you’re not using Netplan, it’s even easier, since you just edit the .network file directly.

So IPv4 is pretty simple, once we know what file to create (I’ve put a complete example below), but IPv6 is a little more mysterious.

For IPv6, you might expect to be able to add a similar setting under one of the IPv6-related sections, but the specific setting we’re after can only be specified under [Route] and [DHCPv4], at least as of systemd 2613.

We do have a solution, thankfully, but it requires understanding how routing is configured for dynamic IPv6:

The IPv6 address is assigned by stateful DHCPv6, where systemd-networkd runs a client to lease the public /128 address from its allocated subnet. GCE does not use SLAAC for address assignment. However, DHCPv6 only provisions addresses, not routes.

The IPv6 route is assigned by watching for periodic ICMPv6 Router Advertisement broadcasts4.

When an RA packet arrives, systemd-networkd processes it and normally installs the default route. RA packets only supply IPv6 routing parameters (and we can’t change what GCE sends anyway), and as noted above, systemd-networkd doesn’t have any obvious place to attach route metrics like initcwnd to the result.

Fortunately, we can switch off the automatic creation of that default route and instead configure a route of our own that will become the default route.

To cut to the chase, we can set initcwnd for both IPv4 and IPv6 by creating a single drop-in file with a name matching the network file shown by networkctl status:

# /etc/systemd/network/10-netplan-all-en.network.d/initcwnd.conf
[DHCPv4]
InitialCongestionWindow=20

[IPv6AcceptRA]
UseGateway=no

[Route]
Destination=::/0
Gateway=_ipv6ra
InitialCongestionWindow=20

Setting UseGateway=no stops systemd-networkd from creating its own default route when receiving the RA; instead, it finds the [Route] that defines Gateway=_ipv6ra, and installs that route using the advertised gateway. Since our route also defines Destination=::/0, it becomes the default route.

While my focus was on initcwnd, you can obviously also use this to set any other per-route settings, such as InitialAdvertisedReceiveWindow (initrwnd). Some other per-route settings can already be set directly in the [IPv6AcceptRA] section, which is probably a better place for them. (If a future version of systemd adds support for the missing settings to [IPv6AcceptRA], the above would become a little simpler.)

sudo networkctl reload reloads the configuration and updates the IPv4 route immediately from the cached DHCP lease. For IPv6, it’ll drop the default route (the one created by UseGateway=yes) and request a new RA by sending out a Router Solicitation packet. That should trigger a new RA, installing the new IPv6 default route within a second. (Alternatively, you can just reboot the VM.)

Once everything’s been updated, we can see the results with ip route:

$ ip -4 route show default
default via 10.154.0.1 dev ens4 proto dhcp src 10.154.15.203 metric 100 initcwnd 20

$ ip -6 route show default
default nhid 886644799 via fe80::4001:aff:fe9b:1 dev ens4 proto ra metric 100 expires 85sec initcwnd 20 pref medium

The two via addresses are the local gateways, and proto dhcp and proto ra show what protocols provided the routes: DHCP for the IPv4 route, and a Router Advertisement broadcast for the IPv6 route. Both now have initcwnd 20, which was what I wanted to achieve.

Well, it took a bit of digging through man pages to figure out what was going on here, but the end result is that I now have a much better understanding of modern networking clients, and of IPv6 — which was really the main reason I wanted to look into this in the first place. Writing it up here should also give me a place to look the next time I want to remember what’s going on.

Oh, and this server is very slightly faster now too.


  1. 10 segments is the recommended value from RFC 6928, at least if we assume non-jumbo frames. 

  2. The maximum number of segments in the initial window is just ⌊65535/MSS⌋, and the MSS is the MTU minus either 20 (IPv4) or 40 (IPv6) bytes for the IP header, 20 bytes for the TCP header, and usually 12 bytes for TCP timestamps, so the largest common MSS is 1460 bytes (1500 byte MTU, IPv4, no timestamps), with at most 44 segments in the initial TCP window. For GCE, the smallest common MSS is 1388 bytes (1460 byte default MTU, IPv6, timestamps), leading to at most 47 segments. 

  3. It would make sense for systemd-networkd to add support for route-related settings like InitialCongestionWindow to [IPv6AcceptRA]; some, like QuickAck, are already there, so this just looks like an omission. 

  4. Technically, Router Advertisements are sent to the link-local multicast address, so I guess “broadcast” is slightly inaccurate, but I’m not going to noun “multicasts”. 

Unhappy eyeballs

Here’s the tl;dr (also, spoilers): if you’re using Prometheus’ blackbox exporter to probe a target, you should probably explicitly specify the IP protocol (IPv4 or IPv6) to use for each probe, especially if you have an IPv4-only target.

I have a very small home monitoring setup using Prometheus and various exporters to monitor this website, plus a few other sites and machines that I care about. For HTTP monitoring, I’m using Prometheus’ blackbox exporter to probe a target URL (the front page of this site) every so often.

The probe configuration is a little bit confusing, as the targets (what to monitor, i.e. the target URL) are set in the main Prometheus config, while the probe definition and expectations (how to monitor) are set in the blackbox exporter config, but the latter looks something like this:

modules:
  farblog_https:
    prober: http
    http:
      headers:
        accept-encoding: br
      compression: br
      fail_if_not_ssl: true
      follow_redirects: false
      fail_if_body_not_matches_regexp:
        - "farblog"

This probe succeeds if the site is responding reasonably, and each probe attempt also records various metrics that I can use elsewhere (probe_ssl_earliest_cert_expiry, probe_success, etc).

The automated checks I’m running are pretty basic: verifying that the site is up, that the TLS certificate isn’t about to expire, that kind of thing, but I also have some basic monitoring for probe latency, mostly to detect accidental regressions (or just let me know if something weird is going on).

In the past I’ve used this historical record to test the impact of server changes (for example, I now pre-generate compressed gzip and Brotli representations of most content, which saves a few milliseconds), but as with most things here, it’s just a fun little project to play with.

For some time though, I’d been aware that the response times I was seeing were a little spiky, which bothered me: my home internet connection should be fairly stable, and the site itself (which is running nginx on a small GCE node) should be very predictable, and yet my graphs still looked like this:

Line chart of HTTP probe duration over 24
hours. The metric sits at a noisy baseline of roughly 40–50ms, interrupted
by frequent, dense spikes reaching 80+ms across the entire day.
Probe duration over a 24-hour period: a steady baseline punctuated by mysterious latency spikes.

I’d assumed that this was down to something in my home networking setup that I hadn’t taken account of, but then more recently I noticed that most of the probes to other sites were much more consistent; so what was the difference with this one?

Zooming in to show an hour rather than the whole day (the highlighted part of the graph above) initially didn’t help much, but breaking down the probe duration by the different probe stages (the probe_http_duration_seconds metric) does show something interesting:

Stacked line graph showing a 1-hour window
of probe duration broken into stages: resolve, connect, tls, processing, and
transfer. Most stages stay uniform, while the resolve stage shows sharp
upward spikes.
Probe duration over a one-hour period, broken down by probe stage.

This graph is fairly unsurprising: as a static site, most of the time is determined by the number of roundtrips we need to make from the probe origin to the target1, but the resolve stage is clearly the issue here.

Isolating just resolve makes this even clearer:

Line chart isolating the DNS resolve phase
over a single hour. Latency hovers at a flat baseline around 1–2ms, jumping
to 7ms at almost exactly five-minute intervals.
Probe duration over a one-hour period, “resolve” stage only.

This is why my probe latencies were spiky: address resolution usually takes a little over a millisecond, but every so often, it takes about 7ms instead. What’s going on here?

1–2ms is about2 the time it takes to resolve an address that’s cached by my router, while 7ms is about3 the time it takes to resolve the target address via my upstream nameservers, so something was clearly making extra queries to my upstream.

My router should definitely be caching the DNS results, and while we absolutely would expect to see a spike when that cached result expires, that should manifest as an extra delay once every 12 hours (the TTL for www.farside.org.uk) — not what we’re seeing here.

One possibility I considered was that my router was prematurely evicting results from its cache, but it’s pretty clear from the graph that “every so often” is actually “every five minutes”, and that’s a big clue: the MINIMUM time for my domain (in the SOA record) is 300 seconds, or five minutes, so this is likely a negative cache result expiring and triggering a query to an upstream nameserver (per RFC 2308). But what non-existent name could it be trying to resolve?

The answer turns out to be IPv6.

At the time I was looking into this, this site only served over IPv4 — GCE didn’t support IPv6 until 2022 or so, and I’d never gone back to turn it on — and so it only had an A record for the IPv4 address, not the AAAA record needed for IPv6.

Somewhat unexpectedly, blackbox exporter’s default behaviour is to resolve both A and AAAA records in parallel, and then only continue once both queries have an answer.4

Since in my case there was no AAAA record to resolve, my upstream nameserver would return a NODATA response, and my router would dutifully cache that for the amount of time specified in the SOA record.

The solution is to tell blackbox exporter to resolve only one address family by disabling ip_protocol_fallback, and then to pick a protocol (or use the default of IPv6) using preferred_ip_protocol:

modules:
  farblog_https:
    prober: http
    http:
      # (as before…)

      # Prefer IPv4, and disable protocol fallback so that we only
      # attempt to resolve an address for one address family.
      preferred_ip_protocol: "ip4"
      ip_protocol_fallback: false

This is also what you’ll want to do if you have an IPv4-only client probing a target that does have an IPv6 address, though in that case it’ll be a lot more obvious, as blackbox exporter will resolve both and then fail to connect via IPv6.

With this change, blackbox exporter only issues a single A record query, which will usually be cached locally. The effect was immediately visible:

Line chart of total probe duration over 24
hours after disabling IPv6 address fallback. The recurring spikes have
completely disappeared, leaving a clean, flat line.
Probe duration over a 24-hour period, after disabling IPv6 address resolution. Without the uncached AAAA queries, the baseline variance decreases and the spikes mostly vanish.

This was also the explanation as to why only some of my probes were noticeably affected by this problem: the others were probing dual-stack targets, where the IPv4 and IPv6 addresses would both exist, and would typically have comparable TTLs. This still wasn’t ideal, since the probes would be waiting for answers to DNS queries they wouldn’t use, but in practice it didn’t have anywhere near the same effect on the probe duration.

I do think that blackbox exporter’s default behaviour is a little unfortunate here, since it abstracts away too many of the details (and until fairly recently, was mostly undocumented). However, the documentation has now been updated to describe the address selection process, and to recommend that dual-stack targets, at least, should have separate IPv4 and IPv6 probes.

Does any of this matter to my little website? Of course not, but I had fun digging into the problem, and learned a fair amount along the way. And maybe this write-up will be useful to someone else too.

Coda

In related news, this finally prompted me to put in the work to enable IPv6 on this site, which turned out be to pretty trivial to do. So now I have two separate sets of latency metrics to look at!

(The title of this post is derived from RFC 8305, the “Happy Eyeballs” algorithm that helps applications to connect to dual-stack hosts. While a low-level prober isn’t the kind of application that would actually want to use this algorithm, I thought the reference was too good to pass up!)


  1. I’m about 7ms away from the machine I’m monitoring, so connect is one roundtrip, and processing plus transfer is two roundtrips, due to the size of the response5. Surprisingly, tls is actually one roundtrip plus client-side processing6, not the two roundtrips that the 16ms measurement would suggest! 

  2. Local address resolution via my router actually seems to take closer to 800μs from the low-powered machine that’s running these probes (according to dig -u); I’m assuming that blackbox exporter (or more likely, the Go runtime) is doing some extra work somewhere to account for the extra time. 

  3. The fact that resolving an address via my upstream nameservers takes almost exactly the same amount of time as the roundtrip time to the machine I’m monitoring is pure coincidence, but confused me quite a bit initially. 

  4. Even more surprisingly, this is actually the default behaviour for Go’s net.Resolver, regardless of whether or not the local system has usable IPv4 and IPv6 addresses. In contrast, glibc and macOS implement getaddrinfo() as if the AI_ADDRCONFIG flag were passed by default, meaning that IPv4-only and IPv6-only clients will only resolve addresses that are likely to be routable. 

  5. processing (TTFB) and transfer (TTLB-TTFB) seemed to be apportioned unpredictably, with occasionally all the time being accounted to processing and almost nothing in transfer. My best guess is that something (task scheduling, client TCP implementation) sometimes caused the prober to not read the first byte until the whole response was ready. I’ve subsequently fixed this by increasing TCP initcwnd so that the server no longer waits for an ACK from the client partway; this means the response only takes a single roundtrip, and this weird response-coalescing behaviour no longer occurs. 

  6. Temporarily disabling client-side certificate verification with insecure_skip_verify reduces the tls time down to one roundtrip plus a millisecond or so of processing. I don’t know why local certificate verification is taking so long. 

Collecting coupons

I’ll freely admit that I don’t need to use a lot of maths1 in my day job: pretty much the only things I really need to understand are basic statistics and, occasionally, how to run a t-test to see whether an optimisation has actually improved things (though I guess I don’t even need to understand the maths for that one).

But sometimes maths ends up happening anyway, and it can be interesting when that happens. Here’s one such story:

Let’s say that we have some frontend service that we’d like to understand the performance of in a production configuration. Maybe we want to understand its RAM usage under load, or the latency for some given types of requests, that kind of thing.

In our example system, we’ll send requests to a single frontend task, but there’s a wrinkle: that frontend uses a sharded or replicated backend RPC service to service these requests, and the frontend will only connect2 to a given backend task when needed. That connection setup will take a little extra time whenever it occurs, which we don’t want to count in our test.

If we keep sending requests, then eventually the frontend task should have connected to all the backend tasks, and we’ll have a steady-state we can start analysing.

The question then is, how long will it take for our system to set up all of these connections and reach that steady-state? Or, more precisely, how many warm-up requests do we need to send before we can start our real measurements?

Let’s say that our frontend has a hundred backend tasks that it can connect to, and that we’ll model the selection of each backend task as uniformly random for each request. Do we need to send a few hundred requests to our frontend before all of those backends are connected, or a few thousand? Or more?

Incidentally, when I did something like this for real, I used a pragmatic solution instead: keep sending requests until the metric I was monitoring looked like it had settled down, then start measuring. But it did make me wonder how we might go about calculating this mathematically.

Back-of-the-envelope maths time: we know that it has to be at least 100 requests, since we can only make at most one connection to a backend task for any given frontend request. We also know that doing it in exactly 100 requests should both be possible and extremely unlikely:

  1. The first request we send will definitely trigger a connection to some backend task, since we haven’t made any connections at all yet.
  2. The second request will probably make another connection. Since we have 100 backend tasks, the chance is only $1 / 100$ that we pick the same backend as we used for the first request, so in $99 / 100$ cases we’ll connect to a new backend task.
  3. The third request will connect to a third backend in $98 / 100$ cases, and so on.

After we’ve sent that 100th request, the chance that we’ve connected to all 100 backends is therefore:

$$ \eqalign{ &\frac{100}{100} \times \frac{99}{100} \times \frac{98}{100} \times \dots \times \frac{2}{100} \times \frac{1}{100}\cr = &\frac{100!}{100^{100}} } $$

… which is a very small number: about 1 in $10^{42}$.

So if we want to have a high chance of having connected to all of the backend tasks, we’ll definitely need more than 100 requests, but it’s still not very clear how many: in theory, it could take an unbounded number of requests, since we might never pick all 100!

It turns out that this problem has a name: it’s the coupon collector’s problem, based on the idea of collecting all of $n$ different coupons by drawing randomly, one-by-one.

Calculating the expected number of requests we need is fairly straightforward: we can just sum the expected number of requests we need to make each new connection:

  1. For the first connection, we always need exactly one request.
  2. For the second connection, the probability of making a new request is $99 / 100$, so the expected number of requests we need is $100 / 99$.
  3. For the third connection, the expected number of requests is $100 / 98$, and so on.

To make the last connection, we expect that it’ll take $100 / 1$ or 100 requests, which makes sense: we have to pick the one unconnected backend from 100 options (in the same way that on average you’d expect to need to make six die rolls on a regular six-sided die in order to roll some pre-chosen number).

So, as Wikipedia tells us, if we calculate3

$$ \sum_{k=1}^{n} \frac{n}{k} $$

for n=100, we get about 518.74, which is the expected number of requests we’d need to send.

But what does this number actually tell us? If we send 519 requests, does that mean that we’ll probably have connected to all of the backends? Or that we have exactly a 50% chance of having done so? How many requests do we need if we want to have a 95% chance of having connected to all of the backends, say?

What we’ve just calculated is the expected value, and is an average of the number of requests.

In other words, if we were to run this experiment a large number of times, counting the number of requests until we successfully connected to all the backends ($C_1, C_2, C_3, \dots$), then the expected value is simply the average of those counts.

Since some of these counts might be quite large, this expected value isn’t actually that useful for thinking about the number of requests we’re going to need to send. What we really want to know is what the actual distribution looks like.

Since randomly picking those last few backends takes quite a few requests, the distribution is heavily skewed to the right. To cut to the chase, it looks like this:

A right-skewed probability distribution graph for the coupon collector's problem with 100 coupons. The curve rises sharply to a peak at 460 attempts, then slowly tapers off into a long tail to the right.
Probability distribution (PMF) for the coupon collector’s problem, for n=100 coupons

For every point $x$ on the x-axis, the height of the curve shows the probability of collecting 100 coupons — or connecting to all 100 backends, in our example — in exactly $x$ attempts, so for $x = 100$ we have the very small (but non-zero) probability we calculated earlier, while for $x \lt 100$ (the hatched area) the probability is zero, as those values are impossible.

The expected value (518.74) we calculated above is marked, as are the median (497) and 95th percentile (754). All of these are larger than the mode (460), which is the most-likely number of attempts we’d need.

This graph also has non-zero values everywhere for $x \geq 100$, though it drops off quickly: the probability that we’d take more than 1000 attempts is only 0.4% or so, for example.

Wikipedia goes into the details of how to calculate this distribution exactly, though I don’t feel competent enough to explain it here. One useful nugget, though, is that we can estimate some of the statistics if we map this distribution to what is apparently called a Gumbel distribution:

For example, for $n=100$ and $p=0.95$, this gives a mode of 460, which is correct, and a 95th percentile of 758, which is pretty close to the real value of 754.

In this example, we could be very sure (~99.57%) that if we were to send about a thousand requests, we’ll have connected to all of our backends.

So, the next time you need to work out how many requests it takes to warm up a cache or fill a connection pool, you could use this to work out some magic numbers. But there’s an important caveat for this specific example: all of the above relies on our initial assumption that picking a backend is uniformly random.

In our non-spherical-cow world, load balancers don’t pick backends at random, and instead use a round-robin or similar strategy that would most likely end up actively picking the unconnected backends, significantly reducing the number of requests we actually need to connect to everything.

Maths gives us an interesting worst-case upper bound, but it also suggests why the pragmatic solution was a better idea: sometimes it’s just easier to watch your latency metrics until they flatten out, and let the maths happen in the background.


  1. or “math”, for y’all in en-US. 

  2. Perhaps you’re wondering whether our hypothetical frontend could just pre-connect to all possible backend tasks upon startup? Sure, it could, but let’s pretend that there are good reasons that it doesn’t. In any case, the same principles here can also apply to other areas like cache warming, for example, so let’s just go with the maths for now. 

  3. We can actually compute a closed-form approximation of the expected value by way of harmonic numbers. It turns out that the exact value we want is $n H_n$, where $H_n$ is the nth harmonic number. We can approximate $H_n$ as $\ln{n} + \gamma + \frac{1}{2n}$, where $\gamma$ is Euler’s constant, approximately $0.577$, and so multiplying that approximation of $H_n$ by $n$ produces a reasonable approximation of the expected value for the coupon collector’s problem. 

A moderately deep dive into filesystem times on Linux

This last weekend, I spent a little time noodling around with the static site generator that I wrote to generate this website. One thing I remembered was that a few of its tests were slightly flaky, and this time I really wanted to dig into what was going on.

One such flaky test did the following:

  1. Get the current time, as start.
  2. Do some processing that should update a file.
  3. Get the current time again, as end.
  4. Expand the range of start and end so that they fall on second boundaries, because “some filesystems can only store modification times to the second”.
  5. Check that the target file’s modification time falls within [start, end).

Perhaps you already know where this is going, but while the test above might work on a strictly POSIX-compliant system (more on that later), it doesn’t work in practice on Linux, occasionally failing because the modification time we see in step 2 is slightly before the start time we obtain in step 1, albeit just by a few milliseconds!

Exactly why this happens is down to how Linux updates file modification times when a file is written to.

I’m going to dig into this below, but to avoid making this any longer, I’m also going to limit myself a bit, and make some simplifying assumptions:

First, a little history

Let’s start our story back in 1979.

While earlier Unixen had creation and modification times, it was V7 Unix that introduced the “atime”, “mtime”, and “ctime” names that later made their way into the POSIX standard. The original1 POSIX standard defined these as:

time_t    st_atime   Time of last access. 
time_t    st_mtime   Time of last data modification. 
time_t    st_ctime   Time of last status change. 

I think the first two are pretty clear, but ctime is a little confusing: from the “ctime” name, you might have assumed it was “creation time”, but as you can see from above, it’s a “status change” time. Effectively, mtime is the time that the file contents changed, while ctime is the time the file metadata changed (and since the metadata includes the mtime, any update to mtime must also update ctime).

While POSIX doesn’t include any concept of file creation time, some operating systems and filesystems do. For example, ext2 does, and Linux has a statx() system call that returns an extra btime field with a file’s creation time (“btime” from “birth time”, following prior usage in BSD).

What POSIX says about how file times are updated is actually quite interesting. Rather than just specifying that modification times are updated when a file is written to, the various times are instead “marked for update” whenever certain specific operations complete, so (e.g.) a successful write() will mark mtime and ctime as to-be-updated.

Implementations are then free to actually run that update later on (“At an update point in time, any marked fields shall be set to the current time”, says POSIX), subject to certain operations that must trigger an update (calling stat(), for instance).

All that seems okay for our test, though? While we might not see exactly the right modification time, we should still see a time that sits within the range we’re expecting? Except we don’t.

As far as I can tell, the reason for that is that Linux doesn’t strictly follow POSIX here, though I must admit that it’s not completely clear: POSIX does leave quite a lot of room for ambiguity anyway2.

What times can filesystems actually store?

Before we talk about Linux in general, it’s probably worth talking about the difference between granularity (what the filesystem can store) and accuracy (how close the stored time is to real time), since that’s something that’s confused me in the past.

Different filesystems support different granularities for the times that they can store. To pick a few examples:

To be extremely specific about ext2/3/4 for a moment: the on-disk representation supports nanosecond resolution (and dates after 2038) if the filesystem was created with an inode size of 256 bytes rather than 128 bytes (as shown in tune2fs -l output).

If you have an ext2/3/4 filesystem, it’s almost certainly using 256-byte inodes already: they became the default in 2008 when creating a filesystem of 512MiB or larger, and nowadays are used by default regardless of size. Modern Linux kernels3 support whatever the filesystem supports.

Incidentally, this does make me wonder whether I actually had access to an older filesystem when I was writing the test I mentioned at the start of this post, or whether I’d just misunderstood why the times I was seeing in my test didn’t match up with what I expected4.

So if some filesystems only support certain granularities, what happens if we try to set a finer-grained value ourselves? In this case, the kernel will truncate the value to what the filesystem supports first (see s_time_gran in the kernel source, which is how each filesystem declares its granularity). This truncation (rather than round-to-nearest, say) is also what POSIX requires.

We can see this behaviour easily by creating a file with a fixed modification time on a few different filesystems.

First, ext4. As mentioned above, this supports nanosecond granularity, so the value we specify is used directly for the access and modification times (while the current time is recorded for the status change and file creation times):

$ touch foo -d '1985-10-26T01:21:03,123456789'; stat foo
[...]
Access: 1985-10-26 01:21:03.123456789 -0700
Modify: 1985-10-26 01:21:03.123456789 -0700
Change: 2025-11-27 02:11:45.779834694 -0800
 Birth: 2025-11-27 02:11:45.779834694 -0800

If we explicitly create the ext4 filesystem with 128-byte inodes, we can see that we only have times stored to seconds resolution (and that we no longer have a creation time):

Access: 1985-10-26 01:21:03.000000000 -0700
Modify: 1985-10-26 01:21:03.000000000 -0700
Change: 2025-11-27 02:16:14.000000000 -0800
 Birth: -

While for something like FAT, we get this mix of granularities:

Access: 1985-10-25 17:00:00.000000000 -0700
Modify: 1985-10-26 01:21:02.000000000 -0700
Change: 1985-10-26 01:21:02.000000000 -0700
 Birth: 2025-11-27 02:17:57.490000000 -0800

On FAT filesystems, access times use day granularity (the value above is midnight UTC for the time I gave), modification times use two-second granularity (FAT doesn’t store a separate ctime, so ctime is always equal to mtime), while creation times use 10ms granularity.

Back to ext4: if we create a file with the current time, we can also see the weird behaviour I mentioned right at the start:

$ date --iso-8601=ns; touch foo; stat foo
2025-11-27T02:22:03,520075379-08:00
  File: foo
[...]
Access: 2025-11-27 02:22:03.518535805 -0800
Modify: 2025-11-27 02:22:03.518535805 -0800
Change: 2025-11-27 02:22:03.518535805 -0800
 Birth: 2025-11-27 02:22:03.518535805 -0800

The recorded times are 1.5 milliseconds before the time printed by the preceding date command! What’s going on?

What’s the time, Mister Linux?

So what about accuracy? How does Linux choose what mtime value to write when a file is updated, and why does it seem like time is going backwards here?

On the face of it, this seems like an odd question: why wouldn’t the kernel just set the modification time equal to the current time every time a file is written? I think that’s what most people (myself included) would have assumed was happening.

The main reason (and probably the only reason, as far as Linux is concerned) is efficiency: files are written to a lot, and recording a new modification time on every write would potentially mean doing a lot of extra work, even if everything stays in cache. Not only that, but merely fetching the current time can be surprisingly expensive5, especially on very old hardware without access to CPU cycle counters.

Instead, Linux (up until 6.13) maintains a ‘coarse’ current time that updates every timer ‘tick’ (usually once every 4ms or 10ms6), and then all writes made (to any file) during that interval will record exactly the same modification time, even if the filesystem could store a finer-grained value.

This is absolutely the explanation for my test failures above: the modification time I’m seeing is the time of the last timer tick, which is before the start time that I captured before writing the file. Having tested this explicitly, it looks like in practice the mtime I see is indeed always older than the real current time by somewhere between 0–4ms plus a constant (~700µs), which is exactly what we’d expect.

Is Linux actually being POSIX-compliant here? Not that it matters too much, but I’d say not? While POSIX allows file time updates to be delayed, I don’t see that it allows the time used during an update to be anything other than the current time at the point the update is run7, which is the same time printed by date, etc.

In other words, POSIX seems to require that the recorded time is on-or-after the actual time, and Linux records a time that’s on-or-before the actual time.

Multigrain timestamps!

While I was researching this, I ran across a new feature called multigrain timestamps. This isn’t present in the version of Debian that I’m currently using (it was added to Linux 6.13, so it’ll be in Debian 14), but it does change kernel behaviour in an interesting way.

The kernel documentation is pretty easy to read, but in summary this feature watches to see if a file’s timestamp has been observed (with stat(), e.g.) since its times were last updated, and if so, the next write will use a fine-grained timestamp (i.e. the ‘real’ time), if using a coarse-grained timestamp would make it look like the modification time hadn’t changed.

(There’s a small extra wrinkle in that, to maintain ordering, the act of using a fine-grained timestamp for any file also has to drag along the current ‘coarse’ timestamp, otherwise a file that only needs a coarse-grained timestamp could appear to have been updated before an earlier update that needed a fine-grained timestamp.)

It seems like this was primarily intended for NFSv3 exports, which want to use timestamps to see if a file has changed since they were last read, but I think it should also help in any other cases where you have something watching a file for changes.

(It won’t help fix my tests, though, since the decision about whether to use a fine-grained or coarse-grained timestamp is made when the file is written, and the first write to a file will — I assume — always use a coarse-grained timestamp.)

Fixing (‘fixing’) my tests

So how should I fix my flaky tests?

I could have fixed them by replacing “Get the current time” with “Create a temporary file on the same filesystem as the target file and read its mtime” (or alternatively I’m pretty sure I could read CLOCK_REALTIME_COARSE directly, albeit that might not be easy from Python).

If I were to do that, the three times would then be monotonically increasing. In practice, this test takes much less than 4ms to run, so it’s actually extremely likely that all three times would be exactly the same. (This should not be that much of a surprise.)

But taking a step back, it’s also worth considering why I had this test in the first place.

In this case, the functionality was inherited from an even older incarnation of this tool that relied upon Subversion’s use-commit-times feature to record when each post was last updated, and a desire to have the (HTML) output file have a last-modified timestamp close to that of the (Markdown) source file.

I don’t store posts in Subversion any more, so this whole feature wasn’t really achieving much. And by far the most-sensible thing to do here is to simply to delete the code (and tests) that was playing around with mtime, and just let the kernel pick a reasonable time by itself.

And so that’s how I fixed my flaky test.

References

In the interests of citing my sources, here’s a few more that I used while researching this post.

The Linux source and mailing lists are a great resource for finding out why things work the way they do. In particular, the following were particularly useful:

Also:


  1. I’m quoting from the older (2004) POSIX standard here for simplicity: later versions replace the time_t fields with struct timespec fields with slightly different names, and redefine the existing members as macros. 

  2. Among the other things that POSIX leaves unspecified: whether file time updates must run for all files at once, or whether it’s technically compliant for two files updated one after another to end up with mtimes in the opposite order. (Nobody does that last one in practice — it’d break tools like make that need to see if an input file has changed relative to an output file — but I don’t see how POSIX disallows it.) 

  3. If you happen to be looking at the Linux source code (as I was when writing this post), note that only fs/ext4 supports nanosecond resolution and years past 2038. In contrast, fs/ext2 does not: even though it can use ext2 filesystems with 256-byte inodes, it won’t do anything with the additional data. However, fs/ext4 is almost certainly what any modern kernel will be configured to use for ext2 (and ext3) filesystems, as fs/ext2 is essentially just kept around for reference purposes now. 

  4. That test dates from 2013, but at the time I would have been running them on a small Debian 7 box under my desk (whereas now I’m running on a small Debian 13 box under someone else’s cloud), but I don’t think that’s old enough for the filesystem to have been created with 128-byte inodes. It’s just about possible that I was referring to an even-older Slackware box that I might still have been using at the time. 

  5. Applications reading the current time repeatedly is also why nowadays userspace calls like clock_gettime() and gettimeofday() usually don’t involve a kernel syscall (and so won’t appear in strace output); see man vdso for a detailed explanation. In that case the returned time is always correct, though. 

  6. The kernel tick rate is determined by CONFIG_HZ, which on Debian is set to 250, giving an update every 1s/250 = 4 milliseconds. 

  7. I do see the (non-normative) POSIX rationale, which says “The accuracy of the time update values is intentionally left unspecified so that systems can control the bandwidth of a possible covert channel”, which is a) not what I would have guessed for the primary justification! but also b) reads more to me as not setting an upper-bound on the amount of time that can pass between a write and an (unforced) file time update, rather than allowing an arbitrary (and otherwise unmentioned) jitter to be introduced into update time values.