Avoid recursive chown of the home PVC on every workspace start #2

Merged
GuillaumeHemmen merged 1 commit from fix/fsgroup-change-policy into master 2026-09-29 14:53:52 +00:00
Member

Problem

Workspaces take ~4 minutes to start, and the Coder web UI reports that the agent is slow to come up.

The pod security_context sets fs_group = 1000 with no fs_group_change_policy, which defaults to Always. kubelet therefore recursively chowns the entire /home/coder volume on every single workspace start, before the container is created.

Evidence

Measured on workspace ghe-perso (pod coder-12f3cdc1-...-p9f95, node talos-ovh-worker-one):

Time (UTC) Event Duration
13:58:43 Scheduled + Longhorn volume attached ~0s
13:58:5x → 14:02:41 VolumePermissionChangeInProgress ~3m50s
14:02:41 Image pulled 75ms
14:02:41 Container started, agent up ~1s

Pod events:

Warning  VolumePermissionChangeInProgress  kubelet
  Setting volume ownership for .../pvc-46f76366-.../mount is taking longer
  than expected, consider using OnRootMismatch
  ... processed 375713 files.
  ... processed 642227 files.
  ... processed 1093420 files.

The home volume currently holds 1,144,713 inodes (Projects 871k, .cache 218k, .nvm 35k) on a 100Gi Longhorn RWO volume — several minutes of metadata operations on network-backed storage.

The image is not the bottleneck. The 2.6 GB sindri image pulled from the node's containerd cache in 75ms (the digest check against git.van-hemmen.com is fast from inside the cluster), and the agent logged agent is starting now 1 second after the container started.

Fix

Set fs_group_change_policy = "OnRootMismatch". kubelet then only chowns when the volume root is not already owned by fs_group — i.e. once, on first use. Expected effect: workspace start drops from ~4 min to ~15s.

This also matters going forward: the cost grows linearly with file count, so it degrades as workspaces accumulate node_modules, build caches and git objects.

Notes / follow-ups (not in this PR)

  • image_pull_policy = "Always" with no node affinity means a workspace scheduled onto a node without the tag cached would eat a 2.6 GB pull. All three workers happen to have sindri images cached today, so this is latent rather than active.
  • Clearing ~/.cache would shave ~20% of the inodes, but that is a band-aid; the policy change is the actual fix.

Testing

Not applied to the cluster — investigation was read-only. terraform validate was not run (terraform is not installed in the workspace); the change is a single attribute addition to an existing security_context block. Verification is to push the template and time the next workspace start.

🤖 Generated with Claude Code

https://claude.ai/code/session_015QKnVbbjpAi4TaSmrkNXAE

## Problem Workspaces take ~4 minutes to start, and the Coder web UI reports that the agent is slow to come up. The pod `security_context` sets `fs_group = 1000` with no `fs_group_change_policy`, which defaults to `Always`. kubelet therefore recursively chowns the **entire** `/home/coder` volume on every single workspace start, before the container is created. ## Evidence Measured on workspace `ghe-perso` (pod `coder-12f3cdc1-...-p9f95`, node `talos-ovh-worker-one`): | Time (UTC) | Event | Duration | |---|---|---| | 13:58:43 | Scheduled + Longhorn volume attached | ~0s | | 13:58:5x → 14:02:41 | `VolumePermissionChangeInProgress` | **~3m50s** | | 14:02:41 | Image pulled | 75ms | | 14:02:41 | Container started, agent up | ~1s | Pod events: ``` Warning VolumePermissionChangeInProgress kubelet Setting volume ownership for .../pvc-46f76366-.../mount is taking longer than expected, consider using OnRootMismatch ... processed 375713 files. ... processed 642227 files. ... processed 1093420 files. ``` The home volume currently holds **1,144,713 inodes** (`Projects` 871k, `.cache` 218k, `.nvm` 35k) on a 100Gi Longhorn RWO volume — several minutes of metadata operations on network-backed storage. **The image is not the bottleneck.** The 2.6 GB sindri image pulled from the node's containerd cache in 75ms (the digest check against `git.van-hemmen.com` is fast from inside the cluster), and the agent logged `agent is starting now` 1 second after the container started. ## Fix Set `fs_group_change_policy = "OnRootMismatch"`. kubelet then only chowns when the volume root is not already owned by `fs_group` — i.e. once, on first use. Expected effect: workspace start drops from ~4 min to ~15s. This also matters going forward: the cost grows linearly with file count, so it degrades as workspaces accumulate `node_modules`, build caches and git objects. ## Notes / follow-ups (not in this PR) - `image_pull_policy = "Always"` with no node affinity means a workspace scheduled onto a node without the tag cached would eat a 2.6 GB pull. All three workers happen to have sindri images cached today, so this is latent rather than active. - Clearing `~/.cache` would shave ~20% of the inodes, but that is a band-aid; the policy change is the actual fix. ## Testing Not applied to the cluster — investigation was read-only. `terraform validate` was not run (terraform is not installed in the workspace); the change is a single attribute addition to an existing `security_context` block. Verification is to push the template and time the next workspace start. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_015QKnVbbjpAi4TaSmrkNXAE
The pod security_context sets fs_group = 1000 without an
fs_group_change_policy, so it defaults to "Always" and kubelet
recursively chowns the entire /home/coder volume every time a
workspace starts.

Measured on workspace ghe-perso (pod coder-12f3cdc1-...-p9f95,
node talos-ovh-worker-one):

  13:58:43  Scheduled, Longhorn volume attached      ~0s
  13:58:5x  VolumePermissionChangeInProgress         ~3m50s
  14:02:41  Image pulled (cached)                     75ms
  14:02:41  Container started, agent up               ~1s

kubelet logged "processed 1093420 files" against a home volume
holding 1,144,713 inodes. The 2.6 GB sindri image was not the
bottleneck: it pulled from cache in 75ms and the agent reported
ready one second after the container started.

"OnRootMismatch" only chowns when the volume root is not already
owned by fs_group, i.e. once on first use. This cost also grows
linearly with file count, so it gets worse as workspaces age.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015QKnVbbjpAi4TaSmrkNXAE
GuillaumeHemmen deleted branch fix/fsgroup-change-policy 2026-09-29 14:53:52 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
GuillaumeHemmen-k8s/coder-sindri-deployment!2
No description provided.