The Ops Notebook

Monitoring Kubernetes Volume Usage with kubelet_volume_stats_used_bytes (and Its Gotchas)

Published September 28, 2026 · 3 min read

The kubelet already exposes per-volume usage metrics to Prometheus. Here's the PromQL we use for dashboards and alerts, plus the five gotchas that will mislead you if you trust the numbers blindly.

Monitoring Kubernetes Volume Usage with kubelet_volume_stats_used_bytes (and Its Gotchas)

Every kubelet exposes per-volume filesystem metrics, and if you run Prometheus with kube-state-metrics or the standard Kubernetes scrape configs, you’re probably already collecting them and ignoring them. These two metrics are the foundation of every storage dashboard we’ve built:

Here’s how to turn them into something you’d actually page on.

The essential queries

Utilization per PVC:

kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes

Top 10 volumes by usage right now:

topk(10, kubelet_volume_stats_used_bytes)

Cluster-wide storage efficiency (provisioned vs. used):

sum(kubelet_volume_stats_used_bytes) / sum(kubelet_volume_stats_capacity_bytes)

That last one is the number for your cost review meetings. Ours sat under 45% for months — which is what kicked off our EBS rightsizing project.

Predicting “disk full” with predict_linear:

predict_linear(kubelet_volume_stats_used_bytes[7d], 14 * 24 * 3600)
  > kubelet_volume_stats_capacity_bytes

This fires when the 7-day trend says the volume fills within 14 days. It’s noisy on bursty workloads, so pair it with a minimum-utilization threshold (e.g., only alert above 75% used) to cut false positives during low-usage growth spurts.

Alerting rules we actually run

groups:
  - name: volume-alerts
    rules:
      - alert: VolumeFillingUp
        expr: |
          predict_linear(kubelet_volume_stats_used_bytes[7d], 14*24*3600)
            > kubelet_volume_stats_capacity_bytes
            and
          kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes > 0.75
        for: 6h
        labels:
          severity: warning
        annotations:
          summary: "Volume {{ $labels.persistentvolumeclaim }} filling up within 14 days"

      - alert: VolumeCriticallyFull
        expr: kubelet_volume_stats_used_bytes / kubelet_volume_stats_capacity_bytes > 0.9
        for: 15m
        labels:
          severity: critical
        annotations:
          summary: "Volume {{ $labels.persistentvolumeclaim }} over 90% full"

The for: 6h on the predictive alert matters — predict_linear on a 7-day window swings wildly with short bursts. Six hours of sustained prediction is a real trend; six minutes is a deploy.

The five gotchas

1. Unmounted volumes are invisible. The kubelet only reports stats for volumes mounted into a running pod. Orphaned PVCs — the ones silently billing you — never appear. Audit those with a separate query joining kube_persistentvolumeclaim_info against actual pod mounts.

2. ext4 reserves 5% by default. The metric reports the filesystem’s view, so that reserved space counts as “used.” A volume reporting 96% used on ext4 is effectively full for non-root processes. Set your critical threshold at 90%, not 95%.

3. It lags. Kubelet volume stats update on the order of minutes, and Prometheus scrapes on top of that. Don’t build autoscaling on this metric; it’s a planning and alerting signal, not a control loop.

4. available_bytes ≠ capacity - used. Reserved blocks, inodes, and filesystem overhead mean the three metrics don’t reconcile exactly. Use used / capacity for utilization and available only for “can I write one more segment” sanity checks.

5. Ephemeral storage has its own metrics. These volume_stats metrics cover persistent volumes. For container ephemeral storage (emptyDir, logs, image layers), you want kubelet_volume_stats_* on the ephemeral volume or node-level node_filesystem_* metrics instead. Mixing the two is a classic dashboard bug.

Building the summary view

Raw per-PVC time series get noisy at fleet scale. We aggregate into a cached summary API: per-cluster provisioned vs. used, top volumes by utilization, and week-over-week growth rate. The key design decision: cache aggressively and degrade gracefully. If Prometheus is down, serve the last good snapshot with a staleness banner instead of a blank dashboard — a storage dashboard that errors out during the incident where you need it is worse than no dashboard.

Two rollups cover 90% of questions:

# Provisioned vs used per StorageClass (find your expensive habits)
sum by (storageclass) (kubelet_volume_stats_capacity_bytes)
sum by (storageclass) (kubelet_volume_stats_used_bytes)

# Growth rate per week per PVC (find your future problems)
(sum by (persistentvolumeclaim) (kubelet_volume_stats_used_bytes)
 - sum by (persistentvolumeclaim) (kubelet_volume_stats_used_bytes offset 7d))
 / 7 / 24 / 3600

Start with utilization and the predictive alert. Add the cost rollup when finance starts asking questions — they will, and you’ll have the answer before they finish the sentence.

About the author

The Ops Notebook is written by an operations engineer running Kubernetes data infrastructure (Redpanda, Elasticsearch) in production. Every article is based on real incidents, real cost numbers, and real fixes — not rewritten documentation.