Skip to content

[TRACKING] Guardian auto-restart on path loss can kill shared NVMe subsystems (namespaced volumes) #298

Description

@noctarius

Situation

The CSI node driver's Guardian (kubernetes/csi/pkg/util/guardian.go) tracks which pods own which volumes (lvols). After all NVMe-oF paths to a volume are lost and the storage cluster becomes active again, the Guardian restarts the owning pod(s) to force a CSI restage (remount) — required so filesystem volumes pick up the recovered paths. Auto-restart is opt-in per workload via simplyblock.io/auto-restart-on-pathloss=true (opt-out: simplyblock.io/guardian-disable=true).

Namespaced volumes place multiple volumes (NVMe namespaces) under a single NVMe subsystem, bounded by the max volumes per subsystem setting (default 5 on our e2e cluster). Because the subsystem is shared, restarting a pod to recover one volume disturbs the shared subsystem and tears down the NVMe-oF paths of the other volumes on it — killing paths for parallel, otherwise-healthy pods that merely share the subsystem.

This surfaced in the new e2e tests that exercise automatic pod restart after total path loss: a single test's guardian-triggered restart knocked out the subsystem shared by parallel-running tests. The configuration is valid — this is a real interaction, not a misconfiguration.

Impact

  • A Guardian restart intended for one workload causes path loss for unrelated workloads sharing the subsystem, cascading restarts / I/O errors.
  • Blast radius scales with max volumes per subsystem.

Immediate workaround (intermediate solution)

Block Guardian auto-restart for namespaced volumes — any volume whose subsystem hosts more than one volume — even when the workload opted in via the label. The Guardian must detect the shared-subsystem case and skip (with a log line and a pod Event) rather than restart. This trades auto-restaging convenience for correctness on shared subsystems until a proper solution exists.

Long-term solution

Coordinate the restart of all pods owning all volumes on an affected subsystem, so they restage together and no pod is left on a half-torn-down subsystem. Feasibility is open: the Guardian runs per-node (DaemonSet) while affected pods may span nodes, so coordination likely needs a cluster-wide component (operator-level orchestration) and careful sequencing. Requires investigation before implementation.

Sub-issues

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingcsi

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions