Status note (September 3, 2026): the latest public tkestack/gpu-manager release remains v1.1.0, published in May 2020. Its companion vcuda-controller documentation declares support for CUDA 11.5.1 and earlier. Treat this project as an architecture case study, not a default greenfield choice for a 2026 cluster.

gpu-manager remains useful because it decomposes Kubernetes GPU sharing into problems that still matter: how a scheduler expresses fractional resources, how kubelet delivers a virtual device to a container, how runtime code enforces memory and compute limits, and how allocation state survives node restarts.

The three problems it solves

  1. Resource expression: split a physical GPU into vcuda-core and vcuda-memory extended resources.
  2. Container delivery: use a Device Plugin Allocate response to inject devices, directories, environment variables, and control libraries.
  3. Runtime enforcement: intercept CUDA calls through a vCUDA library for quota checks, device-view rewriting, and usage collection.

All three layers must hold. Scheduler accounting without runtime enforcement still allows overuse. Runtime interception without scheduler accounting leaves the platform unable to place workloads safely.

Component boundaries

Pod / Extended Resources
          |
          v
gpu-admission / scheduler integration
          |
          v
gpu-manager DaemonSet
  - GPU discovery
  - topology and allocation state
  - Device Plugin registration
  - Allocate response
          |
          v
vCUDA controller library in container
          |
          v
NVIDIA driver and physical GPU
  • gpu-manager runs on GPU nodes, discovers devices, tracks allocations, and registers Device Plugins with kubelet.
  • gpu-admission or scheduler integration handles fractional placement and topology constraints.
  • vcuda-controller provides the in-container interception library that makes the scheduler decision enforceable.
  • checkpoint and metrics services restore node allocation state and expose device and workload usage.

This is a tightly coupled system. gpu-manager, the vCUDA library, NVIDIA Driver, CUDA Runtime, and container runtime require a tested compatibility matrix.

Startup flow: discovery to Device Plugin registration

The startup path fits into seven steps:

  1. Read node configuration, device policy, resource granularity, and allocation strategy.
  2. Discover GPUs, memory, and topology through NVML or driver interfaces.
  3. Initialize local allocation state and restore a checkpoint when available.
  4. Prepare driver libraries and the vCUDA control library for container mounts.
  5. Start Device Plugin services for vcuda-core, vcuda-memory, and related resources.
  6. Register extended resources with kubelet and report health through ListAndWatch.
  7. Answer Allocate with device nodes, Mounts, environment variables, and runtime configuration.

The Kubernetes Device Plugin API defines registration and allocation. It does not understand how “20% of a GPU” should be enforced. gpu-manager therefore maintains a separate mapping among physical GPUs, virtual shares, and containers.

The vCUDA resource model

The project defines two primary extended resources:

  • tencent.com/vcuda-core: compute share; 100 represents one full GPU and 20 represents roughly a 20% share.
  • tencent.com/vcuda-memory: memory share; one unit represents 256 MiB, so 32 represents 8 GiB.

For vcuda-core, values below 100 represent a fraction of one card; larger values normally use multiples of 100. This is not a native Kubernetes percentage. gpu-manager and the vCUDA runtime jointly define the semantics.

This manifest shows the legacy resource contract, not a promise that a current cluster can run it unchanged:

apiVersion: v1
kind: Pod
metadata:
  name: vcuda-demo
  annotations:
    tencent.com/vcuda-core-limit: "50"
spec:
  restartPolicy: Never
  containers:
    - name: workload
      image: <your-registry>/cuda-workload:<pinned-version>
      resources:
        requests:
          tencent.com/vcuda-core: "20"
          tencent.com/vcuda-memory: "32"
        limits:
          tencent.com/vcuda-core: "20"
          tencent.com/vcuda-memory: "32"

The Pod requests about 20% compute share and 8 GiB memory; vcuda-core-limit sets a 50% runtime ceiling. Verify exact field behavior against the pinned gpu-manager revision.

Why Allocate is the critical interface

A Device Plugin Allocate response can deliver:

  • device nodes
  • HostPath Mounts
  • environment variables
  • CDI Devices
  • other runtime-specific information

gpu-manager uses this stage to place virtual-device state and the control library inside the container. At startup, dynamic linking loads the vCUDA library before forwarding or limiting calls to the underlying CUDA stack.

This is more capable than reporting several virtual device IDs, but it exposes the main compatibility risk. A library-path change, symbol-version mismatch, uncovered Driver API, or container-runtime change can produce a Pod that schedules successfully and fails during CUDA initialization.

What CUDA interception can and cannot control

vCUDA-style systems commonly handle:

  • device enumeration and visibility
  • memory calls such as cudaMalloc and cuMemAlloc
  • Context, Stream, and Kernel Launch paths
  • usage collection and quota decisions

Their boundaries are equally important:

  • New or uncovered CUDA APIs can bypass enforcement.
  • Static linking, explicit dlopen, and custom Driver API paths increase compatibility risk.
  • User-space interception is not MIG-level hardware isolation; PCIe, memory bandwidth, caches, and failure domains may remain shared.
  • CUDA, Driver, and control-library ABI changes need release-by-release regression testing.

“The Pod sees a GPU” is not acceptance. Prove allocation, over-limit failure, concurrent interference, restart recovery, and metrics behavior.

Why topology awareness matters

Multi-GPU placement cannot stop at device count. Distances among GPUs, CPU NUMA nodes, PCIe Switches, NVLink, and NVSwitch affect Host-to-Device, Peer-to-Peer, and collective performance.

Allocation must answer:

  • whether a multi-GPU workload stays inside one high-speed interconnect domain
  • whether CPU and memory are close to the target GPU NUMA node
  • whether fractional packing blocks future whole-GPU workloads
  • whether restored virtual shares still point to the same physical devices after a restart

These principles remain current even when the old project’s topology APIs and device identifiers do not.

What to evaluate for a 2026 cluster

Option Primary capability Isolation boundary Best fit
NVIDIA GPU Operator Time-Slicing Expose one card as several shared replicas No hard memory boundary; shared failure domain Trusted workloads and low-cost concurrency
NVIDIA MPS Concurrent processes with architecture-dependent resource controls Depends on GPU generation and MPS configuration Same-host inference or HPC concurrency
NVIDIA MIG Hardware partitions with independent resource instances Strongest, with fixed profiles and supported-GPU limits Multi-tenancy and failure isolation
HAMi Device Plugin plus CUDA API-level enforcement Stronger than application self-limiting, weaker than hardware partitioning Heterogeneous shared GPU platforms
KAI-Scheduler GPU Sharing Reservation Pod and scheduler accounting No default hard memory enforcement Fractional scheduling and queue governance

Related articles:

Minimum gate for maintaining a legacy gpu-manager cluster

  1. Pin Linux Kernel, NVIDIA Driver, CUDA, gpu-manager, and vcuda-controller versions together.
  2. Keep canary nodes and test Driver or container-runtime upgrades before fleet rollout.
  3. Validate registration and checkpoint recovery after kubelet, gpu-manager, and node restarts.
  4. Check vcuda-core and vcuda-memory Capacity, Allocatable, and Pod allocation results.
  5. Confirm the actual control-library path and symbol versions inside the container.
  6. Run memory-ceiling, compute-quota, interference, CUDA error-propagation, and OOM tests.
  7. Collect physical GPU, virtual share, and Pod metrics on one correlated timeline.
kubectl get node <gpu-node> \
  -o jsonpath='{.status.capacity.tencent\.com/vcuda-core}{"\t"}{.status.capacity.tencent\.com/vcuda-memory}{"\n"}'

kubectl describe pod -n <ns> <pod>
kubectl logs -n <gpu-manager-namespace> <gpu-manager-pod>

Do not copy broad RBAC or host-access examples from an old README into production unchanged. Re-scope ServiceAccounts, HostPaths, HostPID, and device access for the current Kubernetes release.

Migration path

  1. Inventory every workload’s vcuda-core, vcuda-memory, memory peak, and compute utilization.
  2. Translate the legacy units into the target resource model; replacing resource names is not enough.
  3. Run old and new node pools in parallel and move retryable, lower-priority workloads first.
  4. Compare throughput, P95/P99 latency, memory peak, failure rate, and node density for the same model.
  5. Keep rollback for CUDA initialization failures, performance regressions, and unenforced quotas.
  6. Move core online services last, then remove the old Admission, Device Plugin, and control-library injection path.

gpu-manager is still worth studying because it shows the complete engineering combination of scheduler accounting, Device Plugin delivery, and user-space enforcement. It is no longer a sensible default for new deployment; the reusable lessons are component boundaries, compatibility matrices, and verification methods.

References