gpu-manager Architecture: Device Plugin, vCUDA, and Migration
Review tkestack/gpu-manager device registration, vCUDA resources, library injection, topology allocation, CUDA limits, and migration choices.
Status note (September 3, 2026): the latest public
tkestack/gpu-managerrelease remainsv1.1.0, published in May 2020. Its companionvcuda-controllerdocumentation declares support for CUDA 11.5.1 and earlier. Treat this project as an architecture case study, not a default greenfield choice for a 2026 cluster.
gpu-manager remains useful because it decomposes Kubernetes GPU sharing into problems that still matter: how a scheduler expresses fractional resources, how kubelet delivers a virtual device to a container, how runtime code enforces memory and compute limits, and how allocation state survives node restarts.
The three problems it solves
- Resource expression: split a physical GPU into
vcuda-coreandvcuda-memoryextended resources. - Container delivery: use a Device Plugin
Allocateresponse to inject devices, directories, environment variables, and control libraries. - Runtime enforcement: intercept CUDA calls through a vCUDA library for quota checks, device-view rewriting, and usage collection.
All three layers must hold. Scheduler accounting without runtime enforcement still allows overuse. Runtime interception without scheduler accounting leaves the platform unable to place workloads safely.
Component boundaries
Pod / Extended Resources
|
v
gpu-admission / scheduler integration
|
v
gpu-manager DaemonSet
- GPU discovery
- topology and allocation state
- Device Plugin registration
- Allocate response
|
v
vCUDA controller library in container
|
v
NVIDIA driver and physical GPU
- gpu-manager runs on GPU nodes, discovers devices, tracks allocations, and registers Device Plugins with kubelet.
- gpu-admission or scheduler integration handles fractional placement and topology constraints.
- vcuda-controller provides the in-container interception library that makes the scheduler decision enforceable.
- checkpoint and metrics services restore node allocation state and expose device and workload usage.
This is a tightly coupled system. gpu-manager, the vCUDA library, NVIDIA Driver, CUDA Runtime, and container runtime require a tested compatibility matrix.
Startup flow: discovery to Device Plugin registration
The startup path fits into seven steps:
- Read node configuration, device policy, resource granularity, and allocation strategy.
- Discover GPUs, memory, and topology through NVML or driver interfaces.
- Initialize local allocation state and restore a checkpoint when available.
- Prepare driver libraries and the vCUDA control library for container mounts.
- Start Device Plugin services for
vcuda-core,vcuda-memory, and related resources. - Register extended resources with kubelet and report health through
ListAndWatch. - Answer
Allocatewith device nodes, Mounts, environment variables, and runtime configuration.
The Kubernetes Device Plugin API defines registration and allocation. It does not understand how “20% of a GPU” should be enforced. gpu-manager therefore maintains a separate mapping among physical GPUs, virtual shares, and containers.
The vCUDA resource model
The project defines two primary extended resources:
tencent.com/vcuda-core: compute share;100represents one full GPU and20represents roughly a 20% share.tencent.com/vcuda-memory: memory share; one unit represents 256 MiB, so32represents 8 GiB.
For vcuda-core, values below 100 represent a fraction of one card; larger values normally use multiples of 100. This is not a native Kubernetes percentage. gpu-manager and the vCUDA runtime jointly define the semantics.
This manifest shows the legacy resource contract, not a promise that a current cluster can run it unchanged:
apiVersion: v1
kind: Pod
metadata:
name: vcuda-demo
annotations:
tencent.com/vcuda-core-limit: "50"
spec:
restartPolicy: Never
containers:
- name: workload
image: <your-registry>/cuda-workload:<pinned-version>
resources:
requests:
tencent.com/vcuda-core: "20"
tencent.com/vcuda-memory: "32"
limits:
tencent.com/vcuda-core: "20"
tencent.com/vcuda-memory: "32"
The Pod requests about 20% compute share and 8 GiB memory; vcuda-core-limit sets a 50% runtime ceiling. Verify exact field behavior against the pinned gpu-manager revision.
Why Allocate is the critical interface
A Device Plugin Allocate response can deliver:
- device nodes
- HostPath Mounts
- environment variables
- CDI Devices
- other runtime-specific information
gpu-manager uses this stage to place virtual-device state and the control library inside the container. At startup, dynamic linking loads the vCUDA library before forwarding or limiting calls to the underlying CUDA stack.
This is more capable than reporting several virtual device IDs, but it exposes the main compatibility risk. A library-path change, symbol-version mismatch, uncovered Driver API, or container-runtime change can produce a Pod that schedules successfully and fails during CUDA initialization.
What CUDA interception can and cannot control
vCUDA-style systems commonly handle:
- device enumeration and visibility
- memory calls such as
cudaMallocandcuMemAlloc - Context, Stream, and Kernel Launch paths
- usage collection and quota decisions
Their boundaries are equally important:
- New or uncovered CUDA APIs can bypass enforcement.
- Static linking, explicit
dlopen, and custom Driver API paths increase compatibility risk. - User-space interception is not MIG-level hardware isolation; PCIe, memory bandwidth, caches, and failure domains may remain shared.
- CUDA, Driver, and control-library ABI changes need release-by-release regression testing.
“The Pod sees a GPU” is not acceptance. Prove allocation, over-limit failure, concurrent interference, restart recovery, and metrics behavior.
Why topology awareness matters
Multi-GPU placement cannot stop at device count. Distances among GPUs, CPU NUMA nodes, PCIe Switches, NVLink, and NVSwitch affect Host-to-Device, Peer-to-Peer, and collective performance.
Allocation must answer:
- whether a multi-GPU workload stays inside one high-speed interconnect domain
- whether CPU and memory are close to the target GPU NUMA node
- whether fractional packing blocks future whole-GPU workloads
- whether restored virtual shares still point to the same physical devices after a restart
These principles remain current even when the old project’s topology APIs and device identifiers do not.
What to evaluate for a 2026 cluster
| Option | Primary capability | Isolation boundary | Best fit |
|---|---|---|---|
| NVIDIA GPU Operator Time-Slicing | Expose one card as several shared replicas | No hard memory boundary; shared failure domain | Trusted workloads and low-cost concurrency |
| NVIDIA MPS | Concurrent processes with architecture-dependent resource controls | Depends on GPU generation and MPS configuration | Same-host inference or HPC concurrency |
| NVIDIA MIG | Hardware partitions with independent resource instances | Strongest, with fixed profiles and supported-GPU limits | Multi-tenancy and failure isolation |
| HAMi | Device Plugin plus CUDA API-level enforcement | Stronger than application self-limiting, weaker than hardware partitioning | Heterogeneous shared GPU platforms |
| KAI-Scheduler GPU Sharing | Reservation Pod and scheduler accounting | No default hard memory enforcement | Fractional scheduling and queue governance |
Related articles:
- KAI-Scheduler vs HAMi: Scheduling Accounting and Runtime Enforcement
- GPU Overprovisioning: Overselling, Sharing, Isolation, and Rollback
Minimum gate for maintaining a legacy gpu-manager cluster
- Pin Linux Kernel, NVIDIA Driver, CUDA, gpu-manager, and vcuda-controller versions together.
- Keep canary nodes and test Driver or container-runtime upgrades before fleet rollout.
- Validate registration and checkpoint recovery after kubelet, gpu-manager, and node restarts.
- Check
vcuda-coreandvcuda-memoryCapacity, Allocatable, and Pod allocation results. - Confirm the actual control-library path and symbol versions inside the container.
- Run memory-ceiling, compute-quota, interference, CUDA error-propagation, and OOM tests.
- Collect physical GPU, virtual share, and Pod metrics on one correlated timeline.
kubectl get node <gpu-node> \
-o jsonpath='{.status.capacity.tencent\.com/vcuda-core}{"\t"}{.status.capacity.tencent\.com/vcuda-memory}{"\n"}'
kubectl describe pod -n <ns> <pod>
kubectl logs -n <gpu-manager-namespace> <gpu-manager-pod>
Do not copy broad RBAC or host-access examples from an old README into production unchanged. Re-scope ServiceAccounts, HostPaths, HostPID, and device access for the current Kubernetes release.
Migration path
- Inventory every workload’s
vcuda-core,vcuda-memory, memory peak, and compute utilization. - Translate the legacy units into the target resource model; replacing resource names is not enough.
- Run old and new node pools in parallel and move retryable, lower-priority workloads first.
- Compare throughput, P95/P99 latency, memory peak, failure rate, and node density for the same model.
- Keep rollback for CUDA initialization failures, performance regressions, and unenforced quotas.
- Move core online services last, then remove the old Admission, Device Plugin, and control-library injection path.
gpu-manager is still worth studying because it shows the complete engineering combination of scheduler accounting, Device Plugin delivery, and user-space enforcement. It is no longer a sensible default for new deployment; the reusable lessons are component boundaries, compatibility matrices, and verification methods.