GPU Overprovisioning: Time-Slicing, MIG, and Backfill
Compare queue backfill, NVIDIA Time-Slicing, MIG, and runtime sharing with Kubernetes configuration, validation, isolation, and rollback steps.
Low GPU utilization does not mean that one card can safely be advertised as four. First identify where the waste occurs: poor queue ordering, memory held while compute sits idle, or replicas sized for a peak that rarely arrives. Each cause needs a different fix.
Kubernetes normally exposes GPUs as extended resources such as nvidia.com/gpu. Extended resources are scheduled as integers and cannot be natively overcommitted like CPU. A node with eight GPUs can satisfy at most eight simultaneous nvidia.com/gpu: 1 requests by default. Sharing requires a device plugin, runtime, or hardware partition to expose different resources; the scheduler does not create memory isolation or compute limits on its own.
Separate the four mechanisms
| Mechanism | Same-card concurrency | Isolation | Best fit | Main risk |
|---|---|---|---|---|
| Queue backfill, priority, preemption | No | Keeps full-GPU isolation | Training, batch, interruptible jobs | Preemption cost and checkpoint recovery |
| NVIDIA Time-Slicing | Yes | No memory or fault isolation | Development, low-risk inference, small jobs | Correlated OOMs and latency jitter |
| MIG | Yes | Hardware compute and memory partitioning | Multi-tenant inference and stable small workloads | Fixed profiles, fragmentation, reconfiguration cost |
| MPS, vGPU, or framework-level sharing | Yes | Implementation-dependent | Fixed workloads needing finer controls | Operational complexity, compatibility, uneven observability |
If a large job only blocks smaller jobs behind it, fix queueing before sharing a card. Kueue Cohorts let queues borrow unused quota and combine that with priority and preemption. They redistribute cluster quota; they do not partition one GPU’s memory.
Build a baseline before sharing
Collect at least one week of data covering releases, traffic peaks, and checkpoint phases. A single nvidia-smi sample is not a capacity plan.
nvidia-smi \
--query-gpu=index,uuid,utilization.gpu,memory.used,memory.total,power.draw \
--format=csv
kubectl get pods -A --field-selector=status.phase=Pending
kubectl get events -A --sort-by='.lastTimestamp' | tail -n 100
kubectl describe node <gpu-node> | sed -n '/Capacity:/,/System Info:/p'
With DCGM Exporter, compare sustained utilization with peaks instead of relying on averages:
avg_over_time(DCGM_FI_DEV_GPU_UTIL[1h])
max_over_time(DCGM_FI_DEV_FB_USED[24h])
max_over_time(DCGM_FI_DEV_MEM_COPY_UTIL[1h])
A useful sharing candidate normally satisfies all of these conditions:
- Compute has repeatable idle windows rather than sampling noise.
- Peak memory leaves headroom for model loading and burst batches.
- A failed attempt can be retried without data loss.
- The service can tolerate some tail-latency variation.
- Pod, GPU, OOM, request-latency, and retry records can be correlated.
Core online inference, long training runs, and jobs without checkpoints should not be the first experiment.
Run a minimal Time-Slicing experiment
The following configuration uses a distinct nvidia.com/gpu.shared resource so shared Pods do not accidentally land in the exclusive pool. replicas: 4 exposes four schedulable shares per physical GPU. It does not reserve one quarter of memory or guarantee one quarter of compute for each Pod.
Save as time-slicing-config.yaml:
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
shared: |-
version: v1
flags:
migStrategy: none
sharing:
timeSlicing:
renameByDefault: true
failRequestsGreaterThanOne: true
resources:
- name: nvidia.com/gpu
replicas: 4
Apply the configuration and point the GPU Operator device plugin at it:
kubectl apply -f time-slicing-config.yaml
kubectl patch clusterpolicies.nvidia.com/cluster-policy \
-n gpu-operator --type merge \
-p '{"spec":{"devicePlugin":{"config":{"name":"time-slicing-config","default":"shared"}}}}'
kubectl get pods -n gpu-operator -w
The GPU Operator does not watch ConfigMap changes and automatically restart the device plugin. Restart it during a maintenance window after modifying the configuration:
kubectl rollout restart -n gpu-operator daemonset/nvidia-device-plugin-daemonset
Confirm that the node reports both its physical count and shared capacity:
kubectl get node <gpu-node> \
-o jsonpath='{.metadata.labels.nvidia\.com/gpu\.count}{" physical\n"}{.metadata.labels.nvidia\.com/gpu\.replicas}{" replicas\n"}{.status.allocatable.nvidia\.com/gpu\.shared}{" shared allocatable\n"}'
Verify scheduling with four Pods
Limit the experiment to one node instead of changing the whole GPU pool:
kubectl label node <gpu-node> gpu-mode=shared
apiVersion: apps/v1
kind: Deployment
metadata:
name: gpu-share-check
spec:
replicas: 4
selector:
matchLabels:
app: gpu-share-check
template:
metadata:
labels:
app: gpu-share-check
spec:
nodeSelector:
gpu-mode: shared
containers:
- name: check
image: nvcr.io/nvidia/cuda:12.8.1-base-ubuntu22.04
command: ["bash", "-lc"]
args: ["nvidia-smi -L && sleep infinity"]
resources:
limits:
nvidia.com/gpu.shared: 1
kubectl apply -f gpu-share-check.yaml
kubectl get pods -l app=gpu-share-check -o wide
kubectl exec deploy/gpu-share-check -- nvidia-smi -L
kubectl describe node <gpu-node> | sed -n '/Allocated resources:/,$p'
Four Pods landing on a single-GPU node proves only that schedulable capacity increased. It does not prove memory isolation or higher application throughput. Run the real model next and compare per-physical-GPU throughput, P95/P99 latency, OOMs, Pod restarts, and GPU Xid errors against the exclusive baseline.
Time-Slicing boundaries
- Processes share the same physical GPU and memory fault domain. One process exhausting memory can affect neighbors.
- The scheduler sees logical shares, not each process’s live memory peak.
nvidia.com/gpu.shared: 2does not buy twice the compute.failRequestsGreaterThanOnerejects that misleading request.- DCGM metrics are primarily attributed to the physical GPU. Pod-level analysis still needs process, container, and application data.
- Higher density may add context switching. Tail latency can regress before throughput improves.
Keep shared and exclusive pools separate. Node labels, taints and tolerations, plus a distinct resource name are enough for the first rollout. Do not make latency-sensitive services rely on luck to avoid shared nodes.
When MIG is the better fit
MIG partitions supported NVIDIA GPUs into fixed hardware instances with dedicated memory, cache, and compute resources. It fits multi-tenant inference requiring predictable latency better than Time-Slicing, but it introduces profile fragmentation: leftover capacity may not form the profile a workload needs, and reconfiguration can disrupt existing work.
Answer four questions before choosing MIG:
- Does the model fit in the target profile with enough KV Cache or training peak headroom?
- Do the selected profiles cover common models instead of one temporary workload?
- Can the platform tolerate node reconfiguration, eviction, and rescheduling windows?
- Do monitoring, quota, and capacity reports understand MIG resource names?
If those answers are missing, stay with exclusive GPUs or a small Time-Slicing trial. Hardware partitioning is not an automatic optimizer.
Rollout and rollback gates
Start at replicas: 2 with retryable, low-priority workloads. The following gates are examples; replace their thresholds with the application’s SLOs:
| Signal | Expand | Roll back |
|---|---|---|
| Useful throughput per physical GPU | Clearly beats exclusive baseline | Flat or lower |
| P99 latency | Remains inside the SLO | Repeatedly exceeds the SLO |
| OOMs and Pod restarts | No worse than baseline | Correlated same-card failures appear |
| GPU Xid errors | Zero | Any new error |
| Queue time | Continues to fall | Waiting merely becomes runtime jitter |
For rollback, restore the previously saved ClusterPolicy configuration and restart the device plugin. Do not remove the shared resource name while workloads still use it. Stop new placement, drain the shared nodes, confirm migration, then restore exclusive mode.
Final decision
GPU overprovisioning is not one switch. The usual order is: correct requests and replica counts, add queue backfill, trial Time-Slicing for low-risk workloads with stable idle windows, then evaluate MIG or vGPU when hard isolation is required.
Success is not a node reporting more GPUs. Success means each physical card completes more useful work while latency, OOMs, and failure scope remain inside known limits.