2026-09-03

GPU Overprovisioning: Time-Slicing, MIG, and Backfill

Compare queue backfill, NVIDIA Time-Slicing, MIG, and runtime sharing with Kubernetes configuration, validation, isolation, and rollback steps.

Low GPU utilization does not mean that one card can safely be advertised as four. First identify where the waste occurs: poor queue ordering, memory held while compute sits idle, or replicas sized for a peak that rarely arrives. Each cause needs a different fix.

Kubernetes normally exposes GPUs as extended resources such as nvidia.com/gpu. Extended resources are scheduled as integers and cannot be natively overcommitted like CPU. A node with eight GPUs can satisfy at most eight simultaneous nvidia.com/gpu: 1 requests by default. Sharing requires a device plugin, runtime, or hardware partition to expose different resources; the scheduler does not create memory isolation or compute limits on its own.

Separate the four mechanisms

Mechanism Same-card concurrency Isolation Best fit Main risk
Queue backfill, priority, preemption No Keeps full-GPU isolation Training, batch, interruptible jobs Preemption cost and checkpoint recovery
NVIDIA Time-Slicing Yes No memory or fault isolation Development, low-risk inference, small jobs Correlated OOMs and latency jitter
MIG Yes Hardware compute and memory partitioning Multi-tenant inference and stable small workloads Fixed profiles, fragmentation, reconfiguration cost
MPS, vGPU, or framework-level sharing Yes Implementation-dependent Fixed workloads needing finer controls Operational complexity, compatibility, uneven observability

If a large job only blocks smaller jobs behind it, fix queueing before sharing a card. Kueue Cohorts let queues borrow unused quota and combine that with priority and preemption. They redistribute cluster quota; they do not partition one GPU’s memory.

Build a baseline before sharing

Collect at least one week of data covering releases, traffic peaks, and checkpoint phases. A single nvidia-smi sample is not a capacity plan.

nvidia-smi \
  --query-gpu=index,uuid,utilization.gpu,memory.used,memory.total,power.draw \
  --format=csv

kubectl get pods -A --field-selector=status.phase=Pending
kubectl get events -A --sort-by='.lastTimestamp' | tail -n 100
kubectl describe node <gpu-node> | sed -n '/Capacity:/,/System Info:/p'

With DCGM Exporter, compare sustained utilization with peaks instead of relying on averages:

avg_over_time(DCGM_FI_DEV_GPU_UTIL[1h])
max_over_time(DCGM_FI_DEV_FB_USED[24h])
max_over_time(DCGM_FI_DEV_MEM_COPY_UTIL[1h])

A useful sharing candidate normally satisfies all of these conditions:

  • Compute has repeatable idle windows rather than sampling noise.
  • Peak memory leaves headroom for model loading and burst batches.
  • A failed attempt can be retried without data loss.
  • The service can tolerate some tail-latency variation.
  • Pod, GPU, OOM, request-latency, and retry records can be correlated.

Core online inference, long training runs, and jobs without checkpoints should not be the first experiment.

Run a minimal Time-Slicing experiment

The following configuration uses a distinct nvidia.com/gpu.shared resource so shared Pods do not accidentally land in the exclusive pool. replicas: 4 exposes four schedulable shares per physical GPU. It does not reserve one quarter of memory or guarantee one quarter of compute for each Pod.

Save as time-slicing-config.yaml:

apiVersion: v1
kind: ConfigMap
metadata:
  name: time-slicing-config
  namespace: gpu-operator
data:
  shared: |-
    version: v1
    flags:
      migStrategy: none
    sharing:
      timeSlicing:
        renameByDefault: true
        failRequestsGreaterThanOne: true
        resources:
          - name: nvidia.com/gpu
            replicas: 4    

Apply the configuration and point the GPU Operator device plugin at it:

kubectl apply -f time-slicing-config.yaml

kubectl patch clusterpolicies.nvidia.com/cluster-policy \
  -n gpu-operator --type merge \
  -p '{"spec":{"devicePlugin":{"config":{"name":"time-slicing-config","default":"shared"}}}}'

kubectl get pods -n gpu-operator -w

The GPU Operator does not watch ConfigMap changes and automatically restart the device plugin. Restart it during a maintenance window after modifying the configuration:

kubectl rollout restart -n gpu-operator daemonset/nvidia-device-plugin-daemonset

Confirm that the node reports both its physical count and shared capacity:

kubectl get node <gpu-node> \
  -o jsonpath='{.metadata.labels.nvidia\.com/gpu\.count}{" physical\n"}{.metadata.labels.nvidia\.com/gpu\.replicas}{" replicas\n"}{.status.allocatable.nvidia\.com/gpu\.shared}{" shared allocatable\n"}'

Verify scheduling with four Pods

Limit the experiment to one node instead of changing the whole GPU pool:

kubectl label node <gpu-node> gpu-mode=shared
apiVersion: apps/v1
kind: Deployment
metadata:
  name: gpu-share-check
spec:
  replicas: 4
  selector:
    matchLabels:
      app: gpu-share-check
  template:
    metadata:
      labels:
        app: gpu-share-check
    spec:
      nodeSelector:
        gpu-mode: shared
      containers:
        - name: check
          image: nvcr.io/nvidia/cuda:12.8.1-base-ubuntu22.04
          command: ["bash", "-lc"]
          args: ["nvidia-smi -L && sleep infinity"]
          resources:
            limits:
              nvidia.com/gpu.shared: 1
kubectl apply -f gpu-share-check.yaml
kubectl get pods -l app=gpu-share-check -o wide
kubectl exec deploy/gpu-share-check -- nvidia-smi -L
kubectl describe node <gpu-node> | sed -n '/Allocated resources:/,$p'

Four Pods landing on a single-GPU node proves only that schedulable capacity increased. It does not prove memory isolation or higher application throughput. Run the real model next and compare per-physical-GPU throughput, P95/P99 latency, OOMs, Pod restarts, and GPU Xid errors against the exclusive baseline.

Time-Slicing boundaries

  • Processes share the same physical GPU and memory fault domain. One process exhausting memory can affect neighbors.
  • The scheduler sees logical shares, not each process’s live memory peak.
  • nvidia.com/gpu.shared: 2 does not buy twice the compute. failRequestsGreaterThanOne rejects that misleading request.
  • DCGM metrics are primarily attributed to the physical GPU. Pod-level analysis still needs process, container, and application data.
  • Higher density may add context switching. Tail latency can regress before throughput improves.

Keep shared and exclusive pools separate. Node labels, taints and tolerations, plus a distinct resource name are enough for the first rollout. Do not make latency-sensitive services rely on luck to avoid shared nodes.

When MIG is the better fit

MIG partitions supported NVIDIA GPUs into fixed hardware instances with dedicated memory, cache, and compute resources. It fits multi-tenant inference requiring predictable latency better than Time-Slicing, but it introduces profile fragmentation: leftover capacity may not form the profile a workload needs, and reconfiguration can disrupt existing work.

Answer four questions before choosing MIG:

  1. Does the model fit in the target profile with enough KV Cache or training peak headroom?
  2. Do the selected profiles cover common models instead of one temporary workload?
  3. Can the platform tolerate node reconfiguration, eviction, and rescheduling windows?
  4. Do monitoring, quota, and capacity reports understand MIG resource names?

If those answers are missing, stay with exclusive GPUs or a small Time-Slicing trial. Hardware partitioning is not an automatic optimizer.

Rollout and rollback gates

Start at replicas: 2 with retryable, low-priority workloads. The following gates are examples; replace their thresholds with the application’s SLOs:

Signal Expand Roll back
Useful throughput per physical GPU Clearly beats exclusive baseline Flat or lower
P99 latency Remains inside the SLO Repeatedly exceeds the SLO
OOMs and Pod restarts No worse than baseline Correlated same-card failures appear
GPU Xid errors Zero Any new error
Queue time Continues to fall Waiting merely becomes runtime jitter

For rollback, restore the previously saved ClusterPolicy configuration and restart the device plugin. Do not remove the shared resource name while workloads still use it. Stop new placement, drain the shared nodes, confirm migration, then restore exclusive mode.

Final decision

GPU overprovisioning is not one switch. The usual order is: correct requests and replica counts, add queue backfill, trial Time-Slicing for low-risk workloads with stable idle windows, then evaluate MIG or vGPU when hard isolation is required.

Success is not a node reporting more GPUs. Success means each physical card completes more useful work while latency, OOMs, and failure scope remain inside known limits.

References