Hard

Design a Shared GPU Cluster Scheduler System Design Interview

Design scheduling and capacity management for training, research, and production inference on a large heterogeneous accelerator fleet.

1. Problem Statement

Design a scheduler for a shared fleet used by production inference, multi-day training jobs, and interactive research experiments.

2. Architecture Discussion Map

Use this as one discussion aid, not a single correct answer. Your design should follow from the requirements, scale, and trade-offs you establish.

Rendering architecture diagram...
Mermaid Source (For AI Bots)
graph LR
    A["Design a Shared GPU Cluster Scheduler"]
    A --> F1["Resource model, topology, and gang scheduling"]
    A --> F2["Priority, quotas, preemption, and fairness"]
    A --> F3["Fragmentation, bin packing, and capacity forecasting"]
    A --> F4["Control-plane scalability and failure domains"]
    A --> F5["Checkpointing, retries, observability, and cost attribution"]

3. Key Focus Areas

  • 1
    Resource model, topology, and gang scheduling
  • 2
    Priority, quotas, preemption, and fairness
  • 3
    Fragmentation, bin packing, and capacity forecasting
  • 4
    Control-plane scalability and failure domains
  • 5
    Checkpointing, retries, observability, and cost attribution

4. What Strong Candidates Should Demonstrate

  • Model jobs with different priorities, topology, duration, and failure semantics.
  • Balance utilization against fragmentation, fairness, and production SLOs.
  • Limit control-plane blast radius at thousands-of-node scale.

Evaluation Guide

This is an evaluation framework, not a single model answer. Strong designs may make different choices when their assumptions and trade-offs are explicit.

Workload and resource model

25%

Senior signals

  • Models accelerator type, memory, topology, network locality, gang size, duration, checkpointability, and priority.
  • Separates production reservations from elastic research capacity.

Scheduling policy

30%

Senior signals

  • Explains queueing, gang admission, fair share, quotas, preemption, backfilling, and fragmentation.
  • Defines starvation prevention and safe preemption rules.

Staff-level signals

  • Proposes hierarchical or multi-cluster scheduling that degrades gracefully during partition or control-plane overload.

Reliability and control plane

25%

Senior signals

  • Covers leader failover, reconciliation, idempotent placement, stale node state, and scale-aware rollout.
  • Avoids coupling workload survival and service discovery entirely to one control plane.

Efficiency and measurement

20%

Senior signals

  • Tracks useful accelerator time, queue delay, fragmentation, preemption waste, and cost by team.
  • Uses forecasts and reservations without permanently stranding capacity.

Trade-offs to articulate

  • Utilization versus predictable production headroom.
  • Fairness versus throughput-optimal packing.
  • Fast preemption versus lost training work.

Practice follow-up questions

  1. 1.A telemetry rollout overloads every cluster API server. How do jobs continue?
  2. 2.How would you schedule a 2,000-GPU job without starving smaller experiments?

Want interactive feedback?

Practice drawing this system component-by-component on a live whiteboard while the interviewer probes at your target level.

Start Interview

Core Concepts

GPU SchedulingKubernetesGang SchedulingPreemptionCapacity Planning