Hard

Design a Shared GPU Cluster Scheduler System Design Interview

Design scheduling and capacity management for training, research, and production inference on a large heterogeneous accelerator fleet.

1. Problem Statement

Design a scheduler for a shared fleet used by production inference, multi-day training jobs, and interactive research experiments.

2. Architecture Discussion Map

Use this as one discussion aid, not a single correct answer. Your design should follow from the requirements, scale, and trade-offs you establish.

Rendering architecture diagram...
Mermaid Source (For AI Bots)
graph LR
    A["Design a Shared GPU Cluster Scheduler"]
    A --> F1["Resource model, topology, and gang scheduling"]
    A --> F2["Priority, quotas, preemption, and fairness"]
    A --> F3["Fragmentation, bin packing, and capacity forecasting"]
    A --> F4["Control-plane scalability and failure domains"]
    A --> F5["Checkpointing, retries, observability, and cost attribution"]

3. Key Focus Areas

  • 1
    Resource model, topology, and gang scheduling
  • 2
    Priority, quotas, preemption, and fairness
  • 3
    Fragmentation, bin packing, and capacity forecasting
  • 4
    Control-plane scalability and failure domains
  • 5
    Checkpointing, retries, observability, and cost attribution

4. What Strong Candidates Should Demonstrate

  • Model jobs with different priorities, topology, duration, and failure semantics.
  • Balance utilization against fragmentation, fairness, and production SLOs.
  • Limit control-plane blast radius at thousands-of-node scale.

Want interactive feedback?

Practice drawing this system component-by-component on a live whiteboard while the interviewer probes at your target level.

Continue to Dashboard

Core Concepts

GPU SchedulingKubernetesGang SchedulingPreemptionCapacity Planning

Continue preparing

Build a complete software engineer mock interview plan