Summary

  • Slurm allocates nodes, CPUs, memory, GPUs and other resources, translating partitions, priorities, quality-of-service rules, fair-share and reservations into queue decisions.
  • TRES, GRES, topology and cgroups let operators schedule accelerators as constrained physical resources; a GPU count alone does not guarantee useful placement or efficient execution.
  • Slurm developers founded SchedMD in 2010 to provide commercial engineering, support and training; NVIDIA acquired the company on 15 December 2025 and pledged continued open-source, vendor-neutral development.
  • Post-acquisition credibility will depend on multi-vendor testing, release behaviour and contribution patterns, while each cluster operator remains accountable for the policy its users actually experience.

An idle GPU is not necessarily an available GPU

On a Slurm cluster, a physically idle accelerator does not automatically go to the next person who asks for one. A job must first be eligible for a partition, fit the requested resource shape, satisfy account and quality-of-service rules, avoid reservations that block the required nodes and outrank competing work. Only then does the scheduler allocate resources and allow the job to start.

That sequence makes Slurm more than a queue. It is an admission and allocation control plane for a large class of high-performance-computing and AI systems. Users submit jobs; slurmctld evaluates them against cluster state and policy; slurmd processes on compute nodes launch and monitor work; slurmstepd manages individual job steps. An optional slurmdbd service records jobs, resource use and account relationships in a database.

The consequence is economic as well as technical. An AI cluster may have enough GPUs in aggregate and still be unable to start a large training job because the free devices are fragmented across the wrong nodes, sit behind unsuitable topology or are reserved for another project. A scheduler can reduce that waste by representing the relevant constraints. It cannot create missing GPUs, fix a congested fabric or make an inefficient application useful.

Slurm’s importance therefore comes from the decisions it makes before the application runs. The project has spent more than two decades turning a simple question — who may use which machine now? — into a configurable system for allocating increasingly heterogeneous and expensive infrastructure.

Slurm turns local policy into machine time

Slurm began in a Lawrence Livermore National Laboratory-led collaboration and first appeared in 2002. Its early job was to coordinate independently submitted parallel work across large Linux clusters without requiring every application to understand the full machine. The original architecture separated resource allocation from job execution and exposed a common control layer that could be adapted across institutions.

That basic separation remains visible. The controller maintains a central view of nodes, jobs and scheduling state, while compute-node daemons execute the work that has already been authorised. A backup controller and persisted state can reduce the impact of a controller outage, but high availability still depends on coherent state, working authentication, network reachability and tested recovery procedures. “Fault tolerant” is a design capability, not a guarantee that every control-plane failure is harmless.

The deeper change came as Slurm accumulated policy primitives. Partitions group nodes into service classes or administrative pools. Associations connect users and accounts to shares, limits and historical usage. Quality-of-service rules can change priority, limits or preemption behaviour. Reservations hold resources for maintenance, events or named users. Multifactor priority can combine age, fair-share, job size, partition, QOS and site-defined factors.

There is no universal Slurm definition of fairness. One university may favour projects that have used less than their long-term allocation. A national laboratory may reserve capacity for a campaign. A commercial GPU operator may create differentiated service tiers. The same software can express all three because the site defines the policy.

That flexibility is one of Slurm’s strengths and one of its operational risks. A queue can become difficult to explain when partitions overlap, exceptions accumulate and several weighting systems interact. Users then experience the outcome as arbitrary even when the software is implementing the configuration exactly. The scheduler can calculate priority; the institution must still justify the policy.

Fair-share determines who waits, not what fairness means

Fair-share is often discussed as though it were an objective property of the scheduler. In practice it is a mechanism for carrying an institution’s allocation choices across time. Historical usage, account hierarchy and configured shares can influence future priority so that a group that has consumed less than its entitlement receives an advantage over one that has consumed more.

That makes accounting part of governance. slurmdbd can record jobs, steps, associations and trackable resources across one or more clusters. Administrators use those records for reporting, chargeback, usage limits and fair-share calculations. A database field that looks administrative can therefore affect when a researcher or engineering team next receives scarce compute.

The quality of the ledger matters. If users are mapped to the wrong account, resource use is not recorded consistently or historical data is retained incorrectly, the resulting priority can be technically valid and institutionally wrong. Changes in project membership, shared service accounts and manually corrected records all need governance because the scheduler may treat them as evidence about entitlement.

This is one reason queue disputes are hard to reduce to a software bug. A long wait can be caused by demand, inaccurate wall-time requests, a reservation, a low fair-share factor, a topology requirement, a QOS rule or simply a job that cannot fit into the current free resources. Slurm exposes the mechanisms, but the operator needs enough observability to reconstruct which one mattered.

For users, explainability is therefore part of service quality. A queue is easier to accept when people can see why a job is pending, what policy applies and what would allow it to start. As clusters become more expensive and more commercially important, that transparency becomes a management issue rather than a convenience for researchers.

Backfill turns empty gaps into useful work

A strict priority queue can waste capacity. A large high-priority job may be first in line but unable to start until enough nodes become free. Without additional logic, smaller jobs that could finish before that reservation might also wait, leaving resources idle.

Slurm’s backfill scheduler addresses that problem by estimating when higher-priority jobs can begin and then starting lower-priority work that should finish without delaying them. The scheduler does not simply ask which job comes next. It asks whether a job can use a temporary opening while preserving an expected start for work ahead of it.

The mechanism is powerful because large clusters are frequently fragmented. Some nodes finish early; others remain occupied. A short job may fit into the gap without changing the start time of the job the institution considers more important. Backfill can therefore improve utilisation and reduce waiting time at the same time.

Its effectiveness depends on the information it receives. If users request far more wall time than they need, the scheduler may conclude that a job cannot safely fit. If they request too little, the job may be terminated before it finishes. Topology and accelerator constraints can make a theoretically available gap unusable. Failures can invalidate the expected schedule.

Backfill illustrates Slurm’s operating model in miniature. The software can make a sophisticated decision from declared state, but it cannot know the future perfectly. Better queue outcomes depend on accurate requests, reliable cluster state and policies that give the scheduler enough room to make trade-offs.

GPUs made the shape of an allocation as important as its size

Accelerators changed what “available capacity” means. A request for eight CPUs is often more interchangeable than a request for eight GPUs in a distributed training job. The accelerator model, memory capacity, PCIe or NVLink relationships, network position and node composition can determine whether the allocation performs as expected.

Slurm represents heterogeneous resources through Trackable RESources, or TRES, and Generic RESources, or GRES. GPUs can be counted, typed and associated with nodes. Device and cgroup integration can restrict a job to the accelerators it was allocated. Topology plugins and constraints can help the scheduler place work with some awareness of the physical machine.

That turns the GPU from an attached peripheral into a schedulable economic unit. Administrators can account for accelerator usage, limit access, reserve particular device types and design policies around scarce hardware. For AI infrastructure, this is important because the queue is often deciding access to the most expensive component in the cluster.

A resource model, however, is only as useful as the topology it captures. Eight free GPUs spread across nodes with unsuitable communication paths may not be equivalent to eight GPUs inside two tightly connected servers. A scheduler can choose based on configured topology, but it cannot infer every network, memory or application dependency automatically.

The same limit applies to utilisation. A dashboard may show that GPUs are allocated while the job waits on storage, collective communication, data loading or repeated failures. Slurm can tell an operator who held the resource and when. It does not, by itself, prove that the accelerator was doing productive work.

The scheduler sits above the fabric but still depends on it

Slurm is not in the data path. Once a job starts, application traffic moves through processors, memory, interconnects and storage without passing through the scheduler. That does not make the scheduler independent of the physical infrastructure.

Placement choices can concentrate or spread work across switches, blocks or accelerator domains. A topology-aware allocation may reduce the communication distance for a parallel job. A topology-blind allocation can turn enough raw capacity into a poor-performing job because the useful bandwidth is somewhere else.

The scheduler also depends on accurate node state. A GPU can be present but unhealthy. A node can be reachable to the controller while its storage path is degraded. A network partition can make a running job look different from the controller’s view. Plugins, node health checks and local operations need to translate those physical conditions into states the scheduler can act on.

This creates a boundary that is easy to misread in performance reporting. If a job runs slowly, the root cause may be allocation, application behaviour, storage, network contention, accelerator health or a combination. If the cluster is idle, the cause may be low demand, fragmentation, reservations or failures rather than a bad scheduling algorithm.

For operators, the useful measure is therefore the whole path from request to completed work: queue delay, allocation quality, launch success, execution time, retries, lost work and final completion. Aggregate GPU allocation is informative, but it is not the same as productive output.

SchedMD turned an open project into a support business

As Slurm moved beyond its laboratory origins, organisations needed more than source code. Production clusters required predictable releases, debugging, upgrade help, training and engineers who could work across unusual site configurations. Slurm developers formed SchedMD in 2010 to provide that commercial layer.

The arrangement created a familiar open-source bargain. The code remained openly available under its project licence, while customers paid for expertise, support and development around difficult production systems. Commercial work gave maintainers a way to fund sustained engineering and gave operators an escalation path when the queue controlling a major cluster behaved unexpectedly.

SchedMD also became a concentration point for knowledge. Large scheduling systems accumulate operational detail that is hard to learn from documentation alone: failure recovery, upgrade ordering, plugin interactions, accounting edge cases and the effects of unusual policies. A company that employs key maintainers can turn that experience into a support advantage without owning every contribution or every deployment.

That distinction matters because Slurm policy has always remained local. SchedMD could ship code, patches and guidance; it did not decide the fair-share weights, reservations or account entitlements inside a university, national laboratory or commercial AI service. The user experience of Slurm is partly upstream software and partly the institution’s own constitution.

By the time AI expanded the value of scheduled GPU capacity, SchedMD’s role was therefore larger than that of a conventional software vendor. It was the principal commercial steward of an open control plane that many operators had already embedded into their workflows, scripts, accounting systems and operating procedures.

NVIDIA changed the incentives around stewardship

NVIDIA announced its acquisition of SchedMD on 15 December 2025. It said Slurm would remain open-source and vendor-neutral, while arguing that SchedMD’s developers would gain access to more accelerated systems and engineering resources. The licence did not suddenly become proprietary, and the acquisition did not transfer local scheduling policy from operators to NVIDIA.

What changed was the incentive structure around the project’s principal commercial steward. NVIDIA is not only a software company funding maintainers. It is also a leading supplier of GPUs, networking and systems whose performance can depend on how workloads are discovered, placed and launched.

That creates a plausible benefit and a plausible concern. Earlier access to complex AI systems can improve testing and shorten the path from hardware changes to scheduler support. The same proximity raises a legitimate question about whether competing accelerators, interconnects and system designs continue to receive first-class attention.

The evidence available in the first months after the acquisition does not justify declaring either capture or perfect neutrality. Public development continued. Slurm 26.05 and subsequent patches showed active release work, while July 2026 patch releases addressed crashes and other operational issues. Slinky also continued to develop. Those are stronger indicators of stewardship than an acquisition-day promise, but they do not resolve the long-term comparative question.

NVIDIA also described Slurm as widely used in leading supercomputing systems and said SchedMD supported hundreds of customers at the time of acquisition. Those statements indicate scale but remain company-reported and dated. They are not a complete census of private AI clusters, research systems or every scheduler deployment.

The neutrality question should therefore be framed as an observable engineering test. Are interfaces kept generic where they can be? Are issues affecting competing hardware handled openly and promptly? Do release processes and continuous-integration environments exercise a genuinely heterogeneous hardware base? Can outside contributors still influence the code without moving through a proprietary NVIDIA product?

Open source gives operators an exit right, not a free replacement

Slurm’s open-source licence matters because operators can inspect, modify and redistribute the code under its terms. That creates a formal barrier against a simple conversion into closed software and gives the community a legal route to fork if stewardship becomes unacceptable.

A viable fork, however, is not created by a licence alone. Large-scale scheduling needs maintainers who understand controller state, accounting, plugins, releases, security and a broad hardware matrix. It needs test systems, user confidence and people willing to backport fixes across supported versions.

The practical switching cost is also much larger than replacing one executable. Mature Slurm environments accumulate job scripts, account structures, historical usage, custom plugins, monitoring, operational procedures and user habits. Another scheduler may be technically capable and still require a costly migration of policy and institutional memory.

This is why NVIDIA ownership deserves scrutiny without treating forkability as a complete answer. The strongest form of neutrality is not the theoretical ability to leave after a problem. It is a project that remains useful across heterogeneous infrastructure before leaving becomes necessary.

The same reasoning applies to commercial support. Operators may rely on the expertise of the company that employs key maintainers even when the code is open. If that expertise narrows around one hardware ecosystem, the source can remain available while the practical support boundary becomes less neutral.

Slinky puts two control planes in the same estate

Modern AI infrastructure increasingly combines batch scheduling with Kubernetes. Platform teams may want Kubernetes for provisioning, operators, services and container lifecycle while retaining Slurm’s job model, fair-share, reservations and parallel workload semantics.

Slinky is SchedMD’s attempt to bridge those worlds. Its slurm-operator can deploy and manage Slurm components through Kubernetes-oriented mechanisms, while slurm-bridge coordinates work between Kubernetes and Slurm against shared resources. Version 1.2.0 was released on 2 July 2026 after the first stable line appeared in late 2025.

The attraction is clear. An organisation can preserve established Slurm policy while using cloud-native tooling to manage infrastructure around it. That can reduce the need to build an entirely separate operational environment for batch compute.

The difficulty is authority. Kubernetes and Slurm have different models of desired state, workload ownership and recovery. If both systems believe they control a node, a device or a workload after a failure, the integration needs a clear answer about which state is authoritative and how the other system is reconciled.

This is not a reason to reject the approach. It is the reason Slinky should be judged on operational evidence rather than architectural neatness. Production deployments need to show how upgrades, fencing, RBAC, controller failures and partial network partitions are handled when two orchestration systems are involved.

Slurm’s history has repeatedly expanded the boundary of what the scheduler coordinates. Kubernetes integration continues that pattern, but every new control surface increases the importance of knowing where responsibility moves when something breaks.

The queue can improve utilisation and still produce a bad outcome

Slurm gives operators many ways to make expensive capacity more useful. Backfill can reduce idle gaps. Fair-share can distribute access over time. Topology-aware placement can improve locality. Reservations can protect critical work. Preemption can make room for urgent or premium jobs.

Each mechanism also has a cost. A reservation can strand capacity if the expected job does not arrive. Preemption can destroy useful work when applications cannot checkpoint. A topology rule can preserve performance for one job while increasing fragmentation for others. Fair-share can reward a policy that no longer matches the institution’s priorities.

The risk grows when operators optimise one metric. High GPU allocation can be achieved by keeping devices assigned to work that is stalled elsewhere. Low queue time can be achieved by admitting jobs onto resource shapes that lengthen execution. Aggressive preemption can protect one service level while wasting the power and compute already spent on interrupted jobs.

For AI infrastructure, the better measure is completed useful work per unit of scarce capacity and time. Slurm contributes to that outcome, but it is only one layer. Training frameworks, storage, network design, checkpointing, accelerator health and user request quality all affect whether the allocation creates value.

This is also the boundary between upstream and local responsibility. If a site chooses a policy that privileges one account, overuses reservations or sets unrealistic preemption rules, the result should not automatically be attributed to SchedMD or NVIDIA. If the scheduler miscalculates state, mishandles a device or introduces a regression, upstream behaviour becomes the relevant layer.

A credible operation needs enough auditability to distinguish those cases.

The real control surface is split across several actors

Slurm’s stewardship can look central because one controller schedules the cluster and one company now employs many of the project’s experts. In practice, control is divided.

Upstream maintainers decide which code enters releases. NVIDIA owns SchedMD and can allocate engineering resources. Hardware vendors contribute integration work and provide systems for testing. Cluster administrators choose versions, plugins, topology models, accounts, QOS and limits. Institutional leaders decide who is entitled to scarce compute. Users decide what resources to request and how accurately they describe runtime. The application then determines whether the allocation is used efficiently.

That layered control is the central fact of Slurm. No single actor owns the whole result.

It also explains why queue governance has become strategically important. When accelerators were less scarce and less valuable, a suboptimal rule could be irritating. In a large AI estate, the same rule can change waiting times, fragmentation and the amount of expensive capacity that finishes useful work.

Slurm’s achievement is that a common open system can express very different allocation models without forcing every institution into one definition of fairness. Its limitation is identical: the software cannot guarantee that a chosen model is wise, legible or legitimate.

The long-term test is therefore not whether Slurm continues to schedule jobs. It is whether operators can still reconstruct why a job received or lost access to scarce compute, while the upstream project remains credible across the heterogeneous hardware those operators want to run.