Rafay-Powered SLURM-as-a-Service (SLURMaaS) | Rafay Plaform

Rafay-Powered SLURM-as-a-Service

Rafay-powered SLURM as a Service delivers fully managed, multi-tenant SLURM environments for high-performance computing workloads as a cloud-like, on-demand service. Through automated, BCM-based cluster bring-up with secure per-tenant separation and governance built in, tenants request a cluster and Rafay provisions, schedules, and governs it.

What Is Slurm-as-a-Service?

Slurm as a Service delivers the Slurm workload manager through a fully managed, on-demand platform. Instead of manually deploying and maintaining HPC clusters, organizations can provision per-tenant Slurm clusters, submit jobs to managed queues, and allocate compute resources through self-service, while the platform automates both the underlying Kubernetes cluster's provisioning and the Slurm environment's scheduling, governance, and lifecycle management.

Rafay automates bring-up of the underlying Kubernetes cluster and layers Slurm on top via the open-source Slinky Slurm Operator, enabling providers and enterprises to deliver HPC resources as scalable, self-service Slurm clusters.

HPC Meets Cloud-Native Agility

Familiar HPC scheduling, delivered as a governed, on-demand service.

Simplified Deployment

Automated BCM-based bring-up removes manual management of the underlying Kubernetes cluster.

Ready-to-Use Access

Login nodes let users submit jobs immediately.

Multi-Tenant Isolation

Multiple tenants or teams with full isolation and operational control.

Full Automation

Provisioning, scaling, and teardown handled automatically.

Break Free from HPC Silos

Expand into HPC and research markets on existing GPU infrastructure.

Automated BCM-based bring-up removes manual HPC cluster management.

Per-tenant metering and billing turn HPC scheduling into a priced service.

Serve traditional HPC and modern AI/ML workflows on shared infrastructure.

Benefits of Rafay-Powered Slurm as a Service

Deploy HPC Clusters Faster

Provision fully managed Slurm clusters on demand through a self-service portal or API.

Reduce Operational Overhead

Automate cluster provisioning, scaling, lifecycle management, and governance to eliminate manual administration.

Maximize Infrastructure Utilization

Run HPC and AI workloads on shared CPU and GPU infrastructure to improve resource efficiency.

Support Multiple Teams Securely

Deliver isolated, multi-tenant Slurm environments with centralized governance and policy controls.

Monetize HPC Services

Turn HPC infrastructure into a managed, consumption-based service with built-in usage metering and chargeback.

Scale with Confidence

Deliver production-ready HPC environments that support research, engineering, and AI workloads from a single platform.

Common Use Cases of Rafay’s Slurm-as-a-Service

Whether you're delivering HPC services to customers or supporting internal research and engineering teams, Rafay helps simplify HPC operations while maximizing infrastructure utilization.

For Cloud Providers

For Enterprises

Frequently Asked Questions on SLURM

What counts as a node?

A node is a physical or virtual server/machine.

Do you have any volume discounts?

Yes! As the number of nodes increases the price per cluster or per node decreases.

What about short-lived or ephemeral clusters?

Our customers love to experiment, and we don’t ding them for it. We don’t charge for node count spikes, but look at the running average of nodes in use when calculating usage.

Is there a difference between production and non-production pricing?

The management overhead for helping our customers operate dev vs prod clusters is effectively the same, so we treat all nodes the same.

What if I use more nodes than I’ve licensed?

Rafay has a true-up forward policy, meaning that we don’t carry out chargebacks for scenarios where the consumption in a completed billing cycle exceeded the licensed count. If the new, steady-state number of nodes is expected to be higher, our customer success team will discuss the situation with you, and take steps to adjust billing accordingly for the next billing cycle.

How much does Enterprise Support (24x7x365) cost?

Enterprise Support is available at an additional fee equaling 20% of the cluster or node subscription.

Do you have EDU or GOV discounts?

Yes, please contact sales for more information about discounts for educational institutions and government agencies.

What is Slurm?

Slurm (Simple Linux Utility for Resource Management) is an open-source workload manager used to schedule and manage compute-intensive workloads across high-performance computing (HPC) clusters. It allows users to submit jobs, request compute resources, join queues, and run workloads across CPUs, GPUs, and other infrastructure.

Who uses Slurm as a Service?

Slurm as a Service enables cloud providers, neoclouds, research organizations, and enterprises to offer managed HPC and GPU compute environments without building and operating the entire platform themselves.

How is Rafay different from managing Slurm yourself?

Managing Slurm independently requires significant expertise to deploy, configure, secure, scale, and maintain HPC clusters. Rafay automates bring-up of the underlying Kubernetes cluster and simplifies delivery of per-tenant Slurm clusters on top of it, enabling self-service access and enterprise-grade governance.

Does Rafay support bare metal, virtual machines, Kubernetes, and SLURM?

Yes. Rafay supports provisioning and lifecycle management across bare metal, virtual machines, Kubernetes, and SLURM environments, allowing providers to deliver standardized infrastructure and AI services from a single platform.

Can Slurm run alongside Kubernetes?

Yes. Through the open-source Slinky Slurm Operator, Slurm's scheduler runs on top of the same Kubernetes cluster that Rafay provisions, so Slurm-based HPC jobs and native Kubernetes workloads share the same underlying infrastructure rather than running as separate stacks.

Does Rafay offer a GPU PaaS?

Yes, Rafay provides infrastructure orchestration and workflow automation for cloud-native (Kubernetes) and AI use cases for enterprises, cloud providers, neoclouds, and Sovereign AI clouds.