GPU as a Service (GPUaaS) Platform | Build & Deliver GPUaaS | Rafay

Enterprise GPU as a Service (GPUaaS) Platform

Rafay is the platform that enables enterprises and providers to deliver GPU as a Service.

Organizations are investing heavily in GPU infrastructure, but most struggle to deliver it as a usable service. Access is manual, environments are inconsistent, and utilization remains low. Developers wait for resources while expensive GPUs sit idle.

GPU as a Service (GPUaaS) solves this by enabling on-demand, self-service access to GPU resources. Rafay provides the platform layer that allows enterprises and providers to build and operate GPUaaS offerings—turning raw infrastructure into a scalable, governed service for AI/ML workloads.

.webp)

What is GPUaaS?

GPU as a Service (GPUaaS) delivers on-demand access to GPU compute through APIs or self-service portals, similar to how cloud platforms deliver CPU-based infrastructure. Instead of provisioning clusters manually, users can instantly launch GPU-backed environments with built-in governance, isolation, and usage tracking.

Rafay enables this model by transforming existing GPU infrastructure into a fully operational GPUaaS platform.

How to Build and Operate a GPUaaS Platform

Deliver a Self-Service GPUaaS Experience

Enable developers, data scientists, and customers to provision GPU resources instantly without tickets or manual intervention. Rafay provides a fully automated, self-service experience for AI/ML workloads.

Operate GPUaaS with Multi-Tenant Control

Deliver GPU as a Service securely across teams, business units, or external customers. Rafay provides built-in multi-tenancy, RBAC, and policy enforcement so infrastructure can be shared safely.

Maximize GPU Utilization and Efficiency

GPUaaS platforms only succeed when utilization is high and waste is minimized. Rafay pools GPU resources and dynamically allocates them across workloads to ensure infrastructure is fully utilized.

Unified Orchestration for GPUaaS Infrastructure

Rafay provides centralized orchestration across Kubernetes, GPUs, and hybrid environments—enabling consistent delivery of GPUaaS across cloud and on-prem infrastructure.

Trusted by leading enterprises, neoclouds and service providers

.png)

One Platform - Multiple Deployment Options

Deploy as a SaaS

A majority of Rafay customers consume Rafay in a SaaS form factor. Why? Because the SaaS model lets them start immediately with the Rafay Platform and deliver value to their customers. The Rafay platform is SOC-2 Type compliant, and will address all requirements put forward by your security team.

Deploy in an Air-Gapped Model

Customers in highly regulated industries prefer Rafay’s air-gapped controller model. Team Rafay is ready to help you deploy the Rafay Platform in your data center or in your private/public cloud environment. You get exactly the same experience and all the same features available to our SaaS customers.

Deploy across Data Center and CSP Environments

Whether you plan to deploy GPUs in multiple colos, or lease GPUs in a CSP environment, or both, Rafay can help. With Rafay, all your compute across all private and CSP environments can be managed as a single pool of GPUs and CPUs, reducing operational overhead and enabling cloud-bursting use cases.

GPU PaaS™ FAQs for Enterprises

What counts as a node?

A node is a physical or virtual server/machine.

Do you have any volume discounts?

Yes! As the number of nodes increases the price per cluster or per node decreases.

What about short-lived or ephemeral clusters?

Our customers love to experiment, and we don’t ding them for it. We don’t charge for node count spikes, but look at the running average of nodes in use when calculating usage.

Is there a difference between production and non-production pricing?

The management overhead for helping our customers operate dev vs prod clusters is effectively the same, so we treat all nodes the same.

What if I use more nodes than I’ve licensed?

Rafay has a true-up forward policy, meaning that we don’t carry out chargebacks for scenarios where the consumption in a completed billing cycle exceeded the licensed count. If the new, steady-state number of nodes is expected to be higher, our customer success team will discuss the situation with you, and take steps to adjust billing accordingly for the next billing cycle.

How much does Enterprise Support (24x7x365) cost?

Enterprise Support is available at an additional fee equaling 20% of the cluster or node subscription.

Do you have EDU or GOV discounts?

Yes, please contact sales for more information about discounts for educational institutions and government agencies.

Is GPU virtualization supported?

Yes. The Rafay Platform supports three GPU sharing modes that operators can offer to tenants in self-service: full passthrough (one physical GPU per workload, optimal for large training runs), NVIDIA MIG (Multi-Instance GPU) partitioning (up to seven isolated MIG instances per A100 or H100, each with dedicated memory and compute), and time-slicing (multiple workloads sharing a GPU in time-multiplexed fashion, suited for lower-intensity inference or development workloads). Operators configure which sharing modes are available per SKU through PaaS Studio; tenants select the appropriate GPU size from the catalog without needing to understand the underlying partitioning mechanism. Security and compute isolation between MIG instances is enforced at the NVIDIA hardware level; chargeback data is collected per MIG instance or per time-slice allocation for granular cost attribution across tenants and business units.

Do you provide AI/ML workbenches and other tooling?

Yes. The Rafay Platform offers a variety of workbenches out of the box. These are based on Kubeflow and KubeRay, with end users consuming these platforms “as a service,” without needing to configure or operate any of these tools on their own. Further, the Rafay platform provides a low-code/no-code framework that empowers partners to bring new capabilities to market faster, e.g. verticalized agents, co-pilots, document translation services, and more.

Does your platform also support CPU consumption?

Yes. The Rafay Platform has always supported CPU-based workloads and can easily deliver a PaaS experience that offers CPU+GPU instances to end users.

How does Rafay solve for chargeback and billing?

Rafay offers a comprehensive solution for chargebacks and billing. The platform collects granular chargeback information on resource usage, which can be easily exported to customers’ existing billing systems for further processing and distribution. Rafay allows for customizable chargeback group definitions to align with organizational structures or projects. Both group definition and data collection can be carried out programmatically, enabling efficient and accurate billing processes.

Does Rafay support “infrastructure as code” (IaC) principles?

Yes. Rafay supports a number of IaC frameworks, enabling customers to programmatize every aspect of their cloud. The Platform supports Terraform, OpenTofu, GitOps pipelines, CLI and API workflows out of the box.

Turn your GPU infrastructure into a GPUaaS platform

Deliver self-service GPU access, improve utilization, and scale AI/ML workloads with confidence.