Release Notes (GPU PaaS) - July 2026 - Rafay Product Documentation

Jul

v3.1-40

03 Jul, 2026

Token Factory

Image Generation Support

Token Factory now supports image generation workloads, allowing administrators to configure metering based on either a fixed cost per image or a cost-per-megapixel pricing model.

The Token Usage dashboards have also been enhanced to show cost distribution across input tokens, output tokens, and generated images, giving administrators a clearer view of multimodal usage and spend.

Anthropic API Compatibility

Support has been added for the /v1/messages endpoint, improving compatibility with Anthropic-style APIs and tools such as Claude CLI.

Pre-Configured Provider Catalog

A curated catalog of pre-configured model providers has been added to reduce setup time and simplifying onboarding for commonly used model ecosystems, like Llama, Qwen, DeepSeek, Gemma, Mistral, Falcon, OpenAI, Moonshot, NVIDIA Nemotron, Google.

API for user-level Token Usage

A new API to retrieve Token Factory usage data at the user and API key level is now available, enabling more granular reporting, usage tracking, and chargeback workflows.

Administrators can use the Get API Key Cost Usage API (POST /v1alpha1/apikey/cost) to retrieve cost usage by user. Pass the user’s login email in the users field, and optionally use the orgIds filter to scope results to specific organizations.

Granular Project Sharing for Model Deployments

Model deployments can now be shared with specific projects within an organization, providing finer-grained access control and governance.

Partner Admins can share model deployments with one or more organizations. Org Admins can then further control access by sharing those models with one or more projects within their organization, determining where each model is available for deployment and inference.

Endpoint Creation UX Enhancements

The endpoint creation workflow now provides sensible default values for key parameters such as replicas and memory, reducing manual configuration effort.

Compute Cluster Cleanup Assistance

Administrators can now access a cleanup script from the UI when deregistering a compute cluster, simplifying removal of cluster-related resources.

DRA-based GPU Scheduling

Support for enabling NVIDIA Compute Domains during compute cluster registration for supported multi-node NVIDIA NVL72 environments has now been added. This allows advanced GPU scheduling and allocation using Kubernetes DRA for compatible NVIDIA infrastructure.


Platform Enhancements

Localization Support

Localization support across all end-user portals and the Ops Console, with out-of-the-box support for English, Turkish, Japanese, French, and Spanish. (Arabic is currently supported in the Developer Hub only).

Please refer to localization documentation for more information.

Organization-Level GPU Quotas

Support has been added for configuring GPU based quotas at the organization level in the Operations Console. Providers can now define quota limits for a tenant organization based on GPU model, GPU count and data center.

This enables operators to control how much GPU capacity each organization is allowed to consume within a specific data center. For example, an operator can limit an organization to a fixed number of H100 GPUs in a selected data center while allowing different quota limits for other GPU models or locations.

Please refer to documentation on quotas for more information.

Enhanced Org Creation Workflow

Partner Administrators can now choose the default user role (Tenant Admin or Org Admin) when creating a new organization in the Ops Console.

Global Settings Management via Ops Console

Configurations such as quotas, cost estimates, override variables, and agents previously had to be managed through a YAML file in the default organization, referred to as Global Settings.

This experience has now been simplified. Administrators can configure these settings directly in the Operations Console. When upgrading to this release, existing Global Settings configurations are automatically migrated to the new model, and the previous YAML-based configuration approach will be blocked going forward.

For more information, see Overrides, Rate cards, and Orchestration Agents

Support for International Characters in User Profile Fields

Support has been added for international characters in user-input text fields, such as first name and last name.

Users can now enter names that include Unicode/UTF-8 characters, including accented characters, hyphens, and apostrophes. Examples include é, è, ç, Jean-Pierre, and O'Connor.

Note
Email address validation continues to follow standards-compliant ASCII validation rules, consistent with major identity providers.

Email Template Customization and Localization

Support has been added for customizing and localizing system-generated email templates from the Operations Console.

Partners can now configure email templates under User Experience > Email Templates to customize email subjects, HTML body content, and localized versions of supported system emails. This enables partners to tailor platform-generated emails to match their branding, terminology, and customer communication requirements.

Note
The default language configured on the Localization page is always included in outgoing emails. Additional localized content is appended after the default language content in the configured order.

For more information, please refer to documentation on Email Templates.

Email Notification Delivery Audit

Visibility into email notification activity and SMTP submission status has been added to the Operations Console. Operators can now review whether platform-generated email notifications were successfully submitted to the configured SMTP provider.

Note
Delivery status reflects submission to the configured SMTP provider. Final recipient delivery, spam filtering, mailbox-level rejection, or downstream provider handling may depend on the SMTP provider and recipient mail system.

Inventory and Data Center Management

Support has been added for inventory data center import and export, bulk inventory import/export workflows, and pagination across inventory objects, simplifying inventory management at scale.

Server Type Allocation API

A new API is now available to retrieve server type allocation per data center and across all data centers. This API can be used to monitor inventory capacity by server type, such as SERVER or GPU-SERVER, view allocation status counts, and review aggregate CPU, memory, and storage capacity.

API path: GET /v2/sentry/paas/infrainventory/servertype-allocation

Audit Trails for Server Allocation

Inventory now shows the allocation history for servers under Audit Trails. Each entry records the previous allocation status, the new allocation status, who made the change, and when the change occurred.

Rows and Racks Support

Inventory now supports rows and racks for NVIDIA rack server trays and base systems (for example, GB200), giving users a structured way to organize and manage physical server placement in the data center.

Inventory can be maintained in two ways:


Bare Metal Server SKU

Serial Console Support

Serial console support is now available for BCM and Non-BCM bare metal servers, enabling out-of-band access for troubleshooting and recovery.

Note
Serial console requires a gateway deployed on a machine in the data center with network connectivity to the server BMC.

Cisco Hyperfabric Integration

BMaaS can now be deployed with Cisco Hyperfabric for BCM and Non-BCM bare metal management. Data centers using Cisco Hyperfabric as their network fabric can provision bare metal instances using VPC and subnet resources managed by the Cisco network fabric.

For more information, see Bare Metal Integrations.

Replace Instance

Users can now recover a BCM bare metal instance using the Replace Instance action when hardware issues occur on the same server for example, after replacing faulty network interfaces, a failed BMC, or switch port details — without rebuilding the instance from scratch.

Update the interface, BMC, and switch port details on the existing server, then run Replace Instance to refresh the platform instance while keeping the same configuration intact.

For operator setup steps and the end-user workflow, see Replace Instance.

Note
Replace Instance is currently supported only for BCM bare metal servers. Support for Non-BCM bare metal is planned for a future release.


VM SKU

Serial Console Support

Serial console support is now available for VMs enabling out-of-band access for troubleshooting and recovery.

Note
Serial console access requires a gateway to be deployed on a machine in the data center with network connectivity to the server hosting the VMs.

VM Auto-Migration / Hypervisor Failure Handling

Support has been added for automated VM migration when a hypervisor is determined to be down.

The system uses a dual health-check model before initiating auto-migration:

  1. Hypervisor metrics that is pushed to the controller
  2. The Gateway Agent performing ongoing proactive hypervisor checks, including SSH-based checks

If both signals indicate that the hypervisor is down, and auto-migration is enabled, the platform can automatically initiate VM migration without manual intervention.

Prerequisites:


Kubernetes SKU

Kubernetes v1.36 Support

New and existing Kubernetes SKU clusters can use Kubernetes v1.36 (provision and in-place upgrade).

Worker Node Pool Labels and Taints

Support has been added for configuring labels and taints on worker node pools, enabling finer-grained workload scheduling and node isolation.

For more information, see Worker Node Pool Labels and Taints.

Service Account Issuer and API Audiences

Support has been added for configuring Service Account Issuer and API Audiences for service account OIDC tokens, enabling secure token validation and improved compatibility with OIDC-based workload integrations.

For more information, see Service Account Issuer and API Audiences.


Controller Enhancements and Improvements

Harbor Registry Adoption

Beginning this release, the Rafay Controller image registry is migrating from Nexus to Harbor for improved scalability, reliability, and lifecycle management of container images.

Note
If you need an older or previously unsupported image version after upgrade, it may need to be republished to Harbor before use.

Storage Backend Migration from OpenEBS to Longhorn

For new bare-metal Rafay Controller deployments, the default storage backend is changing from OpenEBS to Longhorn.

This change improves long-term platform sustainability and addresses known security considerations with OpenEBS upstream support.

Data Disk Size Requirement

With this release, the mandatory data disk size for new bare-metal controller deployments is updated from 1 TB to 2 TB.

Requirement Previous Updated (this release)
Data disk (/data) 1 TB 2 TB (mandatory)

The larger /data volume is required so controller application services have sufficient storage capacity and operate as expected. Plan new controller deployments with a 2 TB data disk mounted at /data.