# Jul

## v3.1-40  
**03 Jul, 2026**

### Token Factory  
#### Image Generation Support  
Token Factory now supports image generation workloads, allowing administrators to configure metering based on either a fixed cost per image or a cost-per-megapixel pricing model.

The Token Usage dashboards have also been enhanced to show cost distribution across input tokens, output tokens, and generated images, giving administrators a clearer view of multimodal usage and spend.

#### Anthropic API Compatibility  
Support has been added for the /v1/messages endpoint, improving compatibility with Anthropic-style APIs and tools such as Claude CLI.

#### Pre-Configured Provider Catalog  
A curated catalog of pre-configured model providers has been added to reduce setup time and simplifying onboarding for commonly used model ecosystems, like Llama, Qwen, DeepSeek, Gemma, Mistral, Falcon, OpenAI, Moonshot, NVIDIA Nemotron, Google.

#### API for user-level Token Usage  
A new API to retrieve Token Factory usage data at the user and API key level is now available, enabling more granular reporting, usage tracking, and chargeback workflows.

Administrators can use the **Get API Key Cost Usage** API (`POST /v1alpha1/apikey/cost`) to retrieve cost usage by user. Pass the user’s login email in the `users` field, and optionally use the `orgIds` filter to scope results to specific organizations.

#### Granular Project Sharing for Model Deployments  
Model deployments can now be shared with specific projects within an organization, providing finer-grained access control and governance.

Partner Admins can share model deployments with one or more organizations. Org Admins can then further control access by sharing those models with one or more projects within their organization, determining where each model is available for deployment and inference.

#### Endpoint Creation UX Enhancements  
The endpoint creation workflow now provides sensible default values for key parameters such as replicas and memory, reducing manual configuration effort.

#### Compute Cluster Cleanup Assistance  
Administrators can now access a cleanup script from the UI when deregistering a compute cluster, simplifying removal of cluster-related resources.

#### DRA-based GPU Scheduling  
Support for enabling NVIDIA Compute Domains during compute cluster registration for supported multi-node NVIDIA NVL72 environments has now been added. This allows advanced GPU scheduling and allocation using Kubernetes DRA for compatible NVIDIA infrastructure.

* * *

### Platform Enhancements  
### Localization Support  
Localization support across all end-user portals and the Ops Console, with out-of-the-box support for English, Turkish, Japanese, French, and Spanish. (Arabic is currently supported in the Developer Hub only).

Please refer to [localization documentation](https://docs.rafay.co/aiml/gpupaas/csp/localization/) for more information.

### Organization-Level GPU Quotas  
Support has been added for configuring GPU based quotas at the organization level in the Operations Console. Providers can now define quota limits for a tenant organization based on GPU model, GPU count and data center.

This enables operators to control how much GPU capacity each organization is allowed to consume within a specific data center. For example, an operator can limit an organization to a fixed number of H100 GPUs in a selected data center while allowing different quota limits for other GPU models or locations.

Please refer to documentation on [quotas](https://docs.rafay.co/aiml/gpupaas/csp/localization/) for more information.

#### Enhanced Org Creation Workflow  
Partner Administrators can now choose the default user role (Tenant Admin or Org Admin) when creating a new organization in the Ops Console.

#### Global Settings Management via Ops Console  
Configurations such as quotas, cost estimates, override variables, and agents previously had to be managed through a YAML file in the default organization, referred to as Global Settings.

This experience has now been simplified. Administrators can configure these settings directly in the Operations Console. When upgrading to this release, existing Global Settings configurations are automatically migrated to the new model, and the previous YAML-based configuration approach will be blocked going forward.

For more information, see [Overrides](https://docs.rafay.co/aiml/gpupaas/csp/overrides/), [Rate cards](https://docs.rafay.co/aiml/gpupaas/csp/rate_cards/), and [Orchestration Agents](https://docs.rafay.co/aiml/gpupaas/csp/orchestration_agents/)

#### Support for International Characters in User Profile Fields  
Support has been added for international characters in user-input text fields, such as first name and last name.

Users can now enter names that include Unicode/UTF-8 characters, including accented characters, hyphens, and apostrophes. Examples include é, è, ç, Jean-Pierre, and O'Connor.

Note  
Email address validation continues to follow standards-compliant ASCII validation rules, consistent with major identity providers.

#### Email Template Customization and Localization  
Support has been added for customizing and localizing system-generated email templates from the Operations Console.

Partners can now configure email templates under User Experience > Email Templates to customize email subjects, HTML body content, and localized versions of supported system emails. This enables partners to tailor platform-generated emails to match their branding, terminology, and customer communication requirements.

Note  
The default language configured on the Localization page is always included in outgoing emails. Additional localized content is appended after the default language content in the configured order.

For more information, please refer to documentation on [Email Templates](https://docs.rafay.co/aiml/gpupaas/csp/email_template/).

#### Email Notification Delivery Audit  
Visibility into email notification activity and SMTP submission status has been added to the Operations Console. Operators can now review whether platform-generated email notifications were successfully submitted to the configured SMTP provider.

Note  
Delivery status reflects submission to the configured SMTP provider. Final recipient delivery, spam filtering, mailbox-level rejection, or downstream provider handling may depend on the SMTP provider and recipient mail system.

#### Inventory and Data Center Management  
Support has been added for inventory data center import and export, bulk inventory import/export workflows, and pagination across inventory objects, simplifying inventory management at scale.

#### Server Type Allocation API  
A new API is now available to retrieve server type allocation per data center and across all data centers. This API can be used to monitor inventory capacity by server type, such as SERVER or GPU-SERVER, view allocation status counts, and review aggregate CPU, memory, and storage capacity.

**API path:** `GET /v2/sentry/paas/infrainventory/servertype-allocation`

#### Audit Trails for Server Allocation  
Inventory now shows the allocation history for servers under **Audit Trails**. Each entry records the previous allocation status, the new allocation status, who made the change, and when the change occurred.

#### Rows and Racks Support  
Inventory now supports **rows and racks** for NVIDIA rack server trays and base systems (for example, **GB200**), giving users a structured way to organize and manage physical server placement in the data center.

Inventory can be maintained in two ways:
- **Manual management** — create and manage rows and racks directly in inventory
- **BCM sync** — automatically sync NVIDIA rack inventory from BCM using a template

* * *

### Bare Metal Server SKU  
#### Serial Console Support  
Serial console support is now available for **BCM** and **Non-BCM** bare metal servers, enabling out-of-band access for troubleshooting and recovery.

Note  
Serial console requires a **gateway** deployed on a machine in the data center with network connectivity to the server BMC.

#### Cisco Hyperfabric Integration  
BMaaS can now be deployed with **Cisco Hyperfabric** for **BCM** and **Non-BCM** bare metal management. Data centers using Cisco Hyperfabric as their network fabric can provision bare metal instances using VPC and subnet resources managed by the Cisco network fabric.

For more information, see [Bare Metal Integrations](https://docs.rafay.co/aiml/bm_service/integrations/#cisco-nexus-hyperfabric).

#### Replace Instance  
Users can now recover a **BCM** bare metal instance using the **Replace Instance** action when hardware issues occur on the same server for example, after replacing faulty network interfaces, a failed BMC, or switch port details — without rebuilding the instance from scratch.

Update the interface, BMC, and switch port details on the existing server, then run **Replace Instance** to refresh the platform instance while keeping the same configuration intact.

For operator setup steps and the end-user workflow, see [Replace Instance](https://docs.rafay.co/aiml/bm_service/capabilities/#replace-instance).

Note  
**Replace Instance** is currently supported only for **BCM** bare metal servers. Support for **Non-BCM** bare metal is planned for a future release.

* * *

### VM SKU  
#### Serial Console Support  
Serial console support is now available for VMs enabling out-of-band access for troubleshooting and recovery.

Note  
Serial console access requires a gateway to be deployed on a machine in the data center with network connectivity to the server hosting the VMs.

#### VM Auto-Migration / Hypervisor Failure Handling  
Support has been added for automated VM migration when a hypervisor is determined to be down.

The system uses a dual health-check model before initiating auto-migration:
1. Hypervisor metrics that is pushed to the controller
2. The Gateway Agent performing ongoing proactive hypervisor checks, including SSH-based checks

If both signals indicate that the hypervisor is down, and auto-migration is enabled, the platform can automatically initiate VM migration without manual intervention.

Prerequisites:
- Gateway Agent must be deployed
- VMaaS Provisioner must be installed
- Hypervisors must be onboarded with the required collectors/metrics configuration

* * *

### Kubernetes SKU  
#### Kubernetes v1.36 Support  
New and existing Kubernetes SKU clusters can use Kubernetes v1.36 (provision and in-place upgrade).

#### Worker Node Pool Labels and Taints  
Support has been added for configuring **labels** and **taints** on worker node pools, enabling finer-grained workload scheduling and node isolation.

For more information, see [Worker Node Pool Labels and Taints](https://docs.rafay.co/aiml/k8s_service/capabilities/#worker-node-pool-labels-and-taints).

#### Service Account Issuer and API Audiences  
Support has been added for configuring **Service Account Issuer** and **API Audiences** for service account OIDC tokens, enabling secure token validation and improved compatibility with OIDC-based workload integrations.

For more information, see [Service Account Issuer and API Audiences](https://docs.rafay.co/aiml/k8s_service/capabilities/#service-account-issuer-and-api-audiences).

* * *

### Controller Enhancements and Improvements  
#### Harbor Registry Adoption  
Beginning this release, the Rafay Controller image registry is migrating from **Nexus** to **Harbor** for improved scalability, reliability, and lifecycle management of container images.
- Controller container images will be hosted in **Harbor**
- **Nexus** continues to host software artifacts such as RPM packages, APT repositories, and binary downloads
- Supported cluster images for this release are replicated to Harbor during upgrade for backward compatibility
- No user action is required for standard upgrades

Note  
If you need an older or previously unsupported image version after upgrade, it may need to be republished to Harbor before use.

#### Storage Backend Migration from OpenEBS to Longhorn  
For new bare-metal Rafay Controller deployments, the default storage backend is changing from **OpenEBS** to **Longhorn**.
- **New** bare-metal controller deployments use **Longhorn** as the default storage class
- **Existing** OpenEBS-based deployments continue to operate without change
- Migration of existing deployments to Longhorn is planned for a future release

This change improves long-term platform sustainability and addresses known security considerations with OpenEBS upstream support.

##### Data Disk Size Requirement  
With this release, the mandatory **data disk** size for new bare-metal controller deployments is updated from **1 TB** to **2 TB**.

| Requirement | Previous | Updated (this release) |
| --- | --- | --- |
| Data disk (`/data`) | 1 TB | **2 TB** (mandatory) |

The larger `/data` volume is required so controller application services have sufficient storage capacity and operate as expected. Plan new controller deployments with a **2 TB** data disk mounted at `/data`.

* * *

* * *
