May 2026 - Rafay Product Documentation
May
v3.1-39
15 May, 2026
Token Factory
This release enhances Rafay’s Token Factory with powerful new capabilities to simplify model experimentation, governance, and deployment flexibility. Key updates include an integrated Playground for testing and refining prompts, support for leading frontier models, tier-based rate limiting for controlled access, and expanded platform support.
📌 Important Upgrade Message
Existing Token Factory deployments upgrading from v3.1.38 to v3.1.39 may require additional upgrade steps for existing compute clusters.
This applies only to environments with:
- Existing compute clusters created before
v3.1.39 - Existing model deployments running on those compute clusters
Additional upgrade steps may be required when the compute cluster does not have sufficient resources for rolling upgrades.
Review the Upgrade Instructions for Existing Token Factory Compute Clusters.
OpenShift Support
Token Factory can now be deployed on Red Hat OpenShift clusters, in addition to Rafay Kubernetes Service (MKS). This provides service providers and enterprises with greater flexibility to run Token Factory within their existing OpenShift environments.
Integrated Playground
Token Factory now includes an integrated, secure web-based Playground that enables users to test, prototype, and refine prompts and model parameters such as temperature and max tokens, without writing code.
The Playground allows teams to compare models, evaluate outputs, and iterate on AI use cases in a controlled environment before promoting them to production.
Frontier Model Support
Support is being added for deployment of large-scale frontier models requiring more than 8 GPUs across multiple nodes, enabling production-grade execution of next-generation LLM workloads.
Disaggregated Serving and KV Cache Management
- Disaggregated serving allows pre-fill and decode workloads to run across separate GPU groups to improve resource utilization and scalability
- Enhancements to KV cache management enable cache offloading from GPU memory to host memory, while providing visibility into cache performance. This helps reduce GPU memory pressure and improves efficiency for large-scale and multi-tenant deployments
ARM Only GPUs
Token Factory now supports model deployments on ARM only AI Infrastructure (e.g. NVIDIA GB200)
Tier-Based Rate Limiting
This enables the definition of service tiers (for example, Gold, Silver, and Bronze) with predefined token capacity, rate limits, and quotas. These tiers can be assigned to organizations, allowing providers to offer differentiated plans and manage resource usage more effectively.
Organization-Level Token Access Controls
Token Factory now allows providers to control token usage at the organization level. Providers can stop token consumption for a specific organization for use cases such as payment enforcement or free trial limits. External systems can use Rafay APIs to update organization status, which stops further token consumption.
PVC-Based Model Loading for Token Factory
Token Factory now supports loading models directly from Persistent Volume Claims (PVCs) within the cluster, enabling the use of pre-cached or locally stored models without relying on external object storage. This improves model startup times and GPU utilization, especially for large model deployments and environments without high-performance shared storage.
Audit Logs
Expanded audit logging delivers improved visibility into both administrative actions—such as model, deployment, provider, cluster, and endpoint operations—and end-user activity, including login history, API key usage, and token consumption. These enhancements support governance, security monitoring, and compliance requirements.
Enhanced End User Model Card
The Model Card experience for end users has been enhanced to include richer model information such as custom metadata, endpoint type (shared or dedicated), performance metrics, supported modalities, and available model variants. This provides greater transparency and helps users select the most appropriate model for their use case.
Clone Model Deployments
Operators can now clone existing model deployments and modify parameters as needed, making it easier and faster to create similar deployments without configuring them from scratch.
Localization Support
Token Factory now supports localization in the end-user Developer Hub, delivering a consistent multi-language experience with the rest of the end user portal.
Alerts
Token Factory now supports configurable alerts with integrations for Slack, webhooks, and email, enabling teams to get notified of important events through their preferred channels.
Dashboard Enhancements
Additional metrics have been added to the Token Factory dashboard to provide deeper visibility into inference activity, improving observability and operational awareness.
Operations Console — Revamped Navigation & Layout
The Operations Console has been reorganized into clearly defined sections to improve discoverability and streamline administrative workflows.
| Section | Items |
|---|---|
| Inventory | Data Centers |
| Catalog | SKU Configuration, Token Factory |
| Tenants | Organizations, Role Management |
| Usage & Billing | Reservations |
| User Experience | Portal Navigation, Localization, Theme & Branding |
| System | Audit Logs, SMTP, Email Notifications |
Key Changes include:
New Overview landing page
A new Overview page is now available as the home screen for the Operations Console. It provides a real-time snapshot of platform usage and offers quick access to key areas of the console for easier navigation.
SKU sharing consolidated under Catalog → SKU Configuration
SKU sharing for both Compute and Service SKUs is now managed directly within the SKU Configuration section.
GenAI renamed to Token Factory and moved to Catalog
The GenAI section has been renamed Token Factory and relocated under Catalog.
Tenant management grouped under Tenants
Organization management (create, edit, manage orgs) and role-based access control are now consolidated under the Tenants section.
Portal customization moved to User Experience
Portal Navigation, Localization, and Theme & Branding settings are grouped together under User Experience, making it easy to find all end-user portal customization options in one place.
System-level settings moved to System
Audit Logs, SMTP configuration, and other platform-level administrative settings are now located under the System section.
Partner Search
A search capability has been added to the Partners page, allowing administrators to quickly locate partner organizations by name.
Partner Admin Role
Partner Admins can now configure and manage all partner-level settings directly through the UI. Previously, these configurations were restricted and could only be managed by Super Admins.
Workspace Read Only Collaborator Role
Users can now invite users with a workspace read-only collaborator role to view instances, related configurations, and status in a particular workspace. These users cannot create, modify, or delete any instances.
Tenant Dashboard
Instance names in the drawer are now clickable, allowing direct navigation to their details in the Usage Dashboard.
Profile to SKU Update in Operator Consoles
References to “Profiles” were previously updated to “SKU” in tenant organization portals (e.g., Developer Hub). This release extends the update to operator consoles (e.g., Ops Console, SKU Studio), replacing all instances of “Profile” with “SKU.” The label is also now customizable via the profile_label setting in Partner Portal Navigation configuration.
Audit Logs
Enhanced audit logs now capture SKU sharing action (compute or service) with tenant organizations in the Ops Console.
Log Archive Migration
Existing log archive tooling is being migrated to the enhanced log archive framework as part of this release.
Important
During the migration, historical audit logs collected through the legacy log archive mechanism will not be visible in the UI. These logs continue to be retained within the controller and remain accessible through backend access if required. Contact Rafay for instructions on accessing legacy audit logs from the backend.
GPU Reservations
Service providers can now enable tenants to pre-purchase guaranteed capacity at a discounted reserved rate compared to on-demand resources. Tenants can pre-purchase a specific quantity of GPUs (GPU type/count) for each compute-type or service-type. The number of pre-purchased GPUs represents the customer's reserved quota at a set reservation rate.
For example, a customer requires 16 H200 GPUs on bare-metal servers to complete a specific model training or fine-tuning task over a six-month period. They want to ensure that GPUs are available to the team tasked with that goal during the reserved time.
| Aspect | On-Demand | Reserved |
|---|---|---|
| Cost | Charged per time unit of use | Fixed for a three-month, six-month, one-year, etc. |
| Cost variability | Cost fluctuates with usage | Cost is fixed irrespective of usage |
| Availability | Not guaranteed, depends on capacity | Always guaranteed |
Important
Once a reservation request is accepted by the service provider, billing will begin immediately for the accepted quantities, regardless of whether instances are running.
Partner-Level Reservation Limits
Service providers can now configure partner-level reservation limits to control how much GPU capacity can be allocated to reserved contracts for a specific data center, inventory type, and GPU type. These limits apply across all organizations under the partner.
This allows service providers to reserve only a portion of their total GPU inventory for contracts while keeping the remaining capacity available for on-demand usage.
For example, if a partner has 1000 A100 GPUs, they may choose to allow only 600 GPUs to be reserved for contracts and keep the remaining 400 GPUs available for on-demand consumption.
When a reservation is created, the system validates the requested quantity against the configured reservation limit. If the limit has already been reached, the reservation request is blocked.
Reserved-Only Mode
Service providers can enable a reserved-only mode, restricting tenants to launching instances only from their reserved capacity. In this mode, on-demand instance creation is disabled.
OIDC Support
Support is being added for OIDC-based identity providers (IdPs), enabling customers to integrate identity providers that rely on OpenID Connect.
Key characteristics:
- OIDC workflows closely align with existing SAML-based workflows
- Role assignment and access control mechanisms remain unchanged
Global IDP Support
A Global IDP feature is now supported for partners, enabling a single SAML or OIDC Identity Provider configured in the default organization to authenticate users across all organizations under the partner.
Key characteristics:
- Supported for Okta, Keycloak, PingOne, and Auth0 with both SAML and OIDC
- Organization mapping is achieved through the External ID configured in the Ops Console
- Tenant-specific IDPs take precedence over the Global IDP when both are configured
Bare Metal as a Service (BMaaS)
Reinstall OS for Non-BCM Bare Metal
This release adds a new Reinstall OS action for Non-BCM (PXE/Metal3-based) bare metal nodes. Operators can now reimage an existing bare metal server in place by providing the Image URL to a QCOW2 whole-disk image along with the image checksum without having to fully decommission and re-enroll the server.
The action is available directly from the instance Actions menu, alongside Start, Stop, Reboot, and Delete:
Note
Reinstall OS is currently supported only for Non-BCM bare metal nodes.
Upstream Kubernetes (Rafay MKS)
This release introduces several enhancements to Rafay MKS, including:
- Support for provisioning new clusters and upgrading existing clusters to Kubernetes 1.35
- Introduction of Platform Version 1.2.0, which includes updated core components such as etcd 3.5.24, required for compatibility with Kubernetes 1.35
Internationalization & Localization
Info
The scope of localization is limited to portal navigation components (e.g., the left-hand navigation menu) and portal interface content, including UI labels, options, help text, and user guidance within the end-user Dev Hub portal.
A previous release introduced default language support for English, Turkish and Japanese.
With this release, the default supported language list has been expanded to include French, Spanish, and Arabic, including support for right-to-left (RTL) languages.
Service providers can now also control which languages are available to end users in both the End User Portal and the login experience using the Enabled Languages setting. This allows providers to restrict language options based on their audience.
For example, a service provider can choose to make only English and French available to end users, even though additional languages are supported by default.
With this release, the default supported language list is now being expanded to include French and Spanish.
Known Issues
Important
A patch release is actively being worked on to address the known issues.
- Instances launched from one tenant org are consuming reservations belonging to another organization when no matching reservation exists.
- During OIDC based SSO user login, the redirect_uri query parameter is being sent with http instead of https when the controller is setup behind AWS ALB.
- NVIDIA's NGC now blocks chart downloads when the "--insecure-skip-tls-verify" flag is used to download the GPU operator chart when used with controller deployments using self signed certificates.