AI & ML FAQs | Rafay AI Infrastructure Platform

GPU/AI/ML FAQs

Find answers to common questions about Rafay's neocloud, sovereign AI cloud, and enterprise infrastructure orchestration and operationalization offerings and solutions and learn how they can benefit you.

What counts as a node?

A node is a physical or virtual server/machine.

Do you have any volume discounts?

Yes! As the number of nodes increases the price per cluster or per node decreases.

What about short-lived or ephemeral clusters?

Our customers love to experiment, and we don’t ding them for it. We don’t charge for node count spikes, but look at the running average of nodes in use when calculating usage.

Is there a difference between production and non-production pricing?

The management overhead for helping our customers operate dev vs prod clusters is effectively the same, so we treat all nodes the same.

What if I use more nodes than I’ve licensed?

Rafay has a true-up forward policy, meaning that we don’t carry out chargebacks for scenarios where the consumption in a completed billing cycle exceeded the licensed count. If the new, steady-state number of nodes is expected to be higher, our customer success team will discuss the situation with you, and take steps to adjust billing accordingly for the next billing cycle.

How much does Enterprise Support (24x7x365) cost?

Enterprise Support is available at an additional fee equaling 20% of the cluster or node subscription.

Do you have EDU or GOV discounts?

Yes, please contact sales for more information about discounts for educational institutions and government agencies.

What does Rafay do or provide around AI/ML or cloud-native adoption?

Rafay provides infrastructure orchestration and workflow automation for enterprises, cloud providers, neoclouds, and sovereign AI clouds. The Rafay Platform delivers a Platform-as-a-Service (PaaS) experience that enables companies to create customized compute environments for developers and data scientists. Rafay’s platform enables faster development and deployment of new capabilities while maintaining necessary controls and guardrails. By simplifying the process of implementing complex platforms, Rafay reduces the need for large teams of experts. In essence, Rafay streamlines cloud-native and AI/ML adoption by offering a ready-to-use platform that balances speed, efficiency, and security for businesses.

Does Rafay offer a GPU PaaS?

Yes, Rafay provides infrastructure orchestration and workflow automation for cloud-native (Kubernetes) and AI use cases for enterprises, cloud providers, neoclouds, and Sovereign AI clouds. Rafay helps companies deploy a Platform-as-a-Service (PaaS) experience that supports both CPU-only and GPU-accelerated compute environments. Platform teams can quickly set up and deliver customized self-service experiences for developers and data scientists, typically within days or weeks. This flexible platform allows end-users to easily access the computational resources they need, whether it’s standard CPU processing or more powerful GPU capabilities. Rafay’s solution streamlines the deployment and management of diverse computing environments, making it easier for organizations to support a wide range of applications, from standard software to complex AI/ML projects.

What does Rafay offer for ML workbenches?

Rafay provides curated ML workbenches that offer developers and data scientists an experience similar to Amazon SageMaker or Google VertexAI, but at a more competitive price point. The platform includes out-of-the-box services such as Notebooks-as-a-Service, with pre-compiled environments featuring TensorFlow, PyTorch, and other popular libraries for immediate productivity. For those preferring a job-based model, Rafay offers Ray-as-a-Service, allowing data scientists to focus on their work without dealing with infrastructure complexities. Advanced teams can opt for a Kubeflow-based ML workbench, which manages pipelines, experiment tracking, and model repositories. These solutions enable data science teams to work efficiently with their preferred tools while Rafay handles the underlying infrastructure management.

What does Rafay offer for GenAI playgrounds?

Rafay provides a controlled, cost-effective Generative AI playground for organizations new to GenAI. This environment allows data scientists to train, tune, and serve GenAI models, enabling efficient experimentation and development without significant investment or infrastructure complexity. It’s ideal for businesses looking to explore GenAI capabilities while managing costs and maintaining control over their AI initiatives.

Who uses Rafay's platform for AI/ML initiatives?

Rafay’s AI/ML platform is utilized by various organizations, particularly in the financial services sector. We’re also collaborating with major GPU vendors for specialized use cases. A notable public example of a company using our AI/GPU stack is MoneyGram, a global leader in cross-border P2P payments and money transfers.

How does Rafay’s platform accelerate time-to-value for AI/ML projects?

Without Rafay, platform teams implement complex platforms internally over multiple years and with large teams of experts. With Rafay, platform teams can deliver a finely tuned PaaS experience to internal users in weeks.

How does Rafay ensure compliance and governance for enterprise AI initiatives?

Rafay applies its proven governance and control features, originally developed for cloud-native projects, to AI/GPU initiatives. These capabilities include blueprinting, access management, chargebacks, and auditing/logging. This approach ensures that enterprises can maintain compliance and control over their AI projects, just as they do with other cloud-native initiatives. By leveraging these established features, Rafay helps organizations accelerate AI adoption while maintaining the necessary governance standards, ultimately leading to increased revenues and lower total cost of ownership for both cloud-native and AI/ML projects.

How does Rafay's platform streamline AI/ML infrastructure management for enterprise adoption?

Rafay enables enterprise platform teams to deliver a PaaS experience for GPU resources, both on-premises and in the cloud. The platform offers a cost-effective alternative to services like Amazon SageMaker or Google VertexAI, providing ML workbenches with similar functionality. Rafay’s self-service model and hierarchical experience sharing allow platform teams to selectively offer compute and ML workbench experiences to different teams, optimizing access to expensive GPU resources. Additionally, the platform includes chargeback capabilities to ensure fair cost allocation among internal teams. This comprehensive approach simplifies AI/ML infrastructure management, accelerating enterprise adoption while maintaining cost control and resource efficiency.

Does Rafay provide AI/ML workbenches and other tooling?

Yes, Rafay offers a comprehensive suite of AI/ML tools. The platform provides out-of-the-box workbenches based on Kubeflow and KubeRay, delivered as fully managed services. This allows users to access sophisticated AI/ML platforms without dealing with infrastructure complexities. Additionally, Rafay includes a low-code/no-code framework that enables partners to rapidly develop and deploy specialized AI solutions such as verticalized agents, co-pilots, and document translation services. This combination of ready-to-use workbenches and a flexible development framework streamlines the adoption and customization of AI/ML tools for various enterprise needs, accelerating time-to-market for new AI capabilities.