Mohan Atreya | Rafay

Mohan Atreya

One-Click Digital Twins: Deploying the NVIDIA Omniverse DSX Blueprint using Rafay

See how Rafay transforms the NVIDIA Omniverse DSX Blueprint into a one-click, self-service digital twin offering with governed GPU access, multi-tenancy, and automated session management.
Read Now

NUMA-Aware GPU VMs: How NVIDIA's Reference Architecture Actually Fixes It

Learn how NVIDIA's NUMA-aware GPU VM reference architecture improves GPU performance by aligning CPUs, memory, and GPUs to reduce latency and maximize throughput.
Read Now

Private Token Factories: How Rafay and Protopia AI Let Sensitive Workloads Run on Shared GPU Capacity

Rafay and Protopia AI eliminate plaintext exposure, letting regulated enterprises finally run sensitive workloads on shared GPU inference infrastructure.
Read Now

Why is your GPU VM slower Than Its Twin?

Two identically configured GPU VMs can deliver very different performance because the GPU, CPU cores, and memory may not be placed on the same NUMA node
Read Now

Serving LLMs on Arm: Running Rafay Token Factory on NVIDIA DGX Spark

Learn how Rafay Token Factory turns NVIDIA DGX Spark into a managed, multi-tenant LLM serving endpoint with Arm-native Kubernetes, metering, governance, and OpenAI-compatible API access.
Read Now

Automated GPU Health Monitoring with NVIDIA NVSentinel on the Rafay Platform

Every GPU node monitored. Faulty nodes automatically quarantined and remediated. The Rafay Platform and NVIDIA NVSentinel make that a fleet-wide guarantee, not a per-cluster aspiration.
Read Now

AI Factories Will Be Won on Efficiency: Why the Rafay + Kubex Partnership Matters

GPU costs are rising. Workloads are unpredictable. Platform teams are stretched. The next frontier of enterprise AI is operating efficiently at scale. That is the problem Rafay and Kubex are solving together.
Read Now

Compute Domains: Bringing Multi-Node NVLink Awareness to Kubernetes

Learn how compute domains and multi-node NVLink enable high-performance, distributed GPU workloads in Kubernetes, improving scalability, resource utilization, and AI infrastructure efficiency.
Read Now

NVIDIA Dynamo: Turning Disaggregated Inference Into a Production System

Discover how NVIDIA Dynamo turns disaggregated inference into a production-ready system, enabling scalable, efficient AI services with better resource utilization and operational control.
Read Now

Introduction to Disaggregated Inference: Why It Matters

Learn how disaggregated inference improves GPU utilization, scalability, and cost efficiency by separating compute, memory, and serving layers—enabling more flexible, self-service AI infrastructure.
Read Now

Fine Tuning as a Service using Rafay and Unsloth Studio

Turn Unsloth Studio into a repeatable, app store style fine tuning experience with Rafay, enabling one click deployment without Kubernetes or MLOps complexity.
Read Now

From Docker Image to 1-Click App: Enabling Self-Service for Custom Apps

Turn Docker images into 1-click, self-service apps, securely delivered across multi-tenant Kubernetes environments with built-in governance and control.
Read Now

Adding New Language Support to the Self Service Portal in 5 Mins

Add any language to your Self-Service Portal in minutes. Deliver a localized, frictionless experience for global AI teams with Rafay.
Read Now

OpenClaw on Kubernetes: A Platform Engineering Pattern for Always-On AI

A deep dive into OpenClaw as a gateway-centric AI runtime and how platform teams can deploy, secure, and scale it as a governed service on Kubernetes.
Read Now

Developer Pods for Platform Teams: Designing the Right Self-Service GPU Experience

A deep dive into how platform teams use SKU design to transform raw GPU infrastructure into intuitive, self-service Developer Pod experiences for AI builders.
Read Now

Developer Pods: A Self-Service GPU Experience That Feels Instant

From a simple form to a live SSH session in ~30 seconds. This is what self-service GPU access actually looks like with Rafay Developer Pods, no Kubernetes knowledge required.
Read Now

Instant Developer Pods: Rethinking GPU Access for AI Teams

Rafay Developer Pods eliminate the ticket queues and bloated VMs holding AI teams back — GPU-ready in ~30 seconds, Kubernetes-powered, complexity-free.
Read Now

Scaling Trust: The Fortanix and Rafay Integration for Enterprise Confidential AI

Learn how the Fortanix and Rafay integration enables confidential AI for enterprises—protecting sensitive data while running AI workloads on secure, governed GPU platforms.
Read Now

NVIDIA AICR Generates It. Rafay Runs It. Your GPU Clusters, Finally Under Control

NVIDIA AI Cluster Runtime (AICR) simplifies AI infrastructure deployment. Learn how Rafay operationalizes GPU clusters with governance, self-service access, and platform automation.
Read Now

Run nvidia-smi on Remote GPU Kubernetes Clusters Using Rafay Zero Trust Access

See how infrastructure operators can securely validate GPU health in remote Kubernetes clusters by running nvidia-smi using Rafay’s Zero Trust Kubectl Access workflow.
Read Now

How Rafay Helps GPU Clouds Run Complex Hackathons at Scale

Discover how Rafay enables GPU cloud providers to run large-scale hackathons by instantly provisioning secure, ready-to-use GPU developer environments for thousands of participants.
Read Now

How GPU Clouds Deliver NVIDIA Run:ai as Self-Service with Rafay GPU PaaS

Learn how Rafay GPU PaaS enables GPU Clouds to offer NVIDIA Run:ai as a fully automated, multi-tenant managed service delivered through self-service with lifecycle management and turnkey deployment.
Read Now

GPU Cloud Billing: From Usage Metering to Billing

In this blog, we take the next step toward a complete billing workflow—automatically transforming usage into billable cost using SKU-specific pricing.
Read Now

Goodbye to Ingress NGINX – What Happens Next?

The community Ingress NGINX project is entering end-of-life in March 2026. Discover what this means for Kubernetes users and why you’ll need to migrate, what alternatives exist (Gateway API, Traefik, etc.), and how to plan your transition smoothly with minimal disruption.
Read Now