Token Delivery Network for AI Inference | Rafay

Operate Token Delivery Networks for Distributed AI Inference

AI inference is becoming increasingly distributed as applications, agents, and intelligent systems demand faster responses, stronger data sovereignty, and lower latency. However, GPU infrastructure is often fragmented across data centers, cloud regions, telco edge locations, and sovereign environments.

A Token Delivery Network is a distributed AI inference architecture that brings AI model endpoints closer to users and solves this issue. However, Rafay's Token Delivery Network (TDN) takes this further. It enables providers to transform distributed compute into a unified AI inference platform that delivers governed, token-metered AI services, bringing inference closer to where it is consumed while creating new opportunities to monetize AI services.

The Rafay Platform delivers the operational workflows and controls that make it easy for providers to centrally deploy generative AI models across the network, manage endpoint lifecycle, meter token usage, etc. The result is a highly performant Token Delivery Network.

What Is a Token Delivery Network?

A Token Delivery Network, or TDN, is a distributed AI inference network that brings model endpoints closer to users, applications, agents, physical AI systems, and other consumers of generative AI models.

AI factory platforms make GPU infrastructure consumable through self-service access, governance, and automation. Token Delivery Networks extend these platforms by enabling distributed AI inference, delivering token-metered AI services from the optimal location based on latency, cost, capacity, and sovereignty requirements.

With all applications beginning to leverage generative AI models to deliver improved user experiences, the need for model endpoints to be closer to devices is driving many providers to invest in TDNs. Tokens can be allotted centrally but consumed across the network, resulting in the best of both worlds: Simplified governance with improved performance.

How TDNs Extend AI Factory Platforms by Enabling Distributed AI Inference

Token Delivery Networks are not standalone infrastructure layers; they build on an AI factory platform by extending AI services across distributed environments. This means the AI factory platform makes GPU infrastructure consumable through self-service access, governance, and automation. Then, Token Factory enables monetization through token-based usage models, and the Token Delivery Network ensures inference is delivered from the optimal location based on performance, capacity, cost, and sovereignty requirements.

TDNs vs. CDNs: From Content Delivery to AI Delivery

A Content Delivery Network (CDN) distributes static content such as images, videos, web pages, and application assets closer to users to reduce latency and improve performance. Raw data transfer is tracked, and the size of the data transferred serves as the usage meter for these interactions.

A TDN applies the same distributed architecture principles to AI inference, delivering model responses from the optimal endpoint based on latency, capacity, cost, and sovereignty requirements.

This is how the two differ:

Content Delivery Networks Token Delivery Networks
Deliver static or pre-generated content Deliver real-time AI inference
Optimize model response performance Cache content at edge locations
Deploy model endpoints across programmable edges Measure usage by data transfer volume
Route requests based on proximity and availability Route requests based on proximity, capacity, cost, and policy
Improve page load times Improve AI application responsiveness

GPUs vs Tokens: Why Token Economics Are Replacing GPU Economics

GPU hours measure infrastructure consumption, whereas tokens measure the value AI services deliver. Here's a quick overview of how they compare:

GPU-Hour Model Token-Based Model
Infrastructure-centric Service-centric
Low visibility Per-request visibility
Difficult chargeback Granular attribution
Limited pricing options Flexible consumption models
Capacity-driven Outcome-driven

With AI moving from experimentation to production, the need for monetization models that align with how customers actually consume AI will grow. Charging for GPU capacity treats inference as a hardware resource, while token-based pricing reflects the real unit of value: model interactions.

As such, token economics provide more granular usage visibility, flexible pricing models, and clearer cost attribution, enabling providers to package and monetize AI services through APIs rather than simply reselling infrastructure.

This development is good news for telcos, neoclouds, and sovereign AI providers, as this shift creates an opportunity to move beyond GPU utilization metrics and participate more directly in the economics of AI inference.

TDNs ♥️ Rafay

The Rafay Platform delivers a suite of capabilities – from edge cluster bringup and lifecycle management to multi-edge inference workload deployment across a dynamic set of programmable edges – that are required to power Token Delivery Networks. The Rafay Platform also tracks token usage at a granular level, enabling transparent monetization models that drive new revenue streams for providers.

With Rafay, providers can deploy, govern, meter, and operate a distributed set of inference endpoints across many locations.

Applications Are Becoming Model-Reliant

AI is becoming embedded across applications, devices, agents, and workflows. As this trend continues, more digital interactions will involve applications calling AI models to deliver better user experiences, automate work, and power real-time intelligence.

Tokens become the meter for how those model interactions are measured, governed, and monetized.

Model Interaction Performance Depends on Proximity

As more applications interact with AI models, the quality of the user experience depends on how quickly and reliably those interactions happen.

TDNs are designed to make model interactions more responsive, resilient, and scalable by distributing inference capacity closer to where AI applications are used.

GPU Supply Is Non-Contiguous

AI compute is increasingly getting deployed wherever power is available: in small metro data centers, telco edge locations, traditional carrier facilities, and sovereign sites.

Sub-1MW power sites are easier to secure across a geography than 100MW+ campuses. For inference, that distributed footprint can become an advantage because compute will organically get placed closer to where AI is used.

Providers Have the Right Assets

Telcos, neoclouds, and Sovereign AI providers already have pieces of the required footprint: distributed locations, regional infrastructure, network access, power, and customer relationships. The missing layer is software to make those assets programmable for AI inference.

The Monetization Model Is Shifting

GPU hours are an infrastructure metric. Tokens are a service metric. Providers that move from raw GPU resale to governed, token-metered AI services can participate more directly in the economics of AI inference.

Not Everyone Should be Forced to Invest in Hyperscaler Infrastructure

Anthropic and OpenAI may be able to invest in global GPU infrastructure, but most model builders would prefer to partner with TDNs to deliver their models to the market.

How Rafay Enables Token Delivery Networks

Rafay takes distributed GPU infrastructure into a unified AI inference platform that supports distributed AI inference, token-based pricing, and AI service monetization.

But how does this work in practice?

Rafay Capability What It Does in a TDN
Token Factory Converts GPU inference infrastructure into governed, token-metered AI services exposed through APIs. The product bridge from TDN thought leadership to deployable capability.
Programmable Edge Orchestration Deploys and manages inference endpoints across distributed data centers, sovereign regions, and edge-adjacent sites, making non-contiguous compute consumable as one coordinated platform.
Self-Service Portals and APIs Lets developers and customers consume AI services and model endpoints without manual provisioning. OpenAI-compatible APIs reduce integration friction.
SKU and Service Catalog Management Packages compute, model endpoints, agents, notebooks, and blueprints into catalog-based offerings operators can sell, tier, or white-label under their own brand.
Multi-Tenancy and Governance Enforces isolation, RBAC, quotas, policy, and secure access across teams, tenants, customers, and regions without sacrificing shared infrastructure efficiency.
Usage Metering, Chargeback, and Billing APIs Tracks token consumption, attributes cost and revenue, and feeds billing workflows — the commercial layer that makes token-metered AI services monetizable.

Industry-Specific TDN Use Cases

Telcos: From Connectivity to Tokens

Many telcos across the globe own points of presence with sufficient power and connectivity in place to be transformed into programmable edges to address AI use cases. Rafay helps telcos turn these assets into inference focused compute hubs that collectively form a TDN.

Sovereign AI Clouds: Local AI Services, Governed

Sovereign AI clouds require model inference to remain within jurisdictional boundaries, with data residency, compliance, and tenant isolation at the core. Rafay delivers the operational layer for local AI services from in-country infrastructure — with policy controls, tenant isolation, usage visibility, and data residency baked in.

Neoclouds: Turn distributed GPU capacity into an inference network

For training, contiguous GPU clusters matter. For inference, distributed pockets of compute can become an advantage. Rafay helps neoclouds pool non-contiguous compute across regions, deploy model endpoints consistently, and operate a TDN from a central control plane.

Frequently Asked Questions

What counts as a node?

A node is a physical or virtual server/machine.

Do you have any volume discounts?

Yes! As the number of nodes increases the price per cluster or per node decreases.

What about short-lived or ephemeral clusters?

Our customers love to experiment, and we don’t ding them for it. We don’t charge for node count spikes, but look at the running average of nodes in use when calculating usage.

Is there a difference between production and non-production pricing?

The management overhead for helping our customers operate dev vs prod clusters is effectively the same, so we treat all nodes the same.

What if I use more nodes than I’ve licensed?

Rafay has a true-up forward policy, meaning that we don’t carry out chargebacks for scenarios where the consumption in a completed billing cycle exceeded the licensed count. If the new, steady-state number of nodes is expected to be higher, our customer success team will discuss the situation with you, and take steps to adjust billing accordingly for the next billing cycle.