Blog / AI Infrastructure

What an inference layer for India looks like

An inference layer for India routes AI for companies, builders, and government in rupees, without lock-ins, leaked data, or imported intelligence.

By Guru Vishwas · 7 min read · Updated 8 Sep 2026

India's inference layer

An inference layer for India is a sourcing and routing system that sits above model providers. It lets us pick and pay for AI that’s tailored for India’s needs to protect sovereignty and re-risk from usual conundrums related to AI deployment.

Most Indian companies are still figuring out how to transform themselves with AI. As are solo developers, agencies, and government services that have to answer citizens in a myriad of languages and systems. They run a patchwork: a closed-source API here, a global aggregator there, a half-empty private GPU rack in a data centre. That patchwork will not survive the growing demand India will have for AI. Here’s how we infer today, why it fails, what a layer for India has to do, and why the country cannot keep importing intelligence.

How India runs AI inference today

India mostly infers in three ways: a direct key from a frontier lab such as OpenAI, Anthropic, or Google; a marketplace in the OpenRouter mould; or self-hosting open-weight models on private GPUs. None of these is an inference layer tailored for India. They are procurement choices.

The first setup is a direct key or an enterprise account at a frontier lab. Closed-source, billed in dollars, one model family, one terms-of-service. It is the fastest way to ship a demo for a Bengaluru tech team, a freelancer with a client bot, or a company piloting a workflow.

The second is an aggregator: many models, one API, priced in dollars, running on foreign infra, designed for a developer with a credit card and no option to pay through UPI, no rupee invoice, no passing on discounts for aggregated demand.

The third is self-hosting. Take an open-weight model, put it on private GPUs, keep the data on prem. This is the residency answer for a bank or government service. It requires upfront capital that a startup, a hospital or a local business cannot afford to run without a dedicated IT team.

The layer is what should sit above all three: a way to pick, route and pay for AI that satisfies your requirements without tradeoffs and worrying about sourcing at best price and performance every quarter.

Setup Typical user Pays in Where data goes Lock-in
Direct lab key Startups, enterprises USD Provider cloud High
Global aggregator Developers, agencies USD Provider cloud Medium
Self-host Banks, Healthcare Capex On prem High (hardware)
Inference layer Companies, builders, government INR Set by policy Low

What is wrong with today's inference stack

Today’s stack routes buyers to foreign labs, leaks IP, adds silent forex on top of platform margins, and ignores India's unique needs and legal recourse for data breach.

Closed endpoints train on proprietary data (despite them promising otherwise, unless proven cryptographically). Prompts are not only chat. For a company - they are intellectual property, confidential contracts, customer information, product plans.

For an individual - they are client work, medical notes, exam prep, a family firm’s books.

For the government - they are Aadhaar-linked records, welfare claims, land files, grievance text.

Sending them to a vendor you do not control is a data decision. India’s privacy regime is about to make it a sovereignty policy and a national security issue.

They also have you hooked into their ecosystem like they did with operating systems like Windows, macOS and Linux. The network effects grow with each integration into your systems. New models arrive every quarter, with different costs and performances. The team that bet on one lab spends the next quarter redoing evals, renegotiating and waiting on IT clearance to push to production.

Dependency is not only commercial. It is geopolitical. The United States and China are treating models and chips as bloc assets. A production path that sits on a foreign API can be cut off by a terms change or an export rule nobody in India controls. That could be catastrophic like it already played out in the limited release of Mythos.

Then there is the bill. You pay the lab. You pay the platform. You pay forex on top of both margins. None of those parties is optimising for India: pay in rupee rails, aggregate models for India native workflows, negotiate lower prices with providers and genuinely look out for the general benefit of an Indian customer.

Self-hosting looks like a solution but the answer is more nuanced. GPUs that are not utilised to capacity are among the most expensive idle assets a company can own. They freeze you to the installed fine-tuned model and when a better open-weight lands, the upgrade is a capex cycle, not a config change. And one box still cannot cover the mix. Document processing, hard reasoning, voice, video generation: different jobs need different models. Cost and performance are a routing problem.

What an inference layer should do for India

An inference layer for India is not just “sovereign models served from India”. It is a system with five jobs: India-specific models, neutrality across providers, a graded catalog, traffic for local inference and privacy controls. Consumers get those jobs from the same layer.

It has to offer models for India’s actual work: multiple languages, credit and insurance, hospitals, manufacturing, agriculture, public services, consumer apps that are highly personalised. A generic global catalog will force every institution to become a model host. IndiaAI mission has already mapped hundreds of public-sector use cases. Those only ship if we can call a layer instead of standing up a GPU farm for each deployment.

It has to be neutral across models and providers. American frontier when the task is hard. Open-weight, including Chinese, when the task is volume. Indian sovereign when the data, or the application, or the regulator requires it. Neutrality is how you become model agnostic and resilient. Public procurement cannot pick a bloc. Neither should a builder whose app has to keep working if one lab changes its terms.

It has to curate a graded catalog, not dump every weight on the internet. Grade by cost, reliability, latency, residency, and sovereignty. The buyer should just pick a point on that gradient rather than managing vendor relationships with multiple providers.

It has to boost local inference providers, not route around them. Indian data centres only become a national advantage if production traffic actually lands here. That includes traffic from consumer apps and services, not only from enterprise contracts.

It has to give users the controls they now demand, such as privacy and data retention without turning those controls into a lock-in. The customer should be able to change models without losing the guarantee. OpenAI-compatible access is the on-ramp. The product is everything after the base URL changes.

Why India can't keep importing inference

India cannot keep importing inference because demand is about to jump, the data is too sensitive, forex reserves deplete and geopolitics can cut the pipe. The country already knows a similar pattern: import crude, refine at home, serve the domestic market, export the surplus. Inference can follow oil’s playbook.

Inference is about to become the main AI workload. McKinsey expects global inference compute to grow from about 21 GW in 2025 to 93 GW by 2030. In India, Google and Inc42 put the AI opportunity at $126 billion by 2030, with enterprise AI rising from $11 billion to $71 billion. That number still understates the load. It does not fully count all the emerging use cases that will have to run at 1.5 billion population scale.

The money is in the middle of the chain. Import the best weights, run them on India’s data centres, serve a huge home market first and then export surplus intelligence the way we already export software through more than 1,700 GCCs.

Data-centre capacity has quadrupled to around 1.7 GW, the largest market in Asia-Pacific outside China. A first wave of sovereign models including Sarvam, BharatGen and Gnani are in the market. Until we also make hardware locally, we can capitalize on serving intelligence. A country that lets another bloc own its production intelligence will eventually pay for it in currency and capability.

The way through is not to pretend India will train every frontier model. It is to embrace open-source, build sovereign models where India is distinct, and put a layer over both so everyone can use them. Solve for India: language, sectors, rupees, residency. The same layer can be sold to the world.

That is what an inference layer for India looks like. It is what Indos is building: sourcing, routing, and serving intelligence, taking the complexity off the plate of anyone consuming AI.