← All AI Notes
India AI JourneyIssue #03

Sarvam Epoch 2026: Models, Infrastructure and Applications

A complete report on India's first frontier AI developer conference — every announcement, every benchmark claim, and what the whole-stack bet means for India's AI journey

Bhaveshkumar Choithram Dharmani

Founder, AIVidhya4Sarvam | AI Mentor, Researcher & Ecosystem Builder

1 August 202634 min read8,800 words

HOW TO READ THIS NOTE

This is a conference report, not an explainer. It is built from the full keynote of the Sarvam Epoch 2026 Builder Edition and organised the way the day itself was organised — framing, models, infrastructure, applications, and the sovereign/assistive work that closed it.

Two things to keep in mind throughout. First, almost every performance figure here is Sarvam's own claim, made from the stage, and has not been independently verified — the Benchmark Scorecard collects all of them in one place precisely so they can be scrutinised together rather than absorbed one at a time. Second, this note deliberately separates what shipped from what was promised; the "What Is New vs. What Is Updated" table exists for that reason.

Want to watch it yourself? The References — Primary Source section at the end links the full keynote livestream, and the Appendix — Jump to the Video maps every section of this report to its exact timestamp in the 3 hour 21 minute keynote, so you can verify any claim against the source in one click.


Sarvam Epoch 2026 — Key Facts

DetailValue
EventSarvam Epoch, Chapter 1 — Builder Edition
Date & venue30 July 2026, Bengaluru (Enterprise Edition followed on 31 July)
SponsorsPowered by NVIDIA, in association with AWS, supported by HCLTech
Keynote lengthApproximately 3 hours 21 minutes
Headline commitmentA trillion-plus parameter frontier model, trained from scratch in India
ComputeLargest Blackwell cluster in India, scaling toward ~10,000 accelerators
Products launchedIndus — a six-product agentic platform under one credit system
New organisational movesSan Francisco research office; Shri Devendra Singh Chaplot as advisor

Table 1: The Epoch 2026 Builder Edition at a glance (Sources: Sarvam Epoch keynote; Inc42; ANI; Business Standard, July 2026)


Executive Summary

On 30 July 2026, Sarvam AI held the Builder Edition of Epoch, its first developer conference, before a room of developers, researchers and founders in Bengaluru. Over roughly three hours of keynote, the company made one of the densest sets of product announcements in the short history of Indian AI — and, more importantly, made an argument about why they belong together.

The argument is that the unit of competition in AI is no longer a single model but a complete stack: the infrastructure that trains and serves models, the models themselves, the harnesses that make models useful, and the applications people actually touch. Sarvam announced across every one of those layers on the same day. Cofounder Dr. Pratyush Kumar opened by naming three commitments — build at the frontier, create value for everyone, and think in decades rather than quarters — noting that in five decades of technology revolutions, India has not built a platform that outlasted multiple decades, and that AI is the opportunity to change that.

What followed was organised into three movements. On models, Sarvam upgraded its flagship 105B, released a new state-of-the-art speech recognition family including the first credible multi-speaker Indic model, pushed its document-vision model into a new performance region, launched an expressive text-to-speech model, opened a custom model training service, and committed to building a trillion-parameter frontier model in India. On infrastructure, it launched Sarvam Inference — an India-hosted "token factory" running on what it says is the largest Blackwell cluster in the country — under an explicit goal of token sovereignty. On applications, it launched Indus, a six-product agentic AI platform spanning voice, work, content, documents, coding and inference, alongside on-premise deployments for the armed forces, Kaze AI smart glasses for the visually impaired, and a voice-interaction product called Kivi.

Read together, the day was Sarvam declaring that it intends to be the horizontal AI platform of India — not a model vendor, and not a single-product company.

THE SINGLE LINE THAT CAPTURES THE DAY

"Sovereignty should not be a tax." Running AI in India, on Indian infrastructure, with Indian-language coverage, should cost less than importing it — not more. Nearly every pricing claim made on stage was built to prove that point.


Part I — The Framing: Why Build Models At All?

Dr. Kumar opened the keynote by taking on a question Sarvam evidently still hears often: with excellent open models like Kimi K3 freely available, why should India train its own? He called the suggestion that it should not "absolutely the wrong thing to do," and gave two reasons.

The first is what he called accumulated judgment — the taste a team builds only by repeatedly training models at the frontier. He traced Sarvam's own three-year arc as evidence: the company built on Llama and was, he said, the only company in Asia to contribute to the Llama 3 series, learning how language skills get added to models; it then pre-trained a 2-billion-parameter model from scratch when the first thousand Hopper GPUs landed in India, and worked with Infosys to make it useful for IT services; it built Sarvam-M on top of a Mistral small model, which he described as the first thinking Mistral variant; and over the past year it pre-trained a 100-billion-plus parameter model from scratch on 4,000 Hopper GPUs. Each step, he argued, deposited judgment that cannot be bought.

The second reason is strategic. "You cannot bet the country's future on fair access to open models," he said, pointing out how quickly the ground moves: one quarter is about token-maxing, the next about complaining that closed models are expensive, and "maybe next year there's a quarter of closed Chinese models." He was careful to frame this as additive rather than protectionist — India should absolutely use the best models available, closed or open, while also retaining the ability to build its own. And he extended the ambition outward: India has built global companies from India before, and the same can be done in AI.

Dr. Kumar then introduced the mental model that organised the rest of the day — the AI stack, and the claim that a model builder sees every layer of it differently from a company that only rents models.

Figure 1 — The Sarvam Stack

LayerWhat sits there
Users & ApplicationsIndus agentic suite, Kaze smart glasses, Kivi — what people actually touch
HarnessesThe "operating system" around the model: skills, memory, tools, sub-agents, routers
DatasetsCurated Indian law, governance, finance, vernacular and enterprise context
ModelsSarvam 105B, 7B/70B, Saaras V4, Bulbul V4, Vision 2.0, trillion-parameter model
InfrastructureBlackwell GPU clusters, Sarvam Inference "token factory", sovereign data centres

Figure 1 — The layered stack Dr. Kumar presented. Sarvam's argument is that having trained models on thousands of GPUs changes how you build every other layer: what infrastructure needs to do, where a harness must bear load, and which data actually matters.

He closed the framing with an invitation rather than a boast. China's AI progress, he argued, did not come from two labs but from a whole-of-industry effort across universities, startups and organisations, quoting the line that no single string can weave a cloth by itself — noting that Epoch's own visual theme was weaving. Building the frontier, he said, is something India has to do together.


Part II — Models

The models section ran the longest and covered five distinct releases plus a set of forward commitments.

2.1 — Sarvam 105B: Two Variants, and a Router That Cuts Cost by 40%

Sarvam released the 105B checkpoint in February and spent the months since acting on production feedback. The upgraded model now ships in two forms: an instruct model tuned for low-latency voice, and a reasoning model for harder tasks.

For conversational AI, the team built its own benchmark rather than relying on public ones — deliberately constructed from production data to capture what actually breaks in deployed voice agents: ambiguous user messages, noise introduced by speech recognition, and strict instruction-following such as being required to say a specific phrase verbatim on the tenth turn. Against that benchmark, Sarvam positioned 105B ahead of GPT-5.4 mini and Gemini 3.5 Flash, at roughly 5.5x and 11x lower cost respectively.

Two illustrative behaviours were shown: the model answering a policy question succinctly in the user's own code-mixed language while competitors dumped all fifteen retrieved documents, and — in a banking scenario where the system prompt forbade assumptions — the model asking a clarifying question instead of executing a balance-clearing tool call on an ambiguous request. As Dr. Kumar put it afterwards, the GPT model in the comparison is not available in India at all, and the Gemini model is not always available at scale.

The more interesting result came from work agents. Inside Sarvam's own harness, the 105B performs close to GLM 5.2 on tasks like security and tool use — despite GLM 5.2 being roughly seven times larger with far more active parameters. Sarvam then built a router on top: since 105B already handles a large share of everyday requests, the router decides which queries need the bigger model. In internal testing, 36% of day-to-day queries were solvable by 105B alone, cutting harness costs by about 40%.

2.2 — Sarvam Vision 2.0: A New Pareto Region in Document OCR

Sarvam Vision 1.0, launched in February, was state-of-the-art on the global MOCR benchmark and has since digitised more than 35 million pages. Version 2.0 adds native key-value extraction — which the team believes is a first among comparable models, ahead of systems like DeepSeek OCR and Paddle — along with Indic handwritten recognition, better complex-table parsing, and improved general OCR.

The team framed the result in terms of the Pareto frontier across two benchmarks: OmniDocBench, which tests structural extraction under pressure, and a second benchmark testing OCR in the wild on dense, degraded documents. Traditional OCR engines from AWS, Google Cloud Vision and Azure sit far in the lower-left. Existing state-of-the-art models are each heavily optimised for one axis at the expense of the other. Vision 2.0, Sarvam claimed, opens an entirely new Pareto region by refusing that trade-off — tables can now be extracted as structured JSON key-value pairs rather than flat markdown, and forms filled in Hindi or other Indian languages can be read with high accuracy across all 22 languages.

Two things were announced alongside it. Sarvam Vision Edge runs the model locally on a user's own machine. And the team showed a live public-sector project: digitising the Odisha government's land records, some of the hardest documents imaginable — handwritten, poorly maintained, decades old. Accuracy has climbed from roughly 13–14% to about 50% and is still improving.

Dr. Kumar highlighted the shape of the partnership as much as the result: a state government bringing a hard problem and its data to a research team, and hill-climbing on it together. He also noted that key-value extraction only became a priority because working with an insurance company revealed that customers don't actually want OCR — they want the name and the address.

2.3 — Saaras V4: Speech Recognition, Including the World's First Real Multi-Speaker Indic Model

Speech recognition, the team said, has been the soul of Sarvam since day one. Saaras V4 extends the company's Indic dominance and widens the gap on commonly used languages, while supporting all 22 constitutional languages of India — including ultra-low-resource ones where training data is scarce.

The headline claim, though, was about English. For the first time, Saaras V4 is state-of-the-art on global English across seven standard benchmarks, six of which are foreign and global English rather than Indian English. The team called this "the Sarvam moment of build in India, for the world." The model also ships as one model with five transcription modes — including romanised output, code-mixed and translation — and shows large gains on the noisiest benchmark sets.

The genuinely novel release was Saaras V4 Multi-Speaker. Conventional pipelines run diarisation first, slicing audio into single-speaker chunks, then transcribe each chunk — an approach that collapses when people talk over each other. Sarvam rebuilt the architecture so a single model with a speaker-biased encoder emits speaker-attributed transcripts while preserving overlaps and back-channels. The stage demo used Indian television news debates, where multiple anchors and panellists talk simultaneously; the model correctly separated overlapping male and female speakers. As one audience member noted in the livestream chat, Indian TV debates may be the hardest speech-recognition challenge available anywhere.

To make the claim falsifiable, Sarvam open-sourced the first multi-speaker conversational Indic benchmark on Hugging Face: 22 languages, over 1,200 speakers, 108 hours of genuinely hard real-world audio from news broadcasts and dinner-table conversation. On that benchmark, and on the standard English AMI benchmark, Sarvam reported leading results on both CPWER and DER. Both models are live in the API and playground.

2.4 — Bulbul V4: Expressive Speech, and a Voice Program for India

Bulbul V4 was introduced with a Navarasa performance — the nine classical emotions of Indian aesthetics, each demonstrated by the model. The technical leap is steerability. A creator can add emotion tags, insert paralinguistics such as laughter, and emphasise specific words. Most notably, the model can be re-steered mid-generation: the demo took a line about winning the lottery from flat, to excited, to excited-then-turning-wistful at the mention of a late father — all within a single utterance.

Alongside it came a new voice library spanning use cases from edtech to voice agents to cricket commentary, and voice cloning. The team then made the cultural argument explicit, playing a cloned voice delivering a warning that if the digital express of AI leaves the station with a dozen Indian languages stranded on the platform, the country will have surrendered its intellectual sovereignty by speaking exclusively in borrowed tongues.

Out of that came the announcement of a Voice Program: brands can create, onboard and clone a brand voice; creators can reach wider audiences; and voice artists can license their voices to the enterprises and developers building on Sarvam's platform.

2.5 — Custom Model Training: Frontier-Scale Training From a Laptop

The team's view is that while a token factory lets you use models that already exist, the world is missing infrastructure to train your own. Getting from procuring GPUs to a first optimised training run takes weeks or months of specialist work across storage, networking and job orchestration, followed by tuning hundreds of framework flags that change from model to model.

Sarvam's answer is a single Python SDK you run from your laptop and submit to Sarvam's infrastructure: data set in, checkpoint out, in roughly 30 lines of code. It launches with reinforcement learning support built on two deliberate engineering choices — asynchronous RL, so the generation worker and training worker overlap instead of idling while the other works, and LoRA, which shrinks trainable parameters enough that weight synchronisation between nodes doesn't saturate the network.

THE COMMERCIAL POINT THAT MATTERS MOST

"It is your IP. The model that you get is your model. It is not Sarvam's model."

Dr. Kumar drew an explicit comparison to Tinker from Thinking Machines Lab. One unnamed customer brought its own data, trained on Sarvam's infrastructure, and ended up with an optimised model six times cheaper to run — which Sarvam can then serve from its own token factory. The value loop stays inside the company, and inside the country.

2.6 — The Frontier Commitment, and the Arithmetic Behind It

Sarvam announced it is investing in trillion-plus parameter models built from scratch in India, targeting coding, cybersecurity, simulation and science. Dr. Kumar framed this as one of several simultaneous scale-ups.

LeverFromTo
DataTens of trillions of tokensHundreds of trillions of tokens
Model size105B parametersTrillion-plus parameters, trained from scratch in India
Compute4,000 Hopper GPUs (used for 105B)Largest Blackwell cluster in India, scaling to ~10,000 accelerators, run bare-metal in-house
Reinforcement learningPost-training on specific tasksSustained RL and continual learning at much larger scale

Table 2 — The four levers Sarvam said it is cranking simultaneously.

To make the trillion-parameter goal concrete, Sarvam brought on stage Shri Devendra Singh Chaplot, announcing both a new San Francisco office and Chaplot's association as advisor. He was most recently pre-training lead at xAI, and before that was part of the founding teams at Mistral AI — where he worked on Mistral 7B, Mixtral 8x7B and Mistral Large — and at Thinking Machines Lab, where he led Tinker Enterprise. He is an IIT Bombay alumnus with over a decade in AI. (Source: Inc42; Business Standard; The Hans India, 30 July 2026.)

Asked directly what it takes to build at the frontier, Shri Chaplot pushed back on the idea that it is out of reach for India. Take a K2/K3-class model, he said — roughly 3 trillion total parameters with about 100 billion active — and train it on 100 trillion tokens. That works out to around 6 × 10²⁵ FLOPs of pre-training, which at good utilisation on Blackwell GPUs means pre-training in about two months on the 10,000 chips Sarvam had just announced. Reinforcement learning, research and post-training extend that, but the full cycle lands at roughly four to six months. "It's very possible in terms of compute to do this in India," he said. "This is not exorbitantly expensive."

On talent he was equally direct, calling it another misconception that frontier work requires a large number of highly experienced people. You need a handful of experienced people for guidance, he argued, and then a lot of motivated and talented ones — something India has in abundance. The technology is young enough that even his own experience spans only a few years, and capable people can catch up very quickly with the right guidance.

Asked where the world is heading, he reframed the question: would you rather live in a world with three models, or thousands? Competition breeds innovation, specialisation beats one-size-fits-all, and India in particular — with its languages, cultures, education system, legal system and public policy — needs models that can be customised for all of it.

2.7 — The Research Agenda

Beyond shipped models, Dr. Kumar named five research directions Sarvam is publicly committing to:

  1. Dense models on the edge — as memory prices rise, mixture-of-experts architectures work well in the cloud but dense models are what actually fit on local machines.
  2. Live interactive systems — sub-second, high-fidelity multimodal interaction, informed by running the country's largest voice AI platform.
  3. Cyber systems — offensive and defensive security capability, an area Dr. Kumar noted every company in the room had worried about in the preceding month.
  4. Simulation and digital twins for real systems — work being done with the armed forces on simulating and then operating large real-world systems.
  5. Foundation models for customer data — India's strength in digital public goods and large financial institutions means enormous volumes of customer data with no good model paradigm yet.

Dr. Kumar also trailed a first-of-its-kind research partnership with what he called a very consequential Indian company, to be announced on the Enterprise Edition day — an Indian company choosing to absorb the capability to build custom models on its own data.


Part III — Infrastructure: The Token Factory

Cofounder Dr. Vivek Raghavan opened the infrastructure section with the argument underneath it. Technological revolutions, Dr. Kumar had noted earlier, come from being able to convert molecules from one form to another — textiles, steel, semiconductors. The equivalent unit today is the token, and most of the tokens India consumes are manufactured outside the country.

Importing tokens, Dr. Kumar argued, is worse than importing steel: it is a double whammy. You give away control of manufacturing, and you give away your data at the same time. He noted this is not a fringe view, citing public remarks from Microsoft's Shri Satya Nadella and Palantir's CEO. Hence the stated goal: Sarvam wants to produce the largest share of the tokens India consumes, right here in India — and then make that available to the world.

3.1 — Sarvam Inference

Sarvam Inference is a platform serving state-of-the-art open-weights models from Indian infrastructure at globally competitive prices. Alongside Sarvam's own models — Saaras, Bulbul, Vision and the dubbing models — three models were highlighted for serving at scale:

ModelScaleServing commitment
GLM 5.2768B parameters — among the most capable open models in the worldServed from India at 80 tokens/second, globally competitive price
Gemma 4 31BSmaller, high-throughput modelTime to first token under 6 seconds, globally competitive price
Sarvam 105BSarvam's own flagship, instruct + reasoning variantsBest conversational model hosted in India; ~9x cheaper than alternatives for instruction-following, tool-calling and voice

Table 3 — The three models anchoring Sarvam Inference at launch.

Dr. Raghavan was candid that taking an open model from Hugging Face to reliable production serving is where the real work lies. Sarvam built an agentic model-optimisation loop that does this automatically: bespoke kernels, framework and scheduler work, and systems tuning. In certain cases the system achieved up to a 15x inference speed-up in a single day — an AI system profiling the CUDA graph, finding the gaps, and optimising them overnight.

Dr. Kumar's aside on this was one of the more striking moments of the day: all the talk about recursive AI building AI, he said, has marketing to it, but as builders "we feel the joy" — there is a model that looks at the simulator, sees where the gaps are, and optimises the CUDA graph.

Building a competitive token factory, Dr. Raghavan noted, means solving a stack of hard problems at once: routing thousands of GPUs across heterogeneous workloads, disaggregating parts of the compute, handling KV cache extremely well, managing the memory hierarchy both on-chip and across disk and shared filesystems, and doing kernel-level optimisation. And like the model frontier, it moves — models change, GPUs change, methods change.

The commercial offer was blunt: 30% off for anyone committing volumes in August. "You shouldn't have to pay a tax to use models here in India," Dr. Kumar said.


Part IV — Applications: The Indus Platform

The third movement marked what Dr. Kumar called an inflection: Sarvam moving from research and bespoke enterprise engagements to generally available products. The vehicle is Indus, described as the agentic AI platform for India — not one product but six, under a single platform with a unified credit system.

ProductWhat it is
Voice Agents (Samvaad)Build, deploy, monitor and evaluate voice, WhatsApp and web agents — now open to everyone
Work AgentsEnterprise agent workspace with sandbox, connectors, skills and Slack deployment — live and free
Content AgentsVoice library, cloning, translation and motion-picture-grade dubbing
Doc AgentsDigitisation, extraction and translation built on Vision 2.0, available as APIs
Coding Agents (Sarvam Code)Three-agent coding harness with verticalised modes — in early beta
InferenceThe API and serving stack described in Part III

Table 4 — The six products launched under Indus. Sarvam's earlier developer surface, launched four months prior, was described as being 10x supercharged by this release.

4.1 — Voice Agents: 325 Million Minutes of Evidence

Sarvam asked the audience to guess how many minutes its AI agents had spent talking to real customers over twelve months. The answer: 325 million minutes — the equivalent, as the team put it, of 618 years of conversation. Every one of those calls taught the team something about Indian voice AI: customers who start in Kannada and switch to Marathi, interruptions, background noise, latency, voice activity detection. Sarvam claims to be by far the largest voice AI platform in the country.

What was previously available only to enterprises is now open to everyone, at the same models, same technology and same price. Three friction points were attacked directly.

Authoring: three weeks to under 60 minutes. Instead of a scoping document handed from a business owner to a developer, you describe the agent you want. The platform interviews you — when should customers be called, which languages — then plans the conversation design, greeting, tools and dialects itself. A simulate-and-test feature then runs roughly a thousand simulated users against the agent before it ever reaches a real customer.

Telephony: six weeks to 30 seconds. Deploying a phone agent normally means a telecom provider, KYC, minimum commitments, contracts and SIP trunking — six weeks of work that often ends with the discovery that trunking only works one way. Sarvam is, it says, the only platform where you can rent an Indian mobile number in the 080 and 91 series directly inside the platform in 30 seconds, by uploading PAN, Aadhaar and GST.

Workflows and self-improving agents. One agent handles a call; a fleet delivers an outcome. Workflows lets agents hand off to each other with shared memory — the demo showed an abandoned-cart sequence where an explainer agent identifies the objection and the callback time, schedules the follow-up with no human involvement, and a differently-prompted negotiator agent closes the sale two days later.

The more consequential feature is self-improvement. You define a goal — a qualified lead counts as success — and the platform then reviews conversations automatically. In the example shown, the agent worked through 4,812 conversations, identified systematic failures such as mishandling budget objections and callback timing, then ran its own prompt experiments across a control and three variants over a range of days to find the winner, promoting it back with human-in-the-loop approval. Deployed with a large fintech in production, agent conversion rate improved roughly 3.5x between day one and day 28.

The section ended with a quiz: an audio clip of a customer order, with the audience asked which speaker was the AI. The answer was both.

Voice agent pricingRate per minuteComparison
Sarvam (Samvaad)₹3.5Lowest in the industry, all features bundled
Typical market players₹8 – ₹122.3x to 3.4x more expensive
An average human agent₹4 – ₹6Sarvam is cheaper than a human

Table 5 — Voice agent pricing as presented on stage. The Sarvam rate includes AI-based authoring, simulation, phone numbers, bots, evaluations and analytics.

4.2 — Work Agents: Beating Hermes on Harness Bench

Sarvam's framing for work agents is that organisations will increasingly be teams of humans and agents together. It offered itself as the example: its twelve-person marketing team is three people consuming food and nine agents consuming tokens, and roughly 25–30% of Sarvam's entire workforce is now agents contributing to code, data analytics and product decisions.

The flagship demo gave the agent ten Excel files — trial balance, sales and purchase registers, bank statements, cash flow summaries, around three lakh rows of messy unstructured data — and a single request to prepare an investor presentation. The agent wrote and executed Python in a secure cloud sandbox, planned the work into six or seven steps, paused to ask the human for direction on framing and format, and produced a polished interactive artifact covering top customers, monthly sales and product lines. One task, single shot, twelve minutes.

Three design decisions were emphasised. The sandbox is custom-built and agents never receive API credentials or secrets, with a gateway enforcing execution security. Agents can be promoted to a shared workspace so multiple people collaborate with them, rather than a single-player model. And everything — inference, model, harness, product and data — runs on Sarvam inside India; nothing leaves.

Sarvam also built a knowledge layer that ingests multimodal, multilingual, multi-format enterprise data and exposes it to the agent as a navigable file system rather than a vectorised retrieval index — allowing the agent to read, drill down and progressively disclose information, and to ground answers by citing where in the knowledge base they came from. Dr. Kumar's metaphor: the model is the CPU, the harness is the operating system, and this is the pen drive you plug in.

Work Agents is live today and free for everyone. Custom agents can be created in under ten minutes — the demo built a bug-RCA agent that chose its own connectors (GitHub, Sentry, SigNoz, web search), wrote its own diagnostic skill and drafted its own instructions — and deployed to Slack in one click, so that agents come to where employees already work.

4.3 — Sarvam Code: Planner, Worker, Verifier

Sarvam Code entered early beta with a clearly stated problem: coding velocity is not translating into product velocity. Codebases are filling with slop, nobody on the team knows what is actually going to production, and coding agents are becoming expensive. Teams told Sarvam they wanted to move to open-source models inside coding harnesses — and found it didn't work, because benchmark performance on Terminal-Bench or SWE-Bench didn't survive contact with real engineering.

The core fix concerns long-horizon work. Goal mode as it exists in Codex, the team argued, really only works with the largest and costliest closed model; with open models like GLM or Kimi it fails because the model sets its own goal, does the work, and then declares itself finished — open models are not reinforcement-trained enough to reliably judge completion. Sarvam decomposed goal mode into three cooperating agents: a planner, a worker and a verifier. On Terminal-Bench's roughly 89 tasks, Sarvam Code solves a comparable number of tasks at a small fraction of the cost of frontier models in their closed harnesses. Notably, Sarvam Code was itself written to a large extent by Sarvam Code.

On top of the base agent sit three opinionated verticalised modes:

  1. Android application building — the demo built an AI-native motor-insurance app end-to-end, using Sarvam's own vision and voice APIs to bypass the handwritten forms and image uploads that make Indian government processes cumbersome. The agent drove an Android emulator to verify its own work, wrote its own test cases, and iterated until shippable.
  2. Data science and ML — built because banks report investing hundreds and thousands of crores in data infrastructure without ROI. On DataAgentBench, created by UC Berkeley and Hasura, the team hill-climbed from around 60 to 82.
  3. Cyber — chaining multiple vulnerabilities to achieve system takeover, and then patching them.

Two results stood out. On Kaggle, the team entered the Home Credit underwriting competition — $100,000 prize pool, 1.5 million rows, 500 features across seven tables, 30,000 entrants, many of whom had submitted hundreds of times over six months. Sarvam's agent reached the top of the leaderboard in three hours, running its own exploratory analysis, generating hypotheses and training models until it hit a compute, time or accuracy budget.

On security, the cyber agent completed all 41 tasks on ExploitBench, achieving full system takeover on 26. More seriously, the team disclosed that it has found genuine vulnerabilities in applications and software used by tens of millions of people in India, and is working responsibly with those vendors to patch them. The live demo showed the agent black-box pentesting a Grafana deployment, chaining a SQL injection into command execution — and then patching the system it had just broken.

4.4 — Content Studio and Doc Agents

Content Studio bundles a 50-plus voice library, text-to-speech, voice cloning and — the piece the team was most excited about — motion-picture-grade dubbing with correct lip sync, preserved emotion and expressive delivery. A cricket World Cup commentary demo showed the range. The most self-referential proof point: the keynote itself was being dubbed live on YouTube into Gujarati and other languages as it happened.

Doc Agents exposes the Vision 2.0 capabilities through a playground and as configurable APIs, with a large fintech already using them to automate handwritten bank forms at scale.


Part V — Beyond the Cloud: Sovereign, Assistive and Ambient

5.1 — Chanakya: On-Premise AI for Defence and National Security

Sarvam's Chanakya vertical builds for government entities across defence, national security and intelligence — work it has been doing with the Indian armed forces for about eighteen months. The constraint is absolute: in these environments any exposure to the internet is a potential leak, so the entire stack — inference, models, harness, application and data — must run on the customer's own premises and on whatever compute they have, which is usually limited.

The team demonstrated Anveshak, a platform that builds a unified data layer over messy data of any format — video, audio, PDFs, images — and supports multi-hop deep research over it. It buckets similar files into logical structures, derives an ontology and entity schema, and extracts a knowledge graph using efficient sampling rather than reading an entire corpus. An analyst can then interrogate the graph, with the system spawning what the team called a paranoid agent to ask deeper follow-up questions, and every answer traceable to its source documents.

The case study was striking. SEBI approached Sarvam after the Rajesh Exports matter to ask whether AI could have detected the problem in advance. Using only open-source data available on the internet before the incident surfaced, Anveshak produced an analyst-grade report flagging financial fragility and suspicious patterns.

Context for readers: SEBI passed an interim order in June 2026 against Rajesh Exports Ltd and its promoter, alleging revenue misrepresentation across FY21–FY25, with the NFRA opening its own probe in July 2026. The matter is at the investigation stage; nothing in this note should be read as a finding of guilt. Source: Business Today, June–July 2026; Law.asia, 2026.

The entire demonstration — model, harness, application and data — ran on a single NVIDIA Spark device sitting on the podium, using a 30-billion-parameter model. The team's point was that giving a model a bounded, well-structured field to search lets a much smaller model match a much larger one. Representatives of all three armed forces were due to attend the following day.

5.2 — Kaze: AI Smart Glasses Built With the National Association for the Blind

Sarvam showed Kaze, AI smart glasses developed in collaboration with the National Association for the Blind, and brought Ms. Yogita, an NAB outreach executive who has used them for seven months, on stage with twenty of her colleagues in the audience. (Product name per NewsBytes, July 2026.) The demo covered weather, choosing an outfit, identifying which bus is approaching by number, reading a café menu, and coaching her on how to order in Kannada.

Her own account was the most direct evidence of the day: identifying bus numbers had been a persistent daily difficulty, and the glasses solve it. She noted the device supports 22 Indian languages including her mother tongue, and can describe colours.

Technically, the glasses run models optimised for the hardware that handle Indian languages accurately in background noise, and use a cloud-plus-edge split so they still work in low-network areas without escalating cost. They are designed and made in India. Sarvam framed three impact areas — social impact, enterprise deployment, and sovereign responsibilities — and is targeting enterprise use cases in healthcare, banking and tourism; a second demo showed a car dealership sales assistant analysing a customer conversation in real time and recommending a vehicle. What makes this rare, the team argued, is that speech, vision, language and edge deployment normally come from disjointed vendors; here they come from one stack.

5.3 — Kivi: Rethinking How We Talk to Machines

The day closed with Kivi, which began as a part-time passion project by four engineers. The framing was historical: the typewriter keyboard and the telephone were both invented in the 1860s. The telephone went from point-to-point, to rotary, to cellular, to satellite, to video anywhere on earth. Typing has barely changed — because the keyboard is predictable, reliable and free, and nothing else has matched all three. Kivi's argument is that with a state-of-the-art ASR model plus intelligence on top, all three are now solvable.

Beyond dictation, Kivi offers custom dictionaries, speed-dial shortcuts, per-app output styles, and the ability to select text anywhere on screen and command an edit by voice — "make it sound apologetic." A demo showed a form filled entirely by speech. A forward-looking vision demo showed voice-driven control across applications: summarising and translating a message thread, replying and sending, drafting an email to a specific contact and booking a calendar slot at their first free time, and booking a flight conditionally. Kivi is downloadable for Mac today, with Windows and Android to follow, and API and MCP access for developers. It is free, and the Pro tier was made free for everyone until 15 August.


The Benchmark Scorecard

AreaBenchmarkResult claimed
Conversational AISarvam's own production-derived benchmark105B ahead of GPT-5.4 mini and Gemini 3.5 Flash, at 5.5x and 11x lower cost
Work agentsHarness Bench (tool use, long-running tasks, latency)Sarvam Work 79% vs Hermes 76%
Model routingInternal day-to-day query mix36% of queries handled by 105B; ~40% harness cost reduction
Document AIOmniDocBench + OCR-in-the-wildVision 2.0 opens a new Pareto region; native key-value extraction
Speech recognition7 standard English benchmarks (6 non-Indian)Saaras V4 state-of-the-art on global English
Multi-speaker ASRIndic Diarisation Bench (newly open-sourced) + AMILeading on both CPWER and DER
CodingTerminal-Bench 2.1 (~89 tasks)Comparable tasks solved at ~$2/task vs $4.10–$27.80
Data scienceDataAgentBench (UC Berkeley + Hasura)Hill-climbed 60 → 82
Applied MLKaggle Home Credit ($100k, 30,000 entrants)Top of leaderboard in 3 hours
CyberExploitBench41/41 tasks completed; 26 full system takeovers
InferenceInternal agentic optimisation loopUp to 15x speed-up achieved overnight

Table 6 — Every quantitative claim made from the Epoch stage, collected in one place. All figures are as presented by Sarvam and have not been independently verified.


What Is New vs. What Is Updated

TrackGenuinely newUpdated or extended
Language models7B and 70B multilingual models; trillion-parameter model in development; model routerSarvam 105B — two variants (instruct + reasoning), voice support, ~$0.80/M tokens
SpeechSaaras V4 Multi-Speaker; open-sourced Indic multi-speaker benchmark; Bulbul V4 mid-generation re-steering; Voice ProgramSaaras V4 (now SOTA on global English); Bulbul voice library expanded
VisionNative key-value extraction; Indic handwriting; Sarvam Vision EdgeVision 2.0 — new Pareto region; 35M+ pages already digitised with 1.0
InfrastructureSarvam Inference; custom model training SDK; largest Blackwell cluster in India; 30% August volume discountGPU scale-up toward 10,000 accelerators; agentic inference optimisation
ApplicationsIndus platform (6 products); Sarvam Code; Workflows; self-improving agents; Slack deployment; knowledge layerVoice agents opened from enterprise-only to everyone; developer platform 10x supercharged
Frontier & orgSan Francisco research office; Shri Devendra Singh Chaplot as advisor; five-point research agendaThree-year model-building journey from Llama 3 contributions to 105B
Sovereign & edgeAnveshak on-prem platform; single-Spark deployment; Kaze AI smart glasses with NAB; Kivi18 months of armed forces work productised into a deployable on-prem stack

Table 7 — Separating first-time launches from iterations on existing Sarvam products.


Impact Analysis

For Developers

The practical headline is cost and access. A developer building for India can now run frontier open models — GLM 5.2 at 768B parameters, Gemma 4 — from Indian infrastructure at globally competitive prices, with a 30% discount for August volume commitments, rather than routing to overseas clouds. Sarvam's own 105B is positioned at 5.5x to 11x cheaper than the comparable global conversational models, one of which is not available in India at all. Work Agents is free, Kivi Pro is free until mid-August, and voice agents are open to everyone with no waitlist.

The deeper shift is that several capabilities that previously required a team now require an afternoon. A voice agent that took three weeks to author and six weeks to connect to telephony can now go live in an hour with a phone number rented in 30 seconds. A custom fine-tuned model that required GPU procurement and months of infrastructure engineering is now roughly 30 lines of Python submitted from a laptop — and the resulting model belongs to the developer, not to Sarvam. For anyone building Indic products, the open-sourced multi-speaker benchmark and the 22-language coverage remove a data problem that was previously unsolvable at small scale.

For Industry

For enterprises the proposition is consolidation. Sarvam is explicitly not positioning itself against OpenAI so much as against the sprawl — the argument being that an enterprise can now buy infrastructure, models, fine-tuning, harnesses and applications from one domestic vendor, with role-based access control at the MCP layer, agent layer and chat layer, and with inference and data never leaving India.

That matters most in regulated sectors, and the evidence shown was sector-specific rather than generic: a large fintech seeing 3.5x conversion improvement from self-improving voice agents; banks with expensive data infrastructure that isn't returning ROI being targeted by the data science coding mode; insurance workflows that drove the key-value extraction feature; a state government digitising land records; SEBI exploring market surveillance; the armed forces running the whole stack on-premise on constrained hardware. The custom-training offer sharpens this further — a company can bring proprietary data, get a model six times cheaper to run, own the IP, and have Sarvam serve it. For IT services companies currently buying coding licences from foreign vendors, Sarvam Code is a direct cost argument.

The honest caveat is that much of this is early. Sarvam Code is in beta, Epoch Builder Edition enters private preview in August with general availability targeted for Q4, and the trillion-parameter model has an arithmetic case but no shipped result. Monetisation at this scale remains unproven.

For Common Users

Most people will never touch an inference API, and the announcements that matter for them are the ones furthest from the infrastructure. The smart glasses built with the National Association for the Blind are the clearest example: a person who could not previously identify an approaching bus can now do so, in her own language, on a device made in India. That is not a demo, it is seven months of daily use.

More broadly, the speech work is what reaches ordinary users. Support lines and government services that understand code-mixed Hinglish and Kannada-English rather than forcing English; voices that carry emotion instead of sounding robotic; a model that can follow a conversation where two people talk over each other; land records and bank forms that get digitised instead of remaining inaccessible paper. And because voice agents now cost less per minute than a human agent, the economics push organisations toward serving people in their own language rather than rationing that as a premium tier. Content dubbing at motion-picture quality means film, education and news can cross language boundaries — the keynote itself was being dubbed live into Gujarati as it was delivered.


Closing Assessment

Epoch 2026 was unusual for the density of what was shipped rather than promised. Saaras V4, Bulbul V4, Vision 2.0, the upgraded 105B, Work Agents, voice agents and Kivi were all live or available the same day; Sarvam Code was in beta; the custom training SDK and Sarvam Inference were open for business. The genuinely speculative items — the trillion-parameter model, the scale-up to 10,000 accelerators — were presented with arithmetic rather than adjectives, which is a more useful way to be held accountable.

The strategic bet is coherent. If the competitive advantage in AI is shifting from having the best model to owning the whole stack, then a company that trains models, runs its own bare-metal GPU clusters, builds its own harnesses, and ships its own applications has a structural position that a model-only lab or an application-only startup does not. The 40% cost reduction from routing between its own model and a larger one, and the 15x inference speed-up from an agent optimising CUDA graphs overnight, are both examples of an advantage that only exists if you control several layers at once.

The open questions are execution and time. Sarvam is now competing simultaneously with frontier labs on models, with hyperscalers on inference, with Anthropic and OpenAI on coding agents, with Microsoft on work agents, and with hardware companies on glasses. Dr. Kumar's own answer was to ask for the timescale to be measured in decades rather than quarters, and to frame the frontier as a whole-of-industry effort rather than one company's race.

Whether the rest of Indian AI takes up that invitation may matter as much as anything Sarvam ships next.


Appendix — Jump to the Video

Every section of this report is mapped to the point in the Sarvam Epoch 2026 livestream where it was announced. Click any timestamp to open the video at that moment.

Report sectionVideo timestampWhat is announced there
Opening keynote — Dr. Pratyush Kumar takes the stage32:39Welcome; what 'Epoch' means
Part I — The Framing: Why Build Models At All?32:55 · 38:33Three commitments; the case for training our own models
The Sarvam Stack (Figure 1)42:49Infra → models → harnesses → data → users
Part II — Models44:19Preview of every model release
2.1 — Sarvam 105B: two variants and the router58:33 · 1:03:50Conversational benchmark; 36% routed, 40% cost cut
2.2 — Sarvam Vision 2.01:07:29 · 1:12:24New Pareto region; Vision Edge; Odisha land records
2.3 — Saaras V4 + Multi-Speaker1:14:38 · 1:17:35SOTA global English; overlapping-speaker ASR
2.4 — Bulbul V41:24:39 · 1:30:11Navarasa demo; emotion steering; Voice Program
2.5 — Custom Model Training1:32:28 · 1:35:30Python SDK from your laptop; your model is your IP
2.6 — The Frontier Commitment45:55 · 1:37:07 · 1:38:32Scale-up levers; SF office; Shri Chaplot's arithmetic
2.7 — The Research Agenda47:55Edge, live systems, cyber, digital twins, customer data
Part III — Infrastructure: The Token Factory50:25 · 51:26Token sovereignty; the double whammy of importing tokens
3.1 — Sarvam Inference1:43:46 · 1:46:34 · 1:48:42GLM 5.2, Gemma 4, 105B; 15x optimisation; August discount
Part IV — Applications: The Indus Platform1:50:41 · 1:52:42Six agentic products under one platform
4.1 — Voice Agents (Samvaad)1:53:31 · 1:53:49325 million minutes; opened to everyone
— Authoring in under 60 minutes1:55:35Describe-your-agent flow; 1,000 simulated users
— Telephony in 30 seconds1:57:37Rent an Indian number with PAN/Aadhaar/GST
— Workflows & self-improving agents1:59:20 · 2:01:52Agent fleets; 4,812 conversations; 3.5x conversion
— Pricing2:06:10₹3.5/min vs ₹8–12 market and ₹4–6 human
4.2 — Work Agents2:08:41 · 2:17:07Investor-deck demo; Harness Bench 79% vs 76%; Slack
4.4 — Content Studio & Doc Agents2:28:26 · 2:31:36Voice cloning; motion-picture dubbing; doc APIs
4.3 — Sarvam Code2:32:55 · 2:37:44Planner/worker/verifier; Android, ML and cyber modes
Part V — Beyond the Cloud2:47:10On-prem, assistive and ambient computing
5.1 — Chanakya: on-prem for defence2:48:17 · 2:51:05Anveshak on a single Spark box; the SEBI case study
5.2 — Kaze AI Smart Glasses with NAB2:56:06 · 2:59:42Demo film; Ms. Yogita's account of seven months' use
5.3 — Kivi3:06:23 · 3:17:18Voice interaction; free Pro until 15 August
Closing remarks3:17:59Wrap-up of the first Epoch keynote

Table 8 — Section-by-section index to the 3 hour 21 minute Builder Edition keynote. The event proper begins at 32:39; everything before that is the pre-show.


References

Primary Source

  • Sarvam Epoch 2026, Builder Edition — full keynote livestream (approximately 3 hours 21 minutes), 30 July 2026, Bengaluru. youtube.com/live/peO2ReobYSwthis is the primary source for the entire note.
  • Sarvam Epoch official event site. epoch.sarvam.ai
  • Sarvam AI model documentation (Saaras, Saarika, Bulbul). docs.sarvam.ai

Same-Day and Follow-Up Coverage

  • Sarvam To Build Trillion-Parameter AI Model, Launches India-Hosted Inference Service. Inc42, July 2026. inc42.com
  • Sarvam Takes On Claude, Codex With Cheaper, India-Hosted Coding Agent. Inc42, July 2026. inc42.com
  • Sarvam Ropes In Mistral AI Founding Team Member Devendra Chaplot As Advisor. Inc42, 30 July 2026. inc42.com
  • Sarvam ropes in Mistral founding team member Devendra Chaplot as adviser. Business Standard, 30 July 2026. business-standard.com
  • Top AI expert Devendra Chaplot, formerly with Elon Musk, joins Sarvam from xAI. The Hans India, July 2026. thehansindia.com
  • Sarvam AI launches platform to help build India-centric AI models. ANI, 30 July 2026. aninews.in
  • Sarvam Epoch 2026: Saras V4, Bulbul-V4, Kaze smart glasses showcased. NewsBytes, July 2026. newsbytesapp.com (source for the Kaze product name)
  • This week in AI: India's sovereign AI ambitions get bigger. Forbes India, August 2026. forbesindia.com
  • Sarvam AI Launches Epoch Builder Edition. GKToday, July 2026. gktoday.in

Context Sources

  • How a shareholder complaint triggered Sebi's massive probe into Rajesh Exports. Business Today, June 2026. businesstoday.in
  • Rajesh Exports case: NFRA begins probe after Sebi flags alleged ₹15.15 lakh crore mismatch. Business Today, July 2026. businesstoday.in
  • SEBI issues interim order after probe into ₹1tn fraud. Law.asia, 2026. law.asia

RESEARCH METHODOLOGY & AI ASSISTANCE

This AI Note is built primarily from the full keynote transcript of the Sarvam Epoch 2026 Builder Edition livestream, supplemented and fact-checked against same-day and follow-up coverage from Inc42, Business Standard, ANI, The Hans India, Forbes India and Business Today. AI-assisted tools, including Claude, supported transcript synthesis, drafting, verification and editorial refinement. The structure, analytical framing and interpretations reflect the author's independent judgement.

On figures: all performance numbers, pricing and benchmark results in this note are as claimed by Sarvam from the stage and have not been independently verified. They are reported here because they are the company's own public, falsifiable commitments — not because they have been confirmed. The Benchmark Scorecard exists to make them easy to check as evidence emerges.

On names: product and personnel names were transcribed from audio and then verified against published post-event coverage. Where the transcript and published coverage differed, published coverage was followed — the speech model is Saaras V4 (not Saarika, an earlier Sarvam ASR model), the desktop voice tool is Kivi, and the smart glasses are Kaze. Any remaining spelling variance from Sarvam's official branding is unintentional.


About the Author

Bhaveshkumar Choithram Dharmani is the Founder of AIVidhya4Sarvam and works as an AI mentor, researcher, and ecosystem builder. His focus is on AI education, mentorship, and building the conditions for meaningful AI participation across India — in institutions, organisations, and communities that are not yet well-served by the current AI education ecosystem.

AIVidhya4Sarvam (aividhya.in) is an AI mentorship, innovation, and transformation organisation. It works with students, professionals, startups, and institutions to build AI capability with rigour and purpose.

Issue #03 of the India AI Journey series by AIVidhya4Sarvam.

India AISarvam AISarvam EpochFrontier ModelsAI InfrastructureAgentic AIIndic LanguagesToken SovereigntyIndia AI Journey