Researchers at the Indian Institute of Science (IISc) Bengaluru, in collaboration with the AI and Robotics Technology Park (ARTPARK) and Google, have released SraVaani, an open-source multilingual speech recognition model that can convert spoken words into text across 65 Indian languages and dialects. Released on 13 August 2026 under the permissive MIT licence on the Hugging Face platform, SraVaani covers 20 scheduled languages and 45 regional languages and dialects, many of which have no existing speech recognition support. The model aims to bring voice AI to the roughly 25 crore people whose languages are not properly handled by current speech technology systems.
What Is SraVaani?
SraVaani is a multilingual automatic speech recognition (ASR) model that converts spoken language into written text. The name draws from the Sanskrit words “Sra” (to hear) and “Vaani” (speech or voice), reflecting the project’s core mission of making voice technology accessible across India’s diverse linguistic landscape.
The model was developed by the Signal Processing, Interpretation and Representation (SPIRE) Lab at IISc Bengaluru, one of India’s premier research institutions founded in 1909. IISc was established by Jamsetji Tata and is headquartered in Bengaluru, Karnataka. The institute has long been at the forefront of scientific research in India and operates under the aegis of the Ministry of Education.
SraVaani stands apart from existing speech recognition systems because of the sheer breadth of its language coverage. While most commercial and research speech AI models focus on a handful of widely spoken languages like Hindi, Bengali, Tamil, and Telugu, SraVaani extends support to over 40 regional and non-scheduled Indian languages that were previously underserved by speech technology. These include languages like Garo, Angika, Chakma, Kokborok, Tulu, Bundeli, and Bajjika, which are spoken by millions but lack robust digital tools for voice interaction.
The People and Organisations Behind SraVaani
Three organisations collaborated to build SraVaani, each bringing a distinct set of capabilities to the project.
SPIRE Lab, IISc Bengaluru
The Signal Processing, Interpretation and Representation (SPIRE) Lab at IISc is the primary research unit behind SraVaani. Led by Professor Prasanta Kumar Ghosh, who also serves as the Principal Investigator of Project Vaani, the lab specialises in speech processing, signal analysis, and natural language understanding. The lab’s work spans speech recognition in agriculture and finance for the poor, dysarthric speech analysis, and multilingual speech synthesis.
Professor Govindan Rangarajan, the Director of IISc, described the release as part of India’s broader ambition to build sovereign AI capabilities. He stated that inclusive language technology must be part of India’s AI ambitions and that SraVaani, serving more than 60 Indian languages, is a contribution toward that goal.
ARTPARK (AI and Robotics Technology Park)
ARTPARK is a not-for-profit Section 8 company established in 2020 at the IISc campus in Bengaluru. It was set up with seed funding of Rs 230 crore (approximately $30 million) from the Department of Science and Technology (DST) under the National Mission on Interdisciplinary Cyber-Physical Systems (NM-ICPS). Of this amount, Rs 170 crore was contributed by the Central Government and Rs 60 crore by the Government of Karnataka.
ARTPARK serves as India’s first dedicated technology park focused exclusively on AI and robotics. It acts as a bridge between academic research and real-world application, running mission-mode R&D projects across healthcare, education, mobility, infrastructure, agriculture, retail, and cybersecurity. In the language technology space, ARTPARK leads the Bhasha AI programme, which includes Project Vaani along with related initiatives like REPIN (Recognising Speech in Indian languages) and SYSPIN (Synthesising Speech in Indian languages).
Google has been the primary funding partner for Project Vaani and provided technical support for SraVaani’s development. Google Research India collaborated with IISc from the early stages of the project to design the data collection methodology and training pipeline. The partnership reflects Google’s broader investment in Indian language AI, which includes support for the Bhashini platform and other Indic language initiatives.
What Makes SraVaani Different?
Most speech recognition systems available today are trained predominantly on data from a small number of dominant languages. This creates a significant gap for the hundreds of millions of Indians who speak regional languages and dialects that do not have enough digital data to train reliable AI models. SraVaani addresses this gap in three important ways.
Pan-India Language Coverage
The model covers languages from every major region of India. Its coverage spans 19 languages from the Northeast, 16 from eastern India, 9 from the west, 8 from the north, 6 from the south, and 5 from central India, along with English and Sanskrit. This geographic spread ensures that the model is not biased toward any single region or language family.
Automatic Language Detection
One of SraVaani’s most practical features is its ability to automatically detect the language being spoken. This eliminates the need for users to manually select their language before speaking, a step that many existing systems require. For a country where bilingual and multilingual speakers are the norm, this feature makes the technology far more usable in everyday settings.
Support for 10 Scripts
The model produces text output across 10 different scripts, including Devanagari, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Odia, Gurmukhi, and the Roman script. This means the model can generate readable text in the native writing system of most Indian languages, rather than forcing users to read transliterated output.
Open-Source and Freely Available
SraVaani has been released under the MIT licence, one of the most permissive open-source software licences in the world. The licence, originating from the Massachusetts Institute of Technology in the late 1980s, allows anyone to use, copy, modify, merge, publish, distribute, sublicense, and even sell copies of the software with minimal restrictions. The only condition is that the original copyright notice and licence text must be included in all copies.
The model is hosted on Hugging Face, the world’s largest open-source AI platform. Founded in 2016 in New York, Hugging Face hosts over 2 million AI models and serves as the central distribution hub for open-source machine learning development. By placing SraVaani on Hugging Face, the developers have made it accessible to startups, researchers, and developers worldwide who can experiment with, adapt, and build upon the model.
The Training Data: Project Vaani
Behind SraVaani lies Project Vaani, one of the largest speech data collection initiatives ever undertaken in India. Launched in December 2022 as a joint effort between IISc, ARTPARK, and Google, Project Vaani was designed to capture how India really speaks, not just how textbooks say it should.
Scale of Data Collection
Project Vaani has recorded more than 31,000 hours of speech from 156,000 people across 165 districts in 28 states and three Union Territories. The dataset represents one of the most linguistically diverse speech corpora in the world, covering 54 languages and numerous dialect variants.
The project uses a district-anchored approach, where more than 1,000 people were randomly selected from each district. This method was chosen to capture the true colour of local speech, including dialects, accents, and natural variations that standardised data collection methods would miss.
How People Were Recorded
Unlike many speech datasets where speakers read pre-written sentences, Project Vaani asked participants to describe images in their own words. This approach captured spontaneous, natural speech that reflects how people actually communicate in everyday life. The recordings include dialects and regional variations that are often lost in scripted data collection.
The dataset also achieved a relatively balanced gender representation, with 45.57 per cent male audio and 54.43 per cent female audio, ensuring that the model learns from voices across genders.
How the Model Was Built
SraVaani is built on a FastConformer architecture, an advanced neural network design optimised for speech recognition. The model has approximately 430 million parameters and was quantised to FP16, resulting in a total model size of about 900 MB.
The training occurred in three stages. First, the model was pretrained from scratch on the complete Vaani speech corpus of 31,255 hours covering 105 languages. Second, it underwent an audio-image alignment stage using 11.8 million audio-image pairs from the Vaani dataset to learn richer audio representations through multimodal relationships. Finally, the model was fine-tuned on approximately 31,270 hours of transcribed speech data spanning 63 Indian languages and dialects, using a combination of the Vaani dataset and multiple open-source speech datasets, including IndicVoices, REPIN, SPRING-INX, SPICOR, and SYSPIN.
How Well Does SraVaani Perform?
SraVaani was evaluated across eight public benchmark datasets covering Indian languages. The results demonstrate that the model performs competitively with leading speech recognition systems on widely supported languages, while significantly outperforming them on low-resource and regional languages.
Lowest Average Word Error Rate
Across the benchmark datasets evaluated, SraVaani achieved the lowest average word error rate (WER) among all systems tested. Word error rate is a standard metric for speech recognition accuracy, measuring the percentage of words incorrectly transcribed. A lower WER indicates better performance.
Strength in Regional Languages
The model’s most striking results come from regional and low-resource languages that existing systems handle poorly. On the Garo benchmark, SraVaani recorded a word error rate of just 9.5 per cent, compared with 69.4 per cent for the next-best system evaluated. Garo is a Tibeto-Burman language spoken primarily in Meghalaya and parts of Assam, and it has historically been one of the most difficult languages for speech recognition systems.
This dramatic improvement demonstrates the value of collecting large, high-quality speech datasets for underserved languages. Where existing systems have little or no training data for languages like Garo, SraVaani benefits from the extensive recordings collected through Project Vaani.
Hindi Benchmark Performance
On the Vaani Benchmark V1.0 for Hindi, SraVaani achieved a WER of 10.9 per cent, establishing a solid baseline for one of India’s most widely spoken languages. While Hindi has more existing speech recognition resources than most Indian languages, SraVaani’s performance shows that it can compete with dedicated Hindi models while simultaneously supporting dozens of other languages.
Why SraVaani Matters for India
The release of SraVaani carries significance that extends well beyond the technical achievements of the model itself. It addresses a fundamental challenge in India’s digital transformation and connects to several broader policy initiatives.
The Digital Divide and Language
India’s digital landscape has been shaped heavily by English and a handful of major Indian languages. According to Census 2011, only about 10 per cent of the Indian population speaks English. The remaining 90 per cent communicate in hundreds of languages and dialects, many of which lack basic digital tools and services. This language gap creates a significant barrier to accessing digital services, from government portals and banking apps to healthcare platforms and educational resources.
SraVaani, by supporting 65 languages and dialects including over 40 that existing systems do not officially support, has the potential to open speech AI capabilities to approximately 25 crore people whose languages are currently underserved by technology.
Connection to the Bhashini Mission
SraVaani aligns closely with the objectives of the Bhashini platform, India’s AI-led language translation initiative under the National Language Translation Mission (NLTM). Launched by Prime Minister Narendra Modi in July 2022 during Digital India Week at Gandhinagar, Gujarat, Bhashini is managed by the Digital India Corporation under the Ministry of Electronics and Information Technology (MeitY).
Bhashini aims to enable easy access to the internet and digital services in Indian languages, including voice-based access, and to help create content in Indian languages. The platform integrates Automatic Speech Recognition (ASR), Machine Translation (MT), and Text-to-Speech (TTS) technologies to provide real-time translation across 22 scheduled languages.
Project Vaani, which provided the training data for SraVaani, is part of the broader Bhashini ecosystem. In July 2024, IISc announced plans to open-source 16,000 hours of speech data from 80 districts under the Bhashini framework, further strengthening the foundation for Indian language AI development.
Sovereign AI and Strategic Autonomy
The release of SraVaani also reflects India’s growing emphasis on building sovereign AI capabilities. By developing an indigenous speech recognition model trained on Indian speech data and released as open-source, India reduces its dependence on foreign speech AI systems that may not adequately represent Indian languages and dialects.
As Professor Prasanta Kumar Ghosh noted, the goal is not only that India builds its own voice AI models, but that those models understand every Indian who speaks to them. This vision of inclusive, sovereign AI is central to India’s broader digital strategy.
The Way Forward
The release of SraVaani marks the beginning of a new phase in Indian language AI development, not the end. Several important directions lie ahead.
Expanding Project Vaani
Project Vaani continues to scale its data collection efforts. The project’s ultimate goal is to create a corpus of over 150,000 hours of speech from approximately 1 million people across all 773 districts of India. The first phase, covering 165 districts, has been completed. Phase two aims to cover 160 additional districts, with a target of 200 hours from about 1,000 people in each district.
As more data is collected, future versions of SraVaani and related models will be able to support even more languages and dialects, further closing the gap between dominant and underserved languages.
Potential Applications
The open-source nature of SraVaani opens up a wide range of potential applications. Developers can fine-tune the model for specific use cases, including voice-based digital services for government portals and citizen services, healthcare applications such as voice-based telemedicine in local languages, education platforms that allow students to interact in their mother tongue, agricultural advisory systems that provide information in regional dialects, and accessibility tools for people with disabilities who rely on voice interfaces.
Challenges Ahead
Despite its achievements, SraVaani faces several challenges. The amount of fine-tuning data varies significantly across the 63 supported languages, which means that recognition accuracy is likely to be lower for languages with limited training data. The model also does not currently support Urdu or Kashmiri, two widely spoken languages in India.
Building robust speech recognition for all Indian languages will require sustained investment in data collection, model development, and community engagement. Projects like Project Vaani, Bhashini, and the open-source release of models like SraVaani provide a strong foundation, but the journey toward truly inclusive voice AI for India is still in its early stages.
Key Takeaways
- SraVaani is an open-source multilingual speech recognition model developed by the SPIRE Lab at IISc Bengaluru, ARTPARK, and Google, released on 13 August 2026 under the MIT licence on Hugging Face.
- The model supports 65 Indian languages and dialects, including 20 scheduled languages and 45 regional languages, covering over 40 languages not supported by existing speech recognition systems.
- SraVaani was trained on data from Project Vaani, which has collected over 31,000 hours of speech from 156,000 people across 165 districts in 28 states.
- The model uses a FastConformer architecture with approximately 430 million parameters and achieves a word error rate of just 9.5 per cent for Garo, compared with 69.4 per cent for the next-best system.
- ARTPARK was established in 2020 at IISc Bengaluru with seed funding of Rs 230 crore from the Department of Science and Technology under the National Mission on Interdisciplinary Cyber-Physical Systems.
- The model connects to India’s broader Bhashini initiative under the National Language Translation Mission, launched in July 2022 by Prime Minister Narendra Modi to enable digital services in Indian languages.