Ad
13/8/2026, 9:21:01 pm

IISc Unveils SraVaani, Its Latest Voice AI Model

In a milestone for Indian language technology, the Indian Institute of Science (IISc) in Bengaluru has announced the release of SraVaani, a groundbreaking open-source multilingual speech-recognition model. Launched on August 13, 2026, and developed by the Speech Processing, Interpreting, and Recognition Engineering (SPIRE) Lab at IISc in partnership with ARTPARK and Google, SraVaani is accessible under the liberal MIT licence, encouraging wide adoption and further development.

SraVaani is designed to convert spoken language into text across 65 Indian languages and dialects, reflecting the country’s remarkable linguistic diversity. The model covers 20 scheduled languages officially recognized in the Indian constitution, as well as 45 regional languages and dialects such as Garo, Angika, Chakma, Kokborok, Tulu, Bundeli, and Bajjika. The system supports ten different scripts and incorporates automatic language identification. This key feature eliminates the need for speakers to manually select a language before using the speech recognition tool, making it more user-friendly in multilingual environments.

Researchers say the public release of SraVaani on the Hugging Face platform is aimed at enabling businesses, researchers, and developers to freely use, modify, and redistribute the software. The MIT licence ensures minimal restrictions on its use, which is expected to drive innovation in fields like machine learning, natural language processing, and assistive technology. Hugging Face, a global platform for sharing artificial intelligence models and data sets, provides the infrastructure for broad dissemination and collaborative improvement of the tool.

The technology behind SraVaani was trained on an extensive dataset compiled by Project Vaani, a large-scale initiative to collect natural speech data from across India. The project has amassed over 31,000 hours of conversational speech contributed by 156,000 speakers from 165 districts in 28 states. By incorporating such a rich and diverse dataset, SraVaani demonstrates significant improvements in real-world speech recognition - especially in resource-constrained languages that traditionally suffer from a lack of digital data and technological resources.

Initial performance assessments highlight the model’s superior capabilities. For Garo, a language spoken by nearly a million people but with limited prior digital support, SraVaani achieved a word error rate (WER) of just 9.5 percent. In comparison, the next-best tested system recorded a much higher WER of 69.4 percent for the same language, underscoring the progress made in supporting underrepresented tongues.

SraVaani’s release is viewed as a meaningful step toward closing the digital divide for India’s many languages and dialects. Multilingual speech models rooted in comprehensive datasets offer hope for broader inclusion in digital tools and public services, furthering access and representation for millions of speakers in the rapidly digitizing economy.

Current Affairs
IISc Unveils SraVaani, Its Latest Voice AI Model

More Flips

Ad