Friday, 2 Oct 2026 Ashwin Krishna Shashthi · VS 2083
Sensex71,909.70▼0.79%Nifty 5022,421.95▼0.88%Bank Nifty54,450.75▼0.33% USD₹96.34 BTC₹81,39,290▼0.52%
--° Delhi
MISHRA NEWS WORLDIndia · States · Money · Live updates
BREAKING
दिल्ली में Goldy Brar गैंग से जुड़े कथित रंगदारी नेटवर्क का भंडाफोड़, 3 आरोपी गिरफ्तार; कारोबारी से मांगे थे ₹1 करोड़उम्र बढ़ने के साथ बदलती है Nutrition की जरूरत, 20s से 50s तक ऐसी होनी चाहिए आपकी Plateसारण में बाढ़ का कहर: तरैया में टूटा रिंग बांध, कई गांवों में घुसा गंडक का पानीTigress T138 finds a new home at Buxa Tiger ReserveA.P. government to note media commission proposal: Minister Kondapalli SrinivasProtests held across country demanding CEC Gyanesh Kumar’s ouster, hundreds detainedDid Cornell frat brother apologize to Jane Doe after alleged rape? Stunning details outFuel inflation explained: How changes in petrol, CNG, LPG and PNG prices could affect your household budgetHow Ankita, Kirti and Kumkum gave their fathers’ faith a golden finish at Asian GamesUPI volume jumps 27% to touch 145 bn in first halfNetanyahu warns 'very heavy price' if any power behind flydubai incidentManipur: Security forces arrest three cadres linked to PREPAK (Pro), UNLFQuote of the day by Zac EfronIndian doctor says European doctors were shocked to know about India’s quick access to cancer and other treatmentsNEC Secy Bhalla calls for integrated coffee development in NagalandTTC issues directive over ENPO-FNTA process, impersonated executivesEnforcement or Reform? The unending debate over NLTP Act
Technology

Microsoft ने लॉन्च किए नए AI Voice Models, 60 भाषाओं में Real-Time Transcription से बदलेगा Voice AI का अनुभव

Microsoft ने तीन नए AI models लॉन्च किए हैं, जिनमें MAI-Transcribe-2-Streaming real-time speech को text में बदल सकता है, जबकि MAI-Voice-2.1 और Flash natural voice generation पर फोकस करते हैं। नए models voice assistants, customer service, multilingual applications और AI agents के लिए बनाए गए हैं।

Microsoft’s new MAI AI voice models focus on real-time transcription and natural multilingual voice generation. | Photo: Microsoft AI

नई दिल्ली, 2 अक्टूबर 2026: Microsoft ने Artificial Intelligence (AI) के क्षेत्र में बड़ा अपडेट करते हुए तीन नए voice models लॉन्च किए हैं। कंपनी ने MAI-Transcribe-2-Streaming, MAI-Voice-2.1 और MAI-Voice-2.1-Flash पेश किए हैं।

इन नए models का उद्देश्य AI को इंसानों के साथ ज्यादा तेज, natural और real-time बातचीत करने में सक्षम बनाना है। Microsoft के मुताबिक, इनका इस्तेमाल customer-service agents, multilingual assistants, interactive learning और voice-based applications में किया जा सकता है।

MAI-Transcribe-2-Streaming क्या है?

Microsoft का नया MAI-Transcribe-2-Streaming कंपनी का streaming transcription model है, जो किसी व्यक्ति के बोलते समय ही speech को text में बदलना शुरू कर देता है।

Advertisementrajeshwar groupBUILDING INDIA’S FUTURE TOGETHER Rajeshwar Group is a diversified business conglomerate deBUILDING INDIA’S FUTURE TOGETHER Rajeshwar Group is a diversified business conglomerate delivering excellence in Manufacturing, Infrastructure, Technology and BKnow more

आमतौर पर speech-to-text systems में user के बोलना पूरा करने के बाद transcription तैयार होती है। नए streaming model में शुरुआती transcription results बातचीत के दौरान ही मिलने लगते हैं।

Microsoft के अनुसार, model audio प्राप्त होने के 100 milliseconds से कुछ अधिक समय में शुरुआती transcription hypotheses देना शुरू कर सकता है।

“It produces its first hypotheses in just over 100ms of receiving audio.”

— Microsoft AI

60 भाषाओं में Real-Time Transcription

MAI-Transcribe-2-Streaming को 60 भाषाओं में real-time transcription के लिए तैयार किया गया है। इसमें automatic और continuous language detection की सुविधा भी दी गई है।

इसका मतलब है कि voice-based applications बातचीत के दौरान भाषा को पहचानकर speech को text में बदल सकती हैं।

AdvertisementTech RajeshwarNeed a new website?Websites & apps built for your business — contact Tech RajeshwarContact us

Microsoft का कहना है कि streaming transcription की मदद से voice agents user की बात पूरी होने से पहले ही processing या दूसरे actions की तैयारी शुरू कर सकते हैं।

AI Voice Agents के लिए क्यों महत्वपूर्ण है?

Voice-based AI में एक सामान्य बातचीत के दौरान system को कई काम करने पड़ते हैं—पहले आवाज सुनना, फिर उसे समझना, जवाब तैयार करना और अंत में आवाज में जवाब देना।

अगर इनमें से किसी चरण में ज्यादा delay हो तो बातचीत robotic महसूस हो सकती है।

Microsoft के नए models इसी latency को कम करने पर फोकस करते हैं।

  • Speech को real time में text में बदलना
  • User की बात के दौरान processing शुरू करना
  • Natural-sounding voice में जवाब देना
  • कई भाषाओं में बातचीत करना
  • Customer-service applications में तेजी से response देना
  • Live captions और voice interfaces को बेहतर बनाना

MAI-Voice-2.1 में 23 भाषाओं का support

Microsoft ने MAI-Voice-2.1 नाम का नया text-to-speech model भी पेश किया है।

कंपनी के अनुसार, यह model 23 भाषाओं और 26 locales को support करता है। इसमें एक ही voice identity को अलग-अलग भाषाओं में बनाए रखने की सुविधा दी गई है।

Supported languages में Hindi और English (India) भी शामिल हैं।

इससे multilingual AI assistants किसी user की भाषा में जवाब देने के साथ एक consistent voice बनाए रख सकते हैं।

MAI-Voice-2.1-Flash ज्यादा speed के लिए

Microsoft ने MAI-Voice-2.1-Flash को high-volume और latency-sensitive applications के लिए तैयार किया है।

कंपनी के अनुसार, Flash version लगभग 55% faster model inference देता है और comparable models की तुलना में लगभग 60% कम cost के लिए designed है।

Microsoft ने इसकी कीमत $15 प्रति 1 million characters बताई है, जबकि MAI-Voice-2.1 की कीमत $22 प्रति 1 million characters है।

Voice Cloning में भी नई सुविधा

दोनों MAI voice models short reference audio की मदद से voice matching और cloning capabilities को support करते हैं।

Microsoft के अनुसार, इसके साथ consent guardrails भी मौजूद हैं, जिनका उद्देश्य voice technology के गलत इस्तेमाल को रोकने में मदद करना है।

Voice cloning के बढ़ते इस्तेमाल के बीच consent और authorized voice use जैसे मुद्दे AI industry में महत्वपूर्ण बने हुए हैं।

किन क्षेत्रों में होगा इस्तेमाल?

Microsoft के अनुसार, developers इन models का इस्तेमाल कई तरह के applications बनाने के लिए कर सकते हैं।

  • Customer service: AI customer-service agents बातचीत के दौरान request को समझकर response तैयार कर सकते हैं।
  • Multilingual assistants: अलग-अलग भाषाओं में बातचीत करने वाले AI assistants बनाए जा सकते हैं।
  • Education: Interactive learning और tutoring applications में natural voice का इस्तेमाल किया जा सकता है।
  • Live transcription: बातचीत के दौरान speech को text में बदला जा सकता है।
  • Voice interfaces: ऐसे applications बनाए जा सकते हैं जिनमें voice मुख्य interaction method हो।

Developers के लिए उपलब्ध हैं नए Models

Microsoft के अनुसार, तीनों models Microsoft Foundry और MAI Playground के जरिए उपलब्ध हैं। कंपनी ने Vercel को भी supported platforms में शामिल किया है।

MAI-Voice-2.1 और MAI-Voice-2.1-Flash को OpenRouter के जरिए भी access किया जा सकता है, जबकि LiveKit support को आने वाले समय के लिए बताया गया है।

Microsoft का बढ़ता AI Model Ecosystem

Microsoft पिछले कुछ महीनों में अपनी MAI model family का विस्तार कर रहा है। कंपनी ने voice, transcription, image, coding और reasoning जैसे अलग-अलग क्षेत्रों में अपने AI models पेश किए हैं।

सितंबर 2026 में कंपनी ने MAI-Transcribe-2 भी पेश किया था, जबकि अब नया streaming version real-time voice interaction पर फोकस करता है।

इससे Microsoft का AI development अब केवल text-based applications तक सीमित नहीं रह गया है, बल्कि real-time voice interaction की दिशा में भी आगे बढ़ रहा है।

AI बातचीत को ज्यादा Natural बनाने की कोशिश

नए models के जरिए Microsoft voice AI के दो महत्वपूर्ण हिस्सों पर एक साथ काम कर रहा है।

एक तरफ MAI-Transcribe-2-Streaming user की आवाज को तेजी से समझने की कोशिश करता है, वहीं दूसरी तरफ MAI-Voice-2.1 और MAI-Voice-2.1-Flash natural voice में जवाब देने के लिए बनाए गए हैं।

इस combination का इस्तेमाल ऐसे AI agents बनाने में किया जा सकता है जो सुनने, समझने और बोलने के बीच कम से कम delay रखें।

“A voice agent is a loop. It has to hear, understand, decide, and speak.”

— Microsoft AI

मुख्य बातें

  • 3 नए AI models: Microsoft ने MAI-Transcribe-2-Streaming, MAI-Voice-2.1 और MAI-Voice-2.1-Flash लॉन्च किए हैं।
  • 60 भाषाएं: Streaming transcription model 60 भाषाओं में real-time transcription support करता है।
  • 23 भाषाएं: MAI-Voice-2.1 और Flash 23 भाषाओं और 26 locales को support करते हैं।
  • Hindi support: Voice models की supported languages में Hindi और English (India) शामिल हैं।
  • Low latency: Streaming model audio मिलने के 100 milliseconds से कुछ अधिक समय में शुरुआती transcription results देना शुरू कर सकता है।
  • Voice cloning: Voice models short reference audio के आधार पर voice matching capabilities support करते हैं।
  • Developer access: Models Microsoft Foundry और MAI Playground सहित कई platforms पर उपलब्ध हैं।

AI Voice Technology का अगला चरण

Real-time transcription और natural text-to-speech को एक साथ जोड़ने से voice-based AI applications के लिए नई संभावनाएं खुल सकती हैं।

Customer service, education, accessibility और multilingual communication जैसे क्षेत्रों में ऐसे AI agents विकसित किए जा सकते हैं जो बातचीत को ज्यादा natural तरीके से संभालें।

Microsoft के नए MAI models इस दिशा में कंपनी के latest AI developments में शामिल हैं और आने वाले समय में voice-based AI applications के विकास को प्रभावित कर सकते हैं।

निष्कर्ष

Microsoft के नए AI voice models real-time speech recognition और natural voice generation को एक साथ आगे बढ़ाने की कोशिश हैं।

MAI-Transcribe-2-Streaming का फोकस user की speech को बातचीत के दौरान समझने पर है, जबकि MAI-Voice-2.1 और MAI-Voice-2.1-Flash natural और तेज voice responses देने के लिए बनाए गए हैं।

इन technologies का इस्तेमाल AI assistants, customer service, education, multilingual communication और voice-driven applications में किया जा सकता है।

Frequently asked questions

Microsoft ने कौन से नए AI models लॉन्च किए हैं?

Microsoft ने MAI-Transcribe-2-Streaming, MAI-Voice-2.1 और MAI-Voice-2.1-Flash लॉन्च किए हैं।

MAI-Transcribe-2-Streaming क्या करता है?

यह बातचीत के दौरान speech को real time में text में बदलने के लिए बनाया गया streaming transcription model है।

MAI-Transcribe-2-Streaming कितनी भाषाओं को support करता है?

Microsoft के ?

क्या नए Microsoft AI voice models में Hindi support है?

हां, MAI-Voice-2.1 और MAI-Voice-2.1-Flash की supported languages में Hindi और English (India) शामिल हैं।

MAI-Voice-2.1-Flash किसके लिए बनाया गया है?

इसे high-volume और latency-sensitive voice applications के लिए optimize किया गया है।

AdvertisementCallZenixAI Voice Bot, Dialer & IVR SolutionsCloud dialer and contact-center software for your teamKnow more

Spotted a mistake? See our corrections policy or write to grievance@mishraworld.com.

Advertisement320 × 50Advertise hereReach readers across India · Contact us
MISHRA NEWS
Sections
More