Edit

What is MAI-Voice (preview)?

Note

This feature is currently in public preview. This preview is provided without a service-level agreement, and isn't recommended for production workloads. Certain features might not be supported or might have constrained capabilities. For more information, see Supplemental Terms of Use for Microsoft Azure Previews.

MAI-Voice is a family of text-to-speech models available through Azure Speech in Foundry Tools in public preview. Built on Microsoft's in-house speech foundation models, MAI-Voice models produce expressive, natural speech output with consistent voice persona quality.

Speech offers the following MAI-Voice models:

Model Voice Count Key Characteristics Best For
MAI-Voice-2-Flash Prebuilt voices across 15+ languages Ultra-fast low-latency, emotionally rich, highly expressive, multilingual, supports 15 languages and 18 locales, instant voice cloning (gated), fine-grained emotion control via SSML Real-time voice agents and assistants, low-latency call center/IVR flows, multilingual interactive experiences
MAI-Voice-2 Prebuilt voices across 15+ languages Emotionally rich, highly expressive, high-fidelity, multilingual, supports 15 languages and 18 locales, instant voice cloning (gated), long-form generation with speaker consistency, fine-grained emotion control via SSML Expressive long-form content, educational content, Audiobooks/Podcasts, Voice Overs

Model details

MAI‑Voice‑2‑Flash is a text‑to‑speech model built for fast, low‑latency generation. It produces high‑fidelity, natural, and expressive speech across 15 languages and supports gated instant voice cloning, all while being optimized for real‑time responsiveness. Its human‑like intonation, rhythm, and emotional nuance make it ideal for voice agents, assistants, and other interactive scenarios where latency and cost are critical.

You can integrate with MAI-Voice-2-Flash using the Azure Speech SDK via SSML, and also via Voice Live.

Key features

Key features Description
Ultra-fast low-latency synthesis Built for real-time text-to-speech with very low latency, suitable for interactive voice scenarios.
High-fidelity natural synthesis Produces natural, expressive, emotionally rich, and high-clarity voice output with human-like rhythm and intonation.
Multilingual support Supports synthesis across 15 languages and 18 locales.
Emotion and style control Developers can influence speaking style by using SSML with mstts:express-as and style, enabling control over emotions such as joy, excitement, empathy, and more.
Voice prompting with instant cloning (gated) Matches a consented reference voice from a short audio clip (5-60 seconds) without additional training.
Voice library Includes licensed curated voices for 15+ languages that work out of the box for rapid deployment.
Real-time agent optimization Optimized for voice agents, assistants, IVR, and call-center interactions where responsiveness is critical.

Prerequisites

To use MAI related models, complete the following steps:


SSML examples

The examples use the following voices.

Voice ID Gender
en-us-Harper:MAI-Voice-2-Flash Female
en-us-Ethan:MAI-Voice-2-Flash Male

Basic SSML Example (Harper):

<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:mstts='http://www.w3.org/2001/mstts' xml:lang='en-US'>
  <voice xml:lang='en-US' name='en-US-Harper:MAI-Voice-2-Flash'>
    hello world, it's very great
  </voice>
</speak>

Basic SSML Example (Ethan):

<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:mstts='http://www.w3.org/2001/mstts' xml:lang='en-US'>
  <voice xml:lang='en-US' name='en-US-Ethan:MAI-Voice-2-Flash'>
    hello world, it's very great, hello world, it's very great?
  </voice>
</speak>

Expressive Control with SSML mstts:express-as

<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:mstts='http://www.w3.org/2001/mstts' xml:lang='en-US'>
  <voice xml:lang='en-US' name='en-US-Ethan:MAI-Voice-2-Flash'>
    <mstts:express-as style="excited">
       hello world, it's very great, hello world, it's very great?
    </mstts:express-as>
  </voice>
</speak>

Prebuilt voices

Voice Name (ShortName) Locale Language Gender Supported Styles
de-DE-Klaus:MAI-Voice-2-Flash de-DE German (Germany) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
de-DE-Mia:MAI-Voice-2-Flash de-DE German (Germany) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
en-AU-Isla:MAI-Voice-2-Flash en-AU English (Australia) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
en-US-Ethan:MAI-Voice-2-Flash en-US English (United States) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
en-US-Harper:MAI-Voice-2-Flash en-US English (United States) Female angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, regretful, relieved, sad, shouting, softvoice, whispering
en-US-Olivia:MAI-Voice-2-Flash en-US English (United States) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
es-ES-Marta:MAI-Voice-2-Flash es-ES Spanish (Spain) Female adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
es-MX-Alejo:MAI-Voice-2-Flash es-MX Spanish (Mexico) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
es-MX-Valeria:MAI-Voice-2-Flash es-MX Spanish (Mexico) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
fr-FR-Marc:MAI-Voice-2-Flash fr-FR French (France) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
fr-FR-Soleil:MAI-Voice-2-Flash fr-FR French (France) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
hi-IN-Arjun:MAI-Voice-2-Flash hi-IN Hindi (India) Male angry, confused, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, sad, surprised
hi-IN-Dhruv:MAI-Voice-2-Flash hi-IN Hindi (India) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
hi-IN-Kavya:MAI-Voice-2-Flash hi-IN Hindi (India) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
hi-IN-Priya:MAI-Voice-2-Flash hi-IN Hindi (India) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
hu-HU-Bence:MAI-Voice-2-Flash hu-HU Hungarian (Hungary) Male
hu-HU-Levente:MAI-Voice-2-Flash hu-HU Hungarian (Hungary) Male
hu-HU-Lilla:MAI-Voice-2-Flash hu-HU Hungarian (Hungary) Female
hu-HU-Réka:MAI-Voice-2-Flash hu-HU Hungarian (Hungary) Female
it-IT-Luca:MAI-Voice-2-Flash it-IT Italian (Italy) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
it-IT-Rosa:MAI-Voice-2-Flash it-IT Italian (Italy) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
ko-KR-Haena:MAI-Voice-2-Flash ko-KR Korean (Korea) Female angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, regretful, relieved, sad, softvoice, surprised
ko-KR-Junho:MAI-Voice-2-Flash ko-KR Korean (Korea) Male angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, relieved, sad, softvoice
nl-NL-Sander:MAI-Voice-2-Flash nl-NL Dutch (Netherlands) Male adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
pt-BR-Caio:MAI-Voice-2-Flash pt-BR Portuguese (Brazil) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
pt-BR-Luana:MAI-Voice-2-Flash pt-BR Portuguese (Brazil) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
pt-BR-Pedro:MAI-Voice-2-Flash pt-BR Portuguese (Brazil) Male confused, determined, embarrassed, excited, happy, hopeful, joyful, regretful, relieved, sad, softvoice, surprised
pt-BR-Rafael:MAI-Voice-2-Flash pt-BR Portuguese (Brazil) Male angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, regretful, relieved, sad, softvoice, surprised
pt-PT-Rui:MAI-Voice-2-Flash pt-PT Portuguese (Portugal) Male angry, confused, determined, embarrassed, excited, happy, hopeful, joyful, regretful, relieved, sad, softvoice, surprised
ro-RO-Andrei:MAI-Voice-2-Flash ro-RO Romanian (Romania) Male
ro-RO-Elena:MAI-Voice-2-Flash ro-RO Romanian (Romania) Female
ro-RO-Ioana:MAI-Voice-2-Flash ro-RO Romanian (Romania) Female
ro-RO-Radu:MAI-Voice-2-Flash ro-RO Romanian (Romania) Male
ru-RU-Lev:MAI-Voice-2-Flash ru-RU Russian (Russia) Male adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
ru-RU-Masha:MAI-Voice-2-Flash ru-RU Russian (Russia) Female adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
th-TH-Krit:MAI-Voice-2-Flash th-TH Thai (Thailand) Male adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
th-TH-Nattapong:MAI-Voice-2-Flash th-TH Thai (Thailand) Male adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
tr-TR-Aydın:MAI-Voice-2-Flash tr-TR Turkish (Türkiye) Male adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
tr-TR-Elif:MAI-Voice-2-Flash tr-TR Turkish (Türkiye) Female adventurous, caringempathy, curious, encouraging, excited, friendlycheerful, nostalgic, reflective, saddisappointed, serious
zh-CN-Bo:MAI-Voice-2-Flash zh-CN Chinese (Mandarin, Simplified) Male angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
zh-CN-Lan:MAI-Voice-2-Flash zh-CN Chinese (Mandarin, Simplified) Female angry, confused, disgusted, embarrassed, excited, fearful, happy, joyful, sad, surprised
zh-CN-Mei:MAI-Voice-2-Flash zh-CN Chinese (Mandarin, Simplified) Female angry, confused, determined, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, relieved, sad, shouting, softvoice, surprised, whispering
zh-CN-Wei:MAI-Voice-2-Flash zh-CN Chinese (Mandarin, Simplified) Male angry, confused, disgusted, embarrassed, excited, fearful, happy, hopeful, jealous, joyful, regretful, sad, surprised

Usage: Available for third-party developers. Microsoft holds full licensing rights for commercial use.

Use MAI-Voice models

MAI-Voice models use the same Azure Speech APIs and SDKs as other Azure neural and HD voices. Use the voice name in the name attribute of the SSML <voice> element. See the prebuilt voice tables in the preceding sections for available names.

Try a MAI-Voice voice in the Foundry portal:

  1. Go to the Text to speech feature page and select Open in playground.
  2. Select a MAI-Voice voice from the voice dropdown.
  3. Enter sample text in the text box.
  4. Select Play to hear the synthesized speech.

Send an SSML POST request to the cognitiveservices/v1 endpoint of your Speech resource. Replace YourRegion with your Speech resource region and <YourSpeechResourceKey> with your resource key.

curl -X POST \
  "https://YourRegion.tts.speech.microsoft.com/cognitiveservices/v1" \
  --header "Content-Type: application/ssml+xml" \
  --header "X-Microsoft-OutputFormat: audio-24khz-160kbitrate-mono-mp3" \
  --header "Ocp-Apim-Subscription-Key: <YourSpeechResourceKey>" \
  --data '<speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xmlns:mstts="http://www.w3.org/2001/mstts" xml:lang="en-US">
  <voice name="en-US-Harper:MAI-Voice-2">
    Hello, this is a sample from MAI Voice.
  </voice>
</speak>' \
  --output output.mp3

Replace the voice name with any MAI-Voice voice from the prebuilt voice tables above.

The following code synthesizes speech by using the Azure Speech SDK and saves the audio to output.mp3. Replace <key> with your Speech resource key.

import azure.cognitiveservices.speech as speechsdk

speech_config = speechsdk.SpeechConfig(
    subscription="<key>", region="eastus"
)
audio_config = speechsdk.audio.AudioOutputConfig(filename="output.mp3")
speech_config.set_speech_synthesis_output_format(
    speechsdk.SpeechSynthesisOutputFormat.Audio24Khz160KBitRateMonoMp3
)
synthesizer = speechsdk.SpeechSynthesizer(
    speech_config=speech_config, audio_config=audio_config
)

ssml = """
<speak version='1.0' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:mstts='http://www.w3.org/2001/mstts' xml:lang='en-US'>
  <voice name='en-US-Harper:MAI-Voice-2-Flash'>
    <mstts:express-as style="excitement">Hello world.</mstts:express-as>
  </voice>
</speak>
"""

synthesizer.speak_ssml_async(ssml).get()

On success, an output.mp3 file is saved to the current directory. Replace the voice name with any MAI-Voice voice from the prebuilt voice tables above.

Follow the text to speech quickstart. In the SSML, use a MAI-Voice voice name in the name attribute of the <voice> element. See the prebuilt voice tables above for available names.

Follow the text to speech quickstart. In the SSML, use a MAI-Voice voice name in the name attribute of the <voice> element. See the prebuilt voice tables above for available names.

Follow the text to speech quickstart. In the SSML, use a MAI-Voice voice name in the name attribute of the <voice> element. See the prebuilt voice tables above for available names.

Next steps