Edit

Add voice talent consent for professional voice

A voice talent is an individual or target speaker whose voices are recorded and used to create neural voice models.

Before you can fine-tune a professional voice, you must submit a recording of the voice talent's consent statement. The voice talent statement is a recording of the voice talent reading a statement that they consent to the usage of their speech data for professional voice fine-tuning. The consent statement is also used to verify that the voice talent is the same person as the speaker in the fine-tuning data.

Tip

Before you get started in the Microsoft Foundry portal, define your voice persona and choose the right voice talent.

You can find the verbal consent statement in multiple languages on GitHub. The language of the verbal statement must be the same as your recording. See also the disclosure for voice talent.

Add voice talent

Tip

For a sample consent statement and training data, see the GitHub repository.

To add a voice talent profile and upload their consent statement, follow these steps:

These steps continue from the Customize a model page you opened in Set up a professional voice.

If you closed the page, resume your draft customization before continuing.

Tip

If you need sample consent and training files, select Download sample data on the Register voice talent step.

  1. On the Register voice talent step, select Add voice talent.

  2. In the Add voice talent pane, provide the voice talent details and upload the recorded verbal consent statement:

    • For Target scenario, select one or more scenarios that match the intended use of the voice.
    • Optionally, for Voice characteristics, describe the characteristics of the voice.
    • Enter the Voice talent name. The name must match the person who recorded the consent statement, in the same language used in the recording.
    • Enter the Company name. The company name must match what was spoken in the recording, in the same language.
    • Drag the .mp3 or .wav file into the upload area, or select Browse for a file to select it.
    • Make sure the verbal statement was recorded with the same settings, environment, and speaking style as your fine-tuning data.
  3. Select Upload.

  4. Wait for the voice talent status to become Succeeded. If processing fails, review the reported error and check the recording, consent statement, and matching voice talent and company names.

  5. On the Register voice talent step, select the voice talent you just added, and then select Next.

Continue with Add training datasets.

View or delete an existing voice talent

If you uploaded the wrong consent recording, find the voice talent entry on the Data tab under Services. This tab is separate from Data in the left navigation.

  1. Open the Foundry project that contains the voice talent.
  2. Select Build > Services, and then select the Data tab.
  3. Find the entry with the Voice Talent type and Text to Speech tag. Select its name to review the details and play the consent recording.

To delete an unwanted voice talent entry:

  1. Return to Services > Data.
  2. In the voice talent's row, open the Actions menu (three dots), and then select Delete.
  3. In the Delete AI service resource dialog, confirm that the displayed name matches the voice talent you want to remove.
  4. Select Delete to confirm, or Cancel to keep the entry.

Important

Deletion can't be undone. Confirm that you selected the unwanted voice talent entry before deleting it.

To upload the correct recording, return to Register voice talent in your customization and follow Add voice talent.

Next steps

A voice talent is an individual or target speaker whose voices are recorded and used to create neural voice models.

Before you can fine-tune a professional voice, you must submit a recording of the voice talent's consent statement. The voice talent statement is a recording of the voice talent reading a statement that they consent to the usage of their speech data for professional voice fine-tuning. The consent statement is also used to verify that the voice talent is the same person as the speaker in the fine-tuning data.

Tip

Before you get started in Speech Studio, define your voice persona and choose the right voice talent.

You can find the verbal consent statement in multiple languages on GitHub. The language of the verbal statement must be the same as your recording. See also the disclosure for voice talent.

Add voice talent

To add a voice talent profile and upload their consent statement, follow these steps:

  1. Sign in to the Speech Studio.
  2. Select Custom voice > Your project name > Set up voice talent > Add voice talent.
  3. In the Add new voice talent wizard, describe the characteristics of the voice you're going to create. The scenarios that you specify here must be consistent with what you provided in the application form.
  4. Select Next.
  5. On the Upload voice talent statement page, follow the instructions to upload the voice talent statement you've recorded beforehand. Make sure the verbal statement was recorded with the same settings, environment, and speaking style as your fine-tuning data. Screenshot of the voice talent statement upload dialog.
  6. Enter the voice talent name and company name. The voice talent name must be the name of the person who recorded the consent statement. Enter the name in the same language used in the recorded statement. The company name must match the company name that was spoken in the recorded statement. Ensure the company name is entered in the same language as the recorded statement.
  7. Select Next.
  8. Review the voice talent and persona details, and select Submit.

After the voice talent status is Succeeded, you can proceed to train your custom voice model.

Next steps

With the professional voice feature, it's required that every voice be created with explicit consent from the user. A recorded statement from the user is required acknowledging that the customer (Azure Speech in Foundry Tools resource owner) will create and use their voice.

To add voice talent consent to the professional voice project, you get the prerecorded consent audio file from a publicly accessible URL (Consents_Create) or upload the audio file (Consents_Post). In this article, you add consent from a URL.

You need an audio recording of the user speaking the consent statement.

You can get the consent statement text for each locale from the text to speech GitHub repository. See SpeakerAuthorization.txt for the consent statement for the en-US locale:

"I  [state your first and last name] am aware that recordings of my voice will be used by [state the name of the company] to create and use a synthetic version of my voice."

To add consent to a professional voice project from the URL of an audio file, use the Consents_Create operation of the custom voice API. Construct the request body according to the following instructions:

  • Set the required projectId property. See create a project.
  • Set the required voiceTalentName property. The voice talent name must be the name of the person who recorded the consent statement. Enter the name in the same language used in the recorded statement. The voice talent name can't be changed later.
  • Set the required companyName property. The company name must match the company name spoken in the recorded statement. Ensure the company name is entered in the same language as the recorded statement. The company name can't be changed later.
  • Set the required audioUrl property. The URL of the voice talent consent audio file. Use a URI with the shared access signatures (SAS) token.
  • Set the required locale property. This should be the locale of the consent. The locale can't be changed later. You can find the text to speech locale list here.

Make an HTTP PUT request using the URI as shown in the following Consents_Create example.

  • Replace YourResourceKey with your Speech resource key.
  • Replace YourResourceName with your Speech resource name.
  • Replace JessicaConsentId with a consent ID of your choice. The case sensitive ID will be used in the consent's URI and can't be changed later.
curl -v -X PUT -H "Ocp-Apim-Subscription-Key: YourResourceKey" -H "Content-Type: application/json" -d '{
  "description": "Consent for Jessica voice",
  "projectId": "ProjectId",
  "voiceTalentName": "Jessica Smith",
  "companyName": "Contoso",
  "audioUrl": "https://contoso.blob.core.windows.net/public/jessica-consent.wav?mySasToken",
  "locale": "en-US"
} '  "https://YourResourceName.cognitiveservices.azure.com/customvoice/consents/JessicaConsentId?api-version=2026-01-01"

You should receive a response body in the following format:

{
  "id": "JessicaConsentId",
  "description": "Consent for Jessica voice",
  "projectId": "ProjectId",
  "voiceTalentName": "Jessica Smith",
  "companyName": "Contoso",
  "locale": "en-US",
  "status": "NotStarted",
  "createdDateTime": "2023-04-01T05:30:00.000Z",
  "lastActionDateTime": "2023-04-02T10:15:30.000Z"
}

The response header contains the Operation-Location property. Use this URI to get details about the Consents_Create operation. Here's an example of the response header:

Operation-Location: https://YourResourceName.cognitiveservices.azure.com/customvoice/operations/aaaabbbb-0000-cccc-1111-dddd2222eeee?api-version=2026-01-01
Operation-Id: aaaabbbb-0000-cccc-1111-dddd2222eeee

Next steps