Quickstart: Multi-device Conversation

Artikkeli
01/15/2025

In this quickstart, you'll learn how to use the Speech SDK to create a new multi-device conversation with translation support, as well as join an existing conversation.

Note

The Speech SDK for Java, JavaScript, Objective-C, and Swift support multi-device conversation, but we haven't yet included a guide here.

You can view or download all Speech SDK C# Samples on GitHub.

Prerequisites

Before you get started, make sure to:

Add sample code

Open Program.cs, and replace all code in it with the following code:

using Microsoft.CognitiveServices.Speech;
using Microsoft.CognitiveServices.Speech.Audio;
using Microsoft.CognitiveServices.Speech.Transcription;
using System;
using System.Threading.Tasks;

namespace HelloWorld
{
    class Program
    {
        static async Task Main(string[] args)
        {
            await CreateConversationAsync();
        }

        static async Task CreateConversationAsync()
        {
            // Replace these values with the details of your Cognitive Speech subscription
            string subscriptionKey = "YourSubscriptionKey";

            // Replace below with your region identifier from here: https://aka.ms/speech/sdkregion
            string region = "YourServiceRegion";

            // Sets source and target languages.
            // Replace with the languages of your choice, from list found here: https://aka.ms/speech/sttt-languages
            string fromLanguage = "en-US";
            string toLanguage = "de";

            // Set this to the display name you want for the conversation host
            string displayName = "The host";

            // Create the task completion source that will be used to wait until the user presses Ctrl + C
            var completionSource = new TaskCompletionSource<bool>();

            // Register to listen for Ctrl + C
            Console.CancelKeyPress += (s, e) =>
            {
                completionSource.TrySetResult(true);
                e.Cancel = true; // don't terminate the current process
            };

            // Create an instance of the speech translation config
            var config = SpeechTranslationConfig.FromSubscription(subscriptionKey, region);
            config.SpeechRecognitionLanguage = fromLanguage;
            config.AddTargetLanguage(toLanguage);

            // Create the conversation
            using (var conversation = await Conversation.CreateConversationAsync(config).ConfigureAwait(false))
            {
                // Start the conversation so the host user and others can join
                await conversation.StartConversationAsync().ConfigureAwait(false);

                // Get the conversation ID. It will be up to your scenario to determine how this is shared with other participants.
                string conversationId = conversation.ConversationId;
                Console.WriteLine($"Created '{conversationId}' conversation");

                // At this point, you can use the conversation object to manage the conversation. 
                // For example, to mute everyone else in the room you can call this method:
                await conversation.MuteAllParticipantsAsync().ConfigureAwait(false);

                // Configure which audio source you want to use. If you are using a text only language, you 
                // can use the other overload of the constructor that takes no arguments
                var audioConfig = AudioConfig.FromDefaultMicrophoneInput();
                using (var conversationTranslator = new ConversationTranslator(audioConfig))
                {
                    // You should connect all the event handlers you need at this point
                    conversationTranslator.SessionStarted += (s, e) =>
                    {
                        Console.WriteLine($"Session started: {e.SessionId}");
                    };
                    conversationTranslator.SessionStopped += (s, e) =>
                    {
                        Console.WriteLine($"Session stopped: {e.SessionId}");
                    };
                    conversationTranslator.Canceled += (s, e) =>
                    {
                        switch (e.Reason)
                        {
                            case CancellationReason.EndOfStream:
                                Console.WriteLine($"End of audio reached");
                                break;

                            case CancellationReason.Error:
                                Console.WriteLine($"Canceled due to error. {e.ErrorCode}: {e.ErrorDetails}");
                                break;
                        }
                    };
                    conversationTranslator.ConversationExpiration += (s, e) =>
                    {
                        Console.WriteLine($"Conversation will expire in {e.ExpirationTime.TotalMinutes} minutes");
                    };
                    conversationTranslator.ParticipantsChanged += (s, e) =>
                    {
                        Console.Write("The following participant(s) have ");
                        switch (e.Reason)
                        {
                            case ParticipantChangedReason.JoinedConversation:
                                Console.Write("joined");
                                break;

                            case ParticipantChangedReason.LeftConversation:
                                Console.Write("left");
                                break;

                            case ParticipantChangedReason.Updated:
                                Console.Write("been updated");
                                break;
                        }

                        Console.WriteLine(":");

                        foreach (var participant in e.Participants)
                        {
                            Console.WriteLine($"\t{participant.DisplayName}");
                        }
                    };
                    conversationTranslator.TextMessageReceived += (s, e) =>
                    {
                        Console.WriteLine($"Received an instant message from '{e.Result.ParticipantId}': '{e.Result.Text}'");
                        foreach (var entry in e.Result.Translations)
                        {
                            Console.WriteLine($"\tTranslated into '{entry.Key}': '{entry.Value}'");
                        }
                    };
                    conversationTranslator.Transcribed += (s, e) =>
                    {
                        Console.WriteLine($"Received a transcription from '{e.Result.ParticipantId}': '{e.Result.Text}'");
                        foreach (var entry in e.Result.Translations)
                        {
                            Console.WriteLine($"\tTranslated into '{entry.Key}': '{entry.Value}'");
                        }
                    };
                    conversationTranslator.Transcribing += (s, e) =>
                    {
                        Console.WriteLine($"Received a partial transcription from '{e.Result.ParticipantId}': '{e.Result.Text}'");
                        foreach (var entry in e.Result.Translations)
                        {
                            Console.WriteLine($"\tTranslated into '{entry.Key}': '{entry.Value}'");
                        }
                    };

                    // Enter the conversation to start receiving events
                    await conversationTranslator.JoinConversationAsync(conversation, displayName).ConfigureAwait(false);

                    // You can now send an instant message to all other participants in the room
                    await conversationTranslator.SendTextMessageAsync("The instant message to send").ConfigureAwait(false);

                    // If specified a speech to text language, you can start capturing audio
                    await conversationTranslator.StartTranscribingAsync().ConfigureAwait(false);
                    Console.WriteLine("Started transcribing. Press Ctrl + c to stop");

                    // At this point, you should start receiving transcriptions for what you are saying using the default microphone. Press Ctrl+c to stop audio capture
                    await completionSource.Task.ConfigureAwait(false);

                    // Stop audio capture
                    await conversationTranslator.StopTranscribingAsync().ConfigureAwait(false);

                    // Leave the conversation. After this you will no longer receive events
                    await conversationTranslator.LeaveConversationAsync().ConfigureAwait(false);
                }

                // End the conversation
                await conversation.EndConversationAsync().ConfigureAwait(false);

                // Delete the conversation. Any other participants that are still in the conversation will be removed
                await conversation.DeleteConversationAsync().ConfigureAwait(false);
            }
        }
    }
}

In the same file, replace the string YourSubscriptionKey with your Cognitive Speech subscription key.
Replace the string YourServiceRegion with the region associated with your subscription.
From the menu bar, choose File > Save All.

Build and run the application to create a new conversation

From the menu bar, select Build > Build Solution to build the application. The code should compile without errors now.
Choose Debug > Start Debugging (or press F5) to start the helloworld application.
Once you see the Started transcribing message appear, you can start speaking. You'll see the transcriptions appear as you speak.
- If you share the conversation code with the others and they join the conversation, you'll see their transcriptions as well.
Once you're done speaking, press Ctrl+C to stop audio capture, and end the conversation.

Build and run the application to join an existing conversation

Copy and paste the following function into your Program.cs:

static async Task JoinConversationAsync(string conversationId)
{
    // Set this to the display name you want for the participant
    string displayName = "participant";

    // Set the speech to text, or text language you want to use
    string language = "en-US";

    // Create the task completion source that will be used to wait until the user presses Ctrl + c
    var completionSource = new TaskCompletionSource<bool>();

    // Register to listen for Ctrl+C
    Console.CancelKeyPress += (s, e) =>
    {
        completionSource.TrySetResult(true);
        e.Cancel = true; // don't terminate the current process
    };

    // As a participant, you don't need to specify any subscription key, or region. You can directly create
    // the conversation translator object
    var audioConfig = AudioConfig.FromDefaultMicrophoneInput();
    using (var conversationTranslator = new ConversationTranslator(audioConfig))
    {
        // Register for any events you are interested here. For now let's just register for
        // transcription, and instant message events
        conversationTranslator.TextMessageReceived += (s, e) =>
        {
            Console.WriteLine($"Received an instant message from '{e.Result.ParticipantId}': '{e.Result.Text}'");
            foreach (var entry in e.Result.Translations)
            {
                Console.WriteLine($"\tTranslated into '{entry.Key}': '{entry.Value}'");
            }
        };
        conversationTranslator.Transcribed += (s, e) =>
        {
            Console.WriteLine($"Received a transcription from '{e.Result.ParticipantId}': '{e.Result.Text}'");
            foreach (var entry in e.Result.Translations)
            {
                Console.WriteLine($"\tTranslated into '{entry.Key}': '{entry.Value}'");
            }
        };
        conversationTranslator.Transcribing += (s, e) =>
        {
            Console.WriteLine($"Received a partial transcription from '{e.Result.ParticipantId}': '{e.Result.Text}'");
            foreach (var entry in e.Result.Translations)
            {
                Console.WriteLine($"\tTranslated into '{entry.Key}': '{entry.Value}'");
            }
        };

        // To start receiving events, you will need to join the conversation
        await conversationTranslator.JoinConversationAsync(conversationId, displayName, language).ConfigureAwait(false);

        // You can now send an instant message
        await conversationTranslator.SendTextMessageAsync("Message from participant").ConfigureAwait(false);

        // Start capturing audio if you specified a speech to text language
        await conversationTranslator.StartTranscribingAsync().ConfigureAwait(false);
        Console.WriteLine("Started transcribing. Press Ctrl-C to stop");

        // At this point, you should start receiving transcriptions for what you are saying using
        // the default microphone. Press Ctrl+C to stop audio capture
        await completionSource.Task.ConfigureAwait(false);

        // Stop audio capture
        await conversationTranslator.StopTranscribingAsync().ConfigureAwait(false);

        // Leave the conversation. You will stop receiving events after this
        await conversationTranslator.LeaveConversationAsync().ConfigureAwait(false);
    }
}

Replace CreateConversationAsync(); in your public static async Task Main(string[] args) function with:
```
// Set this to the conversation you want to join
JoinConversationAsync("YourConversationId");
```

You'll need to create a conversation that you can join:
1. Launch your browser and navigate to: https://translator.microsoft.com.
2. Select "Start conversation".
3. Sign in using any of the available options.
4. Type in a name (e.g. The host).
5. Select a language (e.g. English).
6. Select the Enter button.
7. Note the conversation code at the top of the page.
Tip

If you use the copy button on the web page, make sure you remove the translate.it/ at the start. You only need the 5 character conversation ID (example: ABCDE).

Go back to Visual Studio and replace the string YourConversationId with the conversation ID you created in the previous step.
From the menu bar, select Build > Build Solution to build the application. The code should compile without errors now.
Choose Debug > Start Debugging (or press F5) to start the helloworld application.
Once you see the Started transcribing message appear, you can start speaking. You'll see the transcriptions appear as you speak.
- If you go back to your browser, you should see your transcriptions appear there as you speak as well.
Once you're done speaking, press Ctrl+C to stop audio capture, and end the conversation.
Go back to your browser and exit the conversation using the exit button in the upper right corner.

Next Steps

Explore C# samples on GitHub

You can view or download all Speech SDK C++ Samples on GitHub.

Prerequisites

Before you get started, make sure to:

Add sample code

From Visual Studio, open the source file helloworld.cpp.

Replace all the code with the following snippet:

#include "pch.h"

#define WIN32_LEAN_AND_MEAN
#include <Windows.h>

#include <iostream>
#include <thread>
#include <future>
#include <string>
#include <speechapi_cxx.h>

using namespace std::chrono_literals;
using namespace Microsoft::CognitiveServices::Speech;
using namespace Microsoft::CognitiveServices::Speech::Audio;
using namespace Microsoft::CognitiveServices::Speech::Transcription;

// Create a promise we will set when the user presses Ctrl-C
std::promise<bool> promise;
// Create the future we will use to wait for the promise to be fulfilled
std::future<bool> future = promise.get_future();

void StartNewConversation()
{
    // Set these
    std::string subscriptionKey("YourSubscriptionKey");

    // Replace below with your own region identifier from here: https://aka.ms/speech/sdkregion
    std::string region("YourServiceRegion");

    std::string speechLanguage("en-US");

    // Create the conversation object you'll need to manage the conversation
    auto speechConfig = SpeechConfig::FromSubscription(subscriptionKey, region);
    speechConfig->SetSpeechRecognitionLanguage(speechLanguage);
    auto conversation = Conversation::CreateConversationAsync(speechConfig).get();

    // Start the conversation so you and others can join
    conversation->StartConversationAsync().get();

    // Get the conversation ID. It will be up to your scenario to determine how this is shared with other participants.
    std::cout << "Created conversation: " << conversation->GetConversationId() << std::endl;

    // You can now call various commands to manage the room. For example:
    conversation->MuteAllParticipantsAsync().get();

    // Create the conversation translator you'll need to send audio, send IMs, and receive conversation events
    auto audioConfig = AudioConfig::FromDefaultMicrophoneInput();
    auto conversationTranslator = ConversationTranslator::FromConfig(audioConfig);

    // Add any event handlers
    conversationTranslator->SessionStarted += [](const SessionEventArgs& args)
    {
        std::cout << "Session started " << args.SessionId << std::endl;
    };
    conversationTranslator->SessionStopped += [](const SessionEventArgs& args)
    {
        std::cout << "Session stopped " << args.SessionId << std::endl;
    };
    conversationTranslator->Canceled += [](const ConversationTranslationCanceledEventArgs& args)
    {
        switch (args.Reason)
        {
            case CancellationReason::EndOfStream:
                std::cout << "End of audio reached" << std::endl;
                break;

            case CancellationReason::Error:
                std::cout << "Canceled due to error. " << (long)args.ErrorCode << ": " << args.ErrorDetails << std::endl;
                break;
        }
    };
    conversationTranslator->ConversationExpiration += [](const ConversationExpirationEventArgs& args)
    {
        std::cout << "Conversation will expire in " << args.ExpirationTime.count() << " minutes" << std::endl;
    };
    conversationTranslator->ParticipantsChanged += [](const ConversationParticipantsChangedEventArgs& args)
    {
        std::cout << "The following participant(s) have ";
        switch (args.Reason)
        {
            case ParticipantChangedReason::JoinedConversation:
                std::cout << "joined";
                break;

            case ParticipantChangedReason::LeftConversation:
                std::cout << "left";
                break;

            case ParticipantChangedReason::Updated:
                std::cout << "been updated";
                break;
        }

        std::cout << ":" << std::endl;

        for(std::shared_ptr<Participant> participant : args.Participants)
        {
            std::cout << "\t" << participant->DisplayName << std::endl;
        }
    };
    conversationTranslator->Transcribing += [](const ConversationTranslationEventArgs& args)
    {
        std::cout << "Received a partial transcription from " << args.Result->ParticipantId << ": " << args.Result->Text << std::endl;
        for (const auto& entry : args.Result->Translations)
        {
            std::cout << "\t" << entry.first << ": " << entry.second << std::endl;
        }
    };
    conversationTranslator->Transcribed += [](const ConversationTranslationEventArgs& args)
    {
        std::cout << "Received a transcription from " << args.Result->ParticipantId << ": " << args.Result->Text << std::endl;
        for (const auto& entry : args.Result->Translations)
        {
            std::cout << "\t" << entry.first << ": " << entry.second << std::endl;
        }
    };
    conversationTranslator->TextMessageReceived += [](const ConversationTranslationEventArgs& args)
    {
        std::cout << "Received an instant message from " << args.Result->ParticipantId << ": " << args.Result->Text << std::endl;
        for (const auto& entry : args.Result->Translations)
        {
            std::cout << "\t" << entry.first << ": " << entry.second << std::endl;
        }
    };

    // Join the conversation so you can start receiving events
    conversationTranslator->JoinConversationAsync(conversation, "Test Host").get();

    // Send an instant message
    conversationTranslator->SendTextMessageAsync("This is a short test message").get();

    // Start sending audio
    conversationTranslator->StartTranscribingAsync().get();
    std::cout << "Started transcribing. Press Ctrl + C to stop" << std::endl;
    future.get(); // wait for Ctrl - C to be pressed

    // Stop audio capture
    conversationTranslator->StopTranscribingAsync().get();

    // Leave the conversation. You will stop receiving events
    conversationTranslator->LeaveConversationAsync().get();

    // Once you are done, remember to delete the conversation.
    conversation->EndConversationAsync().get(); // You will not be able to rejoin after this
    conversation->DeleteConversationAsync().get(); // All participants still in the room will be ejected
}

int main()
{
    // Register a handler for the Ctrl - C callback
    SetConsoleCtrlHandler(
        [](DWORD dwCtrlType) -> BOOL
        {
            if (dwCtrlType == CTRL_C_EVENT)
            {
                // Signal that the user has pressed ctrl + C
                promise.set_value(true);
                return TRUE;
            }

            return FALSE;
        },
        TRUE);

    StartNewConversation();
}

In the same file, replace the string YourSubscriptionKey with your subscription key.
Replace the string YourServiceRegion with the region associated with your subscription.
From the menu bar, choose File > Save All.

Build and run the application to create a new conversation

From the menu bar, select Build > Build Solution to build the application. The code should compile without errors now.
Choose Debug > Start Debugging (or press F5) to start the helloworld application.
Once you see the Started transcribing message appear, you can start speaking. You'll see the transcriptions appear as you speak
- If you share the conversation code with the others and they join the conversation, you'll see their transcriptions as well.
Once you're done speaking, press Ctrl+C on your keyboard to stop audio capture.

Note

You may see a message from Visual Studio about an exception similar to: Exception thrown at 0x76EB90BF (KernelBase.dll) in helloworld.exe: 0x40010005: Control-C. You can safely ignore this.

Press F5 to continue.

Build and run the application to join an existing conversation

Copy and paste the following function into your helloworld.cpp just before the int main() function:

void JoinExistingConversation(const std::string& conversationId)
{
    std::string speechLanguage("en-US");

    // You'll now need to create a ConversationTranslator to send audio, send IMs, and receive conversation events
    auto audioConfig = AudioConfig::FromDefaultMicrophoneInput();
    auto conversationTranslator = ConversationTranslator::FromConfig(audioConfig);

    // Attach event handlers here. For example:
    conversationTranslator->Transcribing += [](const ConversationTranslationEventArgs& args)
    {
        std::cout << "Received a partial transcription from " << args.Result->ParticipantId << ": " << args.Result->Text << std::endl;
        for (const auto& entry : args.Result->Translations)
        {
            std::cout << "\t" << entry.first << ": " << entry.second << std::endl;
        }
    };
    conversationTranslator->Transcribed += [](const ConversationTranslationEventArgs& args)
    {
        std::cout << "Received a transcription from " << args.Result->ParticipantId << ": " << args.Result->Text << std::endl;
        for (const auto& entry : args.Result->Translations)
        {
            std::cout << "\t" << entry.first << ": " << entry.second << std::endl;
        }
    };
    conversationTranslator->TextMessageReceived += [](const ConversationTranslationEventArgs& args)
    {
        std::cout << "Received an instant message from " << args.Result->ParticipantId << ": " << args.Result->Text << std::endl;
        for (const auto& entry : args.Result->Translations)
        {
            std::cout << "\t" << entry.first << ": " << entry.second << std::endl;
        }
    };

    // Join the conversation
    conversationTranslator->JoinConversationAsync(conversationId, "participant", speechLanguage).get();

    // Start sending audio
    conversationTranslator->StartTranscribingAsync().get();
    std::cout << "Started transcribing. Press Ctrl + C to stop" << std::endl;
    future.get(); // wait for Ctrl - C to be pressed

    // Stop audio capture
    conversationTranslator->StopTranscribingAsync().get();

    // Once you are done, leave the conversation
    conversationTranslator->LeaveConversationAsync().get();
}

Replace StartNewConversation(); in your int main() function with:

// Set this to the conversation you want to join
JoinExistingConversation("YourConversationId");

You'll need to create a conversation that you can join:
1. Launch your browser and navigate to: https://translator.microsoft.com.
2. Select "Start conversation".
3. Sign in using any of the available options.
4. Type in a name (e.g. The host).
5. Select a language (e.g. English).
6. Select the Enter button.
7. Note the conversation code at the top of the page.
Tip

If you use the copy button on the web page, make sure you remove the translate.it/ at the start. You only need the 5 character conversation ID (example: ABCDE).

Go back to Visual Studio and replace YourConversationId in your int main() function with the conversation ID from the previous step.
From the menu bar, select Build > Build Solution to build the application. The code should compile without errors.
Choose Debug > Start Debugging (or press F5) to start the helloworld application.
Once you see the Started transcribing message appear, you can start speaking. You'll see the transcriptions appear as you speak.
- If you go back to your browser, you should see your transcriptions appear there as you speak as well.
Once you're done speaking, press Ctrl+C to stop audio capture, and end the conversation.

Note

You may see a message from Visual Studio about an exception similar to: Exception thrown at 0x76EB90BF (KernelBase.dll) in helloworld.exe: 0x40010005: Control-C. You can safely ignore this.

Press F5 to continue.
Go back to your browser and exit the conversation using the exit button in the upper right corner.

Next steps

Explore C++ samples on GitHub

Jaa

Quickstart: Multi-device Conversation

Prerequisites

Add sample code

Build and run the application to create a new conversation

Build and run the application to join an existing conversation

Next Steps

Prerequisites

Add sample code

Build and run the application to create a new conversation

Build and run the application to join an existing conversation

Next steps

Palaute

Lisäresursseja