Samouczek: tworzenie aplikacji do notatek głosowych zamieniającej mowę na tekst

Skompiluj aplikację, która konwertuje dźwięk mówiony na zorganizowane notatki — w całości na urządzeniu. Aplikacja najpierw transkrybuje plik audio przy użyciu modelu zamiany mowy na tekst, a następnie używa modelu czatu do podsumowania i organizowania transkrypcji w czyste notatki.

W tym poradniku nauczysz się, jak:

  • Konfigurowanie projektu i instalowanie lokalnego zestawu SDK rozwiązania Foundry
  • Ładowanie modelu zamiany mowy na tekst i transkrypcja pliku audio
  • Ładowanie modelu czatu i podsumowywanie transkrypcji
  • Łączenie transkrypcji i podsumowania w pełną aplikację
  • Uprzątnij zasoby

Wymagania wstępne

  • Komputer Windows, macOS lub Linux z co najmniej 8 GB pamięci RAM.
  • .wav Plik audio do transkrypcji (w tym samouczku jest używany przykładowy plik).

Instalowanie pakietów

Repozytorium przykładów

Pełny kod przykładowy do tego artykułu można znaleźć w repozytorium GitHub przykładów platformy Foundry. Aby sklonować repozytorium i przejść do przykładu, użyj następującego polecenia:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/csharp/foundry-local/tutorial-voice-to-text

Jeśli programujesz lub wysyłasz na Windows, wybierz kartę Windows. Pakiet Windows integruje się z środowiskiem uruchomieniowym Windows ML — zapewnia ten sam obszar powierzchni interfejsu API z szerszym zakresem przyspieszania sprzętowego.

dotnet add package Microsoft.AI.Foundry.Local.WinML
dotnet add package OpenAI

Przykłady języka C# w repozytorium GitHub to wstępnie skonfigurowane projekty. Jeśli tworzysz od podstaw, zapoznaj się z dokumentacją zestawu SDK lokalnego rozwiązania Foundry , aby uzyskać więcej informacji na temat sposobu konfigurowania projektu w języku C# przy użyciu rozwiązania Foundry Local.

Transkrypcja pliku audio

W tym kroku załadujesz model zamiany mowy na tekst i transkrybujesz plik audio. Lokalny zestaw SDK Foundry używa aliasu whisper modelu, aby wybrać najlepszy wariant Whisper dla twojego urządzenia.

  • Otwórz Program.cs i zastąp jego zawartość następującym kodem, aby zainicjować zestaw SDK, załadować model mowy i transkrypować plik audio:

    // Load the speech-to-text model
    var speechModel = await catalog.GetModelAsync("whisper-tiny")
        ?? throw new Exception("Speech model not found");
    
    await speechModel.DownloadAsync(progress =>
    {
        Console.Write($"\rDownloading speech model: {progress:F2}%");
        if (progress >= 100f) Console.WriteLine();
    });
    
    await speechModel.LoadAsync();
    Console.WriteLine("Speech model loaded.");
    
    // Transcribe the audio file
    var audioClient = await speechModel.GetAudioClientAsync();
    var transcriptionText = new StringBuilder();
    
    Console.WriteLine("\nTranscription:");
    var audioResponse = audioClient
        .TranscribeAudioStreamingAsync("meeting-notes.wav", ct);
    await foreach (var chunk in audioResponse)
    {
        Console.Write(chunk.Text);
        transcriptionText.Append(chunk.Text);
    }
    Console.WriteLine();
    
    // Unload the speech model to free memory
    await speechModel.UnloadAsync();
    

    Metoda GetAudioClientAsync zwraca klienta dla operacji audio. Metoda TranscribeAudioStreamingAsync przesyła strumieniowo fragmenty transkrypcji w miarę ich dostępności. Zakumulujesz tekst, aby móc przekazać go do modelu czatu w następnym kroku.

Uwaga / Notatka

Zastąp "meeting-notes.wav" ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Podsumowanie transkrypcji

Teraz użyj modelu czatu, aby uporządkować surowe transkrypcje w uporządkowane notatki. Załaduj qwen2.5-0.5b model i wyślij transkrypcję jako kontekst z monitem systemowym, który nakazuje modelowi tworzenie czystych, podsumowanych notatek.

Dodaj następujący kod po kroku transkrypcji:

// Load the chat model for summarization
var chatModel = await catalog.GetModelAsync("qwen2.5-0.5b")
    ?? throw new Exception("Chat model not found");

await chatModel.DownloadAsync(progress =>
{
    Console.Write($"\rDownloading chat model: {progress:F2}%");
    if (progress >= 100f) Console.WriteLine();
});

await chatModel.LoadAsync();
Console.WriteLine("Chat model loaded.");

// Summarize the transcription into organized notes
var chatClient = await chatModel.GetChatClientAsync();
var messages = new List<ChatMessage>
{
    new ChatMessage
    {
        Role = "system",
        Content = "You are a note-taking assistant. Summarize " +
                  "the following transcription into organized, " +
                  "concise notes with bullet points."
    },
    new ChatMessage
    {
        Role = "user",
        Content = transcriptionText.ToString()
    }
};

var chatResponse = await chatClient.CompleteChatAsync(messages, ct);
var summary = chatResponse.Choices[0].Message.Content;
Console.WriteLine($"\nSummary:\n{summary}");

// Clean up
await chatModel.UnloadAsync();
Console.WriteLine("\nDone. Models unloaded.");

Monit systemowy kształtuje format danych wyjściowych modelu. Poinstruując go, aby utworzyć "zorganizowane, zwięzłe notatki z punktami punktowymi", otrzymujesz ustrukturyzowaną zawartość, a nie nieprzetworzone parafrazy.

Połącz w kompletną aplikację

Zastąp zawartość Program.cs następującym kompletnym kodem, który transkrypuje plik audio i podsumowuje transkrypcję:

using Microsoft.AI.Foundry.Local;
using Betalgo.Ranul.OpenAI.ObjectModels.RequestModels;
using Microsoft.Extensions.Logging;
using System.Text;

CancellationToken ct = CancellationToken.None;

var config = new Configuration
{
    AppName = "foundry_local_samples",
    LogLevel = Microsoft.AI.Foundry.Local.LogLevel.Information
};

using var loggerFactory = LoggerFactory.Create(builder =>
{
    builder.SetMinimumLevel(
        Microsoft.Extensions.Logging.LogLevel.Information
    );
});
var logger = loggerFactory.CreateLogger<Program>();

// Initialize the singleton instance
await FoundryLocalManager.CreateAsync(config, logger);
var mgr = FoundryLocalManager.Instance;

// Download and register all execution providers.
var currentEp = "";
await mgr.DownloadAndRegisterEpsAsync((epName, percent) =>
{
    if (epName != currentEp)
    {
        if (currentEp != "") Console.WriteLine();
        currentEp = epName;
    }
    Console.Write($"\r  {epName.PadRight(30)}  {percent,6:F1}%");
});
if (currentEp != "") Console.WriteLine();

var catalog = await mgr.GetCatalogAsync();

// Load the speech-to-text model
var speechModel = await catalog.GetModelAsync("whisper-tiny")
    ?? throw new Exception("Speech model not found");

await speechModel.DownloadAsync(progress =>
{
    Console.Write($"\rDownloading speech model: {progress:F2}%");
    if (progress >= 100f) Console.WriteLine();
});

await speechModel.LoadAsync();
Console.WriteLine("Speech model loaded.");

// Transcribe the audio file
var audioClient = await speechModel.GetAudioClientAsync();
var transcriptionText = new StringBuilder();

Console.WriteLine("\nTranscription:");
var audioResponse = audioClient
    .TranscribeAudioStreamingAsync("meeting-notes.wav", ct);
await foreach (var chunk in audioResponse)
{
    Console.Write(chunk.Text);
    transcriptionText.Append(chunk.Text);
}
Console.WriteLine();

// Unload the speech model to free memory
await speechModel.UnloadAsync();

// Load the chat model for summarization
var chatModel = await catalog.GetModelAsync("qwen2.5-0.5b")
    ?? throw new Exception("Chat model not found");

await chatModel.DownloadAsync(progress =>
{
    Console.Write($"\rDownloading chat model: {progress:F2}%");
    if (progress >= 100f) Console.WriteLine();
});

await chatModel.LoadAsync();
Console.WriteLine("Chat model loaded.");

// Summarize the transcription into organized notes
var chatClient = await chatModel.GetChatClientAsync();
var messages = new List<ChatMessage>
{
    new ChatMessage
    {
        Role = "system",
        Content = "You are a note-taking assistant. Summarize " +
                  "the following transcription into organized, " +
                  "concise notes with bullet points."
    },
    new ChatMessage
    {
        Role = "user",
        Content = transcriptionText.ToString()
    }
};

var chatResponse = await chatClient.CompleteChatAsync(messages, ct);
var summary = chatResponse.Choices[0].Message.Content;
Console.WriteLine($"\nSummary:\n{summary}");

// Clean up
await chatModel.UnloadAsync();
Console.WriteLine("\nDone. Models unloaded.");

Uwaga / Notatka

Zastąp "meeting-notes.wav" ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Uruchom program do wykonywania notatek:

dotnet run

Zobaczysz dane wyjściowe podobne do:

Downloading speech model: 100.00%
Speech model loaded.

Transcription:
OK so let's get started with the weekly sync. First, the backend
API is nearly done. Sarah finished the authentication endpoints
yesterday. We still need to add rate limiting before we go to
staging. On the frontend, the dashboard redesign is about seventy
percent complete. Jake, can you walk us through the new layout?
Great. The charts look good. I think we should add a filter for
date range though. For testing, we have about eighty percent code
coverage on the API. We need to write integration tests for the
new auth flow before Friday. Let's plan to do a full regression
test next Tuesday before the release. Any blockers? OK, sounds
like we are in good shape. Let's wrap up.

Downloading chat model: 100.00%
Chat model loaded.

Summary:
- **Backend API**: Authentication endpoints complete. Rate limiting
  still needed before staging deployment.
- **Frontend**: Dashboard redesign 70% complete. New chart layout
  reviewed. Action item: add a date range filter.
- **Testing**: API code coverage at 80%. Integration tests for the
  auth flow due Friday. Full regression test scheduled for next
  Tuesday before release.
- **Status**: No blockers reported. Team is on track.

Done. Models unloaded.

Aplikacja najpierw transkrybuje zawartość audio z danymi wyjściowymi przesyłania strumieniowego, a następnie przekazuje skumulowany tekst do modelu czatu, który wyodrębnia kluczowe punkty i organizuje je w notatki ustrukturyzowane.

Instalowanie pakietów

Repozytorium przykładów

Kompletny przykładowy kod dla tego artykułu jest dostępny w repozytorium foundry-samples GitHub. Aby sklonować repozytorium i przejść do przykładowego projektu, użyj:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/javascript/foundry-local/tutorial-voice-to-text

Jeśli programujesz lub wysyłasz na Windows, wybierz kartę Windows. Pakiet Windows integruje się z środowiskiem uruchomieniowym Windows ML — zapewnia ten sam obszar powierzchni interfejsu API z szerszym zakresem przyspieszania sprzętowego.

npm install foundry-local-sdk-winml openai

Transkrypcja pliku audio

W tym kroku załadujesz model zamiany mowy na tekst i transkrybujesz plik audio. Lokalny zestaw SDK Foundry używa aliasu whisper modelu, aby wybrać najlepszy wariant Whisper dla twojego urządzenia.

  1. Utwórz plik o nazwie app.js.

  2. Dodaj następujący kod w celu zainicjowania zestawu SDK, załadowania modelu mowy i transkrypcji pliku audio:

    // Load the speech-to-text model
    const speechModel = await manager.catalog.getModel('whisper-tiny');
    await speechModel.download((progress) => {
        process.stdout.write(
            `\rDownloading speech model: ${progress.toFixed(2)}%`
        );
    });
    console.log('\nSpeech model downloaded.');
    
    await speechModel.load();
    console.log('Speech model loaded.');
    
    // Transcribe the audio file
    const audioClient = speechModel.createAudioClient();
    const transcription = await audioClient.transcribe(
        path.join(__dirname, 'meeting-notes.wav')
    );
    console.log(`\nTranscription:\n${transcription.text}`);
    
    // Unload the speech model to free memory
    await speechModel.unload();
    

    Metoda createAudioClient zwraca klienta dla operacji audio. Metoda transcribe akceptuje ścieżkę pliku i zwraca obiekt z właściwością text zawierającą treść transkrypcji.

Uwaga / Notatka

Zastąp './meeting-notes.wav' ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Podsumowanie transkrypcji

Teraz użyj modelu czatu, aby uporządkować surowe transkrypcje w uporządkowane notatki. Załaduj qwen2.5-0.5b model i wyślij transkrypcję jako kontekst z monitem systemowym, który nakazuje modelowi tworzenie czystych, podsumowanych notatek.

Dodaj następujący kod po kroku transkrypcji:

// Load the chat model for summarization
const chatModel = await manager.catalog.getModel('qwen2.5-0.5b');
await chatModel.download((progress) => {
    process.stdout.write(
        `\rDownloading chat model: ${progress.toFixed(2)}%`
    );
});
console.log('\nChat model downloaded.');

await chatModel.load();
console.log('Chat model loaded.');

// Summarize the transcription into organized notes
const chatClient = chatModel.createChatClient();
const messages = [
    {
        role: 'system',
        content: 'You are a note-taking assistant. Summarize ' +
                 'the following transcription into organized, ' +
                 'concise notes with bullet points.'
    },
    {
        role: 'user',
        content: transcription.text
    }
];

const response = await chatClient.completeChat(messages);
const summary = response.choices[0]?.message?.content;
console.log(`\nSummary:\n${summary}`);

// Clean up
await chatModel.unload();
console.log('\nDone. Models unloaded.');

Monit systemowy kształtuje format danych wyjściowych modelu. Poinstruując go, aby utworzyć "zorganizowane, zwięzłe notatki z punktami punktowymi", otrzymujesz ustrukturyzowaną zawartość, a nie nieprzetworzone parafrazy.

Połącz w kompletną aplikację

Utwórz plik o nazwie app.js i dodaj następujący pełny kod, który transkrypuje plik audio i podsumowuje transkrypcję:

import { FoundryLocalManager } from 'foundry-local-sdk';
import { fileURLToPath } from 'url';
import path from 'path';

const __dirname = path.dirname(fileURLToPath(import.meta.url));

// Initialize the Foundry Local SDK
const manager = FoundryLocalManager.create({
    appName: 'foundry_local_samples',
    logLevel: 'info'
});

// Download and register all execution providers.
let currentEp = '';
await manager.downloadAndRegisterEps((epName, percent) => {
    if (epName !== currentEp) {
        if (currentEp !== '') process.stdout.write('\n');
        currentEp = epName;
    }
    process.stdout.write(`\r  ${epName.padEnd(30)}  ${percent.toFixed(1).padStart(5)}%`);
});
if (currentEp !== '') process.stdout.write('\n');

// Load the speech-to-text model
const speechModel = await manager.catalog.getModel('whisper-tiny');
await speechModel.download((progress) => {
    process.stdout.write(
        `\rDownloading speech model: ${progress.toFixed(2)}%`
    );
});
console.log('\nSpeech model downloaded.');

await speechModel.load();
console.log('Speech model loaded.');

// Transcribe the audio file
const audioClient = speechModel.createAudioClient();
const transcription = await audioClient.transcribe(
    path.join(__dirname, 'meeting-notes.wav')
);
console.log(`\nTranscription:\n${transcription.text}`);

// Unload the speech model to free memory
await speechModel.unload();

// Load the chat model for summarization
const chatModel = await manager.catalog.getModel('qwen2.5-0.5b');
await chatModel.download((progress) => {
    process.stdout.write(
        `\rDownloading chat model: ${progress.toFixed(2)}%`
    );
});
console.log('\nChat model downloaded.');

await chatModel.load();
console.log('Chat model loaded.');

// Summarize the transcription into organized notes
const chatClient = chatModel.createChatClient();
const messages = [
    {
        role: 'system',
        content: 'You are a note-taking assistant. Summarize ' +
                 'the following transcription into organized, ' +
                 'concise notes with bullet points.'
    },
    {
        role: 'user',
        content: transcription.text
    }
];

const response = await chatClient.completeChat(messages);
const summary = response.choices[0]?.message?.content;
console.log(`\nSummary:\n${summary}`);

// Clean up
await chatModel.unload();
console.log('\nDone. Models unloaded.');

Uwaga / Notatka

Zastąp './meeting-notes.wav' ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Uruchom program do wykonywania notatek:

node app.js

Zobaczysz dane wyjściowe podobne do:

Downloading speech model: 100.00%
Speech model downloaded.
Speech model loaded.

Transcription:
OK so let's get started with the weekly sync. First, the backend
API is nearly done. Sarah finished the authentication endpoints
yesterday. We still need to add rate limiting before we go to
staging. On the frontend, the dashboard redesign is about seventy
percent complete. Jake, can you walk us through the new layout?
Great. The charts look good. I think we should add a filter for
date range though. For testing, we have about eighty percent code
coverage on the API. We need to write integration tests for the
new auth flow before Friday. Let's plan to do a full regression
test next Tuesday before the release. Any blockers? OK, sounds
like we are in good shape. Let's wrap up.

Downloading chat model: 100.00%
Chat model downloaded.
Chat model loaded.

Summary:
- **Backend API**: Authentication endpoints complete. Rate limiting
  still needed before staging deployment.
- **Frontend**: Dashboard redesign 70% complete. New chart layout
  reviewed. Action item: add a date range filter.
- **Testing**: API code coverage at 80%. Integration tests for the
  auth flow due Friday. Full regression test scheduled for next
  Tuesday before release.
- **Status**: No blockers reported. Team is on track.

Done. Models unloaded.

Aplikacja najpierw transkrybuje zawartość audio, a następnie przekazuje ten tekst do modelu czatu, który wyodrębnia kluczowe punkty i organizuje je w notatki ustrukturyzowane.

Instalowanie pakietów

Repozytorium przykładów

Kompletny przykładowy kod dla tego artykułu jest dostępny w repozytorium foundry-samples GitHub. Aby sklonować repozytorium i przejść do przykładowego projektu, użyj:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/python/foundry-local/tutorial-voice-to-text

Jeśli programujesz lub wysyłasz na Windows, wybierz kartę Windows. Pakiet Windows integruje się z środowiskiem uruchomieniowym Windows ML — zapewnia ten sam obszar powierzchni interfejsu API z szerszym zakresem przyspieszania sprzętowego.

pip install foundry-local-sdk-winml openai

Transkrypcja pliku audio

W tym kroku załadujesz model zamiany mowy na tekst i transkrybujesz plik audio. Lokalny zestaw SDK Foundry używa aliasu whisper modelu, aby wybrać najlepszy wariant Whisper dla twojego urządzenia.

  1. Utwórz plik o nazwie app.py.

  2. Dodaj następujący kod w celu zainicjowania zestawu SDK, załadowania modelu mowy i transkrypcji pliku audio:

    # Load the speech-to-text model
    speech_model = manager.catalog.get_model("whisper-tiny")
    speech_model.download(
        lambda progress: print(
            f"\rDownloading speech model: {progress:.2f}%",
            end="",
            flush=True,
        )
    )
    print()
    speech_model.load()
    print("Speech model loaded.")
    
    # Transcribe the audio file
    audio_client = speech_model.get_audio_client()
    transcription = audio_client.transcribe("meeting-notes.wav")
    print(f"\nTranscription:\n{transcription.text}")
    
    # Unload the speech model to free memory
    speech_model.unload()
    

    Metoda get_audio_client zwraca klienta dla operacji audio. Metoda transcribe akceptuje ścieżkę pliku i zwraca obiekt z właściwością text zawierającą treść transkrypcji.

Uwaga / Notatka

Zastąp "meeting-notes.wav" ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Podsumowanie transkrypcji

Teraz użyj modelu czatu, aby uporządkować surowe transkrypcje w uporządkowane notatki. Załaduj qwen2.5-0.5b model i wyślij transkrypcję jako kontekst z monitem systemowym, który nakazuje modelowi tworzenie czystych, podsumowanych notatek.

Dodaj następujący kod po kroku transkrypcji:

# Load the chat model for summarization
chat_model = manager.catalog.get_model("qwen2.5-0.5b")
chat_model.download(
    lambda progress: print(
        f"\rDownloading chat model: {progress:.2f}%",
        end="",
        flush=True,
    )
)
print()
chat_model.load()
print("Chat model loaded.")

# Summarize the transcription into organized notes
client = chat_model.get_chat_client()
messages = [
    {
        "role": "system",
        "content": "You are a note-taking assistant. "
        "Summarize the following transcription "
        "into organized, concise notes with "
        "bullet points.",
    },
    {"role": "user", "content": transcription.text},
]

response = client.complete_chat(messages)
summary = response.choices[0].message.content
print(f"\nSummary:\n{summary}")

# Clean up
chat_model.unload()
print("\nDone. Models unloaded.")

Monit systemowy kształtuje format danych wyjściowych modelu. Poinstruując go, aby utworzyć "zorganizowane, zwięzłe notatki z punktami punktowymi", otrzymujesz ustrukturyzowaną zawartość, a nie nieprzetworzone parafrazy.

Połącz w kompletną aplikację

Utwórz plik o nazwie app.py i dodaj następujący pełny kod, który transkrypuje plik audio i podsumowuje transkrypcję:

from foundry_local_sdk import Configuration, FoundryLocalManager



def main():
    # Initialize the Foundry Local SDK
    config = Configuration(app_name="foundry_local_samples")
    FoundryLocalManager.initialize(config)
    manager = FoundryLocalManager.instance

    # Download and register all execution providers.
    current_ep = ""

    def ep_progress(ep_name: str, percent: float):
        nonlocal current_ep
        if ep_name != current_ep:
            if current_ep:
                print()
            current_ep = ep_name
        print(f"\r  {ep_name:<30}  {percent:5.1f}%", end="", flush=True)

    manager.download_and_register_eps(progress_callback=ep_progress)
    if current_ep:
        print()

    # Load the speech-to-text model
    speech_model = manager.catalog.get_model("whisper-tiny")
    speech_model.download(
        lambda progress: print(
            f"\rDownloading speech model: {progress:.2f}%",
            end="",
            flush=True,
        )
    )
    print()
    speech_model.load()
    print("Speech model loaded.")

    # Transcribe the audio file
    audio_client = speech_model.get_audio_client()
    transcription = audio_client.transcribe("meeting-notes.wav")
    print(f"\nTranscription:\n{transcription.text}")

    # Unload the speech model to free memory
    speech_model.unload()

    # Load the chat model for summarization
    chat_model = manager.catalog.get_model("qwen2.5-0.5b")
    chat_model.download(
        lambda progress: print(
            f"\rDownloading chat model: {progress:.2f}%",
            end="",
            flush=True,
        )
    )
    print()
    chat_model.load()
    print("Chat model loaded.")

    # Summarize the transcription into organized notes
    client = chat_model.get_chat_client()
    messages = [
        {
            "role": "system",
            "content": "You are a note-taking assistant. "
            "Summarize the following transcription "
            "into organized, concise notes with "
            "bullet points.",
        },
        {"role": "user", "content": transcription.text},
    ]

    response = client.complete_chat(messages)
    summary = response.choices[0].message.content
    print(f"\nSummary:\n{summary}")

    # Clean up
    chat_model.unload()
    print("\nDone. Models unloaded.")


if __name__ == "__main__":
    main()

Uwaga / Notatka

Zastąp "meeting-notes.wav" ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Uruchom program do wykonywania notatek:

python app.py

Zobaczysz dane wyjściowe podobne do:

Downloading speech model: 100.00%
Speech model loaded.

Transcription:
OK so let's get started with the weekly sync. First, the backend
API is nearly done. Sarah finished the authentication endpoints
yesterday. We still need to add rate limiting before we go to
staging. On the frontend, the dashboard redesign is about seventy
percent complete. Jake, can you walk us through the new layout?
Great. The charts look good. I think we should add a filter for
date range though. For testing, we have about eighty percent code
coverage on the API. We need to write integration tests for the
new auth flow before Friday. Let's plan to do a full regression
test next Tuesday before the release. Any blockers? OK, sounds
like we are in good shape. Let's wrap up.

Downloading chat model: 100.00%
Chat model loaded.

Summary:
- **Backend API**: Authentication endpoints complete. Rate limiting
  still needed before staging deployment.
- **Frontend**: Dashboard redesign 70% complete. New chart layout
  reviewed. Action item: add a date range filter.
- **Testing**: API code coverage at 80%. Integration tests for the
  auth flow due Friday. Full regression test scheduled for next
  Tuesday before release.
- **Status**: No blockers reported. Team is on track.

Done. Models unloaded.

Aplikacja najpierw transkrybuje zawartość audio, a następnie przekazuje ten tekst do modelu czatu, który wyodrębnia kluczowe punkty i organizuje je w notatki ustrukturyzowane.

Instalowanie pakietów

Repozytorium przykładów

Kompletny przykładowy kod dla tego artykułu jest dostępny w repozytorium foundry-samples GitHub. Aby sklonować repozytorium i przejść do przykładowego projektu, użyj:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/rust/foundry-local/tutorial-voice-to-text

Jeśli programujesz lub wysyłasz na Windows, wybierz kartę Windows. Pakiet Windows integruje się z środowiskiem uruchomieniowym Windows ML — zapewnia ten sam obszar powierzchni interfejsu API z szerszym zakresem przyspieszania sprzętowego.

cargo add foundry-local-sdk --features winml
cargo add tokio --features full
cargo add tokio-stream anyhow

Transkrypcja pliku audio

W tym kroku załadujesz model zamiany mowy na tekst i transkrybujesz plik audio. Lokalny zestaw SDK Foundry używa aliasu whisper modelu, aby wybrać najlepszy wariant Whisper dla twojego urządzenia.

  • Otwórz src/main.rs i zastąp jego zawartość następującym kodem, aby zainicjować zestaw SDK, załadować model mowy i transkrypować plik audio:

    // Load the speech-to-text model
    let speech_model = manager
        .catalog()
        .get_model("whisper-tiny")
        .await?;
    
    if !speech_model.is_cached().await? {
        println!("Downloading speech model...");
        speech_model
            .download(Some(|progress: f64| {
                print!("\r  {progress:.1}%");
                io::stdout().flush().ok();
            }))
            .await?;
        println!();
    }
    
    speech_model.load().await?;
    println!("Speech model loaded.");
    
    // Transcribe the audio file
    let audio_client = speech_model.create_audio_client();
    let transcription = audio_client
        .transcribe("meeting-notes.wav")
        .await?;
    println!("\nTranscription:\n{}", transcription.text);
    
    // Unload the speech model to free memory
    speech_model.unload().await?;
    

    Metoda create_audio_client zwraca klienta dla operacji audio. Metoda transcribe akceptuje ścieżkę pliku i zwraca obiekt z polem text zawierającym zawartość transkrypcji.

Uwaga / Notatka

Zastąp "meeting-notes.wav" ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Podsumowanie transkrypcji

Teraz użyj modelu czatu, aby uporządkować surowe transkrypcje w uporządkowane notatki. Załaduj qwen2.5-0.5b model i wyślij transkrypcję jako kontekst z monitem systemowym, który nakazuje modelowi tworzenie czystych, podsumowanych notatek.

Dodaj następujący kod wewnątrz funkcji main po kroku transkrypcji:

// Load the chat model for summarization
let chat_model = manager
    .catalog()
    .get_model("qwen2.5-0.5b")
    .await?;

if !chat_model.is_cached().await? {
    println!("Downloading chat model...");
    chat_model
        .download(Some(|progress: f64| {
            print!("\r  {progress:.1}%");
            io::stdout().flush().ok();
        }))
        .await?;
    println!();
}

chat_model.load().await?;
println!("Chat model loaded.");

// Summarize the transcription into organized notes
let client = chat_model
    .create_chat_client()
    .temperature(0.7)
    .max_tokens(512);

let messages: Vec<ChatCompletionRequestMessage> = vec![
    ChatCompletionRequestSystemMessage::from(
        "You are a note-taking assistant. Summarize \
         the following transcription into organized, \
         concise notes with bullet points.",
    )
    .into(),
    ChatCompletionRequestUserMessage::from(
        transcription.text.as_str(),
    )
    .into(),
];

let response = client
    .complete_chat(&messages, None)
    .await?;
let summary = response.choices[0]
    .message
    .content
    .as_deref()
    .unwrap_or("");
println!("\nSummary:\n{}", summary);

// Clean up
chat_model.unload().await?;
println!("\nDone. Models unloaded.");

Monit systemowy kształtuje format danych wyjściowych modelu. Poinstruując go, aby utworzyć "zorganizowane, zwięzłe notatki z punktami punktowymi", otrzymujesz ustrukturyzowaną zawartość, a nie nieprzetworzone parafrazy.

Połącz w kompletną aplikację

Zastąp zawartość src/main.rs następującym kompletnym kodem, który transkrypuje plik audio i podsumowuje transkrypcję:

use foundry_local_sdk::{
    ChatCompletionRequestMessage,
    ChatCompletionRequestSystemMessage,
    ChatCompletionRequestUserMessage,
    FoundryLocalConfig, FoundryLocalManager,
};
use std::io::{self, Write};

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Initialize the Foundry Local SDK
    let manager = FoundryLocalManager::create(
        FoundryLocalConfig::new("foundry_local_samples"),
    )?;

    // Download and register all execution providers.
    manager
        .download_and_register_eps_with_progress(None, {
            let mut current_ep = String::new();
            move |ep_name: &str, percent: f64| {
                if ep_name != current_ep {
                    if !current_ep.is_empty() {
                        println!();
                    }
                    current_ep = ep_name.to_string();
                }
                print!("\r  {:<30}  {:5.1}%", ep_name, percent);
                io::stdout().flush().ok();
            }
        })
        .await?;
    println!();

    // Load the speech-to-text model
    let speech_model = manager
        .catalog()
        .get_model("whisper-tiny")
        .await?;

    if !speech_model.is_cached().await? {
        println!("Downloading speech model...");
        speech_model
            .download(Some(|progress: f64| {
                print!("\r  {progress:.1}%");
                io::stdout().flush().ok();
            }))
            .await?;
        println!();
    }

    speech_model.load().await?;
    println!("Speech model loaded.");

    // Transcribe the audio file
    let audio_client = speech_model.create_audio_client();
    let transcription = audio_client
        .transcribe("meeting-notes.wav")
        .await?;
    println!("\nTranscription:\n{}", transcription.text);

    // Unload the speech model to free memory
    speech_model.unload().await?;

    // Load the chat model for summarization
    let chat_model = manager
        .catalog()
        .get_model("qwen2.5-0.5b")
        .await?;

    if !chat_model.is_cached().await? {
        println!("Downloading chat model...");
        chat_model
            .download(Some(|progress: f64| {
                print!("\r  {progress:.1}%");
                io::stdout().flush().ok();
            }))
            .await?;
        println!();
    }

    chat_model.load().await?;
    println!("Chat model loaded.");

    // Summarize the transcription into organized notes
    let client = chat_model
        .create_chat_client()
        .temperature(0.7)
        .max_tokens(512);

    let messages: Vec<ChatCompletionRequestMessage> = vec![
        ChatCompletionRequestSystemMessage::from(
            "You are a note-taking assistant. Summarize \
             the following transcription into organized, \
             concise notes with bullet points.",
        )
        .into(),
        ChatCompletionRequestUserMessage::from(
            transcription.text.as_str(),
        )
        .into(),
    ];

    let response = client
        .complete_chat(&messages, None)
        .await?;
    let summary = response.choices[0]
        .message
        .content
        .as_deref()
        .unwrap_or("");
    println!("\nSummary:\n{}", summary);

    // Clean up
    chat_model.unload().await?;
    println!("\nDone. Models unloaded.");

    Ok(())
}

Uwaga / Notatka

Zastąp "meeting-notes.wav" ścieżką do pliku audio. Obsługiwane formaty to WAV, MP3 i FLAC.

Uruchom program do wykonywania notatek:

cargo run

Zobaczysz dane wyjściowe podobne do:

Downloading speech model: 100.00%
Speech model loaded.

Transcription:
OK so let's get started with the weekly sync. First, the backend
API is nearly done. Sarah finished the authentication endpoints
yesterday. We still need to add rate limiting before we go to
staging. On the frontend, the dashboard redesign is about seventy
percent complete. Jake, can you walk us through the new layout?
Great. The charts look good. I think we should add a filter for
date range though. For testing, we have about eighty percent code
coverage on the API. We need to write integration tests for the
new auth flow before Friday. Let's plan to do a full regression
test next Tuesday before the release. Any blockers? OK, sounds
like we are in good shape. Let's wrap up.

Downloading chat model: 100.00%
Chat model loaded.

Summary:
- **Backend API**: Authentication endpoints complete. Rate limiting
  still needed before staging deployment.
- **Frontend**: Dashboard redesign 70% complete. New chart layout
  reviewed. Action item: add a date range filter.
- **Testing**: API code coverage at 80%. Integration tests for the
  auth flow due Friday. Full regression test scheduled for next
  Tuesday before release.
- **Status**: No blockers reported. Team is on track.

Done. Models unloaded.

Aplikacja najpierw transkrybuje zawartość audio, a następnie przekazuje ten tekst do modelu czatu, który wyodrębnia kluczowe punkty i organizuje je w notatki ustrukturyzowane.

Uprzątnij zasoby

Wagi modelu są przechowywane w Twojej lokalnej pamięci podręcznej po rozładowaniu modelu. Oznacza to, że przy następnym uruchomieniu aplikacji krok pobierania zostanie pominięty, a model ładuje się szybciej. Nie jest wymagane żadne dodatkowe czyszczenie, chyba że chcesz odzyskać miejsce na dysku.