Öğretici: Foundry Local ile çok aşamalı sohbet yardımcısı oluşturma

Bu öğreticide, tamamen cihazınızda çalışan etkileşimli bir sohbet yardımcısı oluşturacaksınız. Yardımcı, konuşma sırasında birden fazla etkileşimdeki bağlamı korur, yani konuşmanın önceki bölümlerinde ne tartıştığınızı hatırlar. Bir model seçmek, sistem istemi tanımlamak ve yanıt belirtecini belirteçle akışla aktarmak için Foundry Yerel SDK'sını kullanırsınız.

Bu eğitimde şunları öğreniyorsunuz:

  • Bir proje ayarlama ve Foundry Yerel SDK'sını yükleme
  • Model kataloğuna göz atın ve bir model seçin
  • Asistanın davranışını belirlemek için sistem istemi tanımlama
  • Mesaj geçmişiyle çok aşamalı konuşmayı uygulama
  • Yanıt veren bir deneyim için yanıtları akışla aktar
  • Konuşma sona erdiğinde kaynakları temizleme

Önkoşullar

  • En az 8 GB RAM'e sahip bir Windows, macOS veya Linux bilgisayar.

Örnekler deposu

Bu makalenin örnek kodunun tamamını Dökümhane örnekleri GitHub deposunda bulabilirsiniz. Depoyu kopyalamak ve örneğe gitmek için şunu kullanın:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/csharp/foundry-local/tutorial-chat-assistant

Paketleri yükleme

Windows üzerinde geliştiriyor veya gönderiyorsanız Windows sekmesini seçin. Windows paketi Windows ML çalışma zamanıyla tümleşir; aynı API yüzey alanını daha geniş bir donanım hızlandırmasıyla sağlar.

dotnet add package Microsoft.AI.Foundry.Local.WinML
dotnet add package OpenAI

GitHub deposundaki C# örnekleri önceden yapılandırılmış projelerdir. Sıfırdan oluşturuyorsanız, C# projenizi Foundry Local ile nasıl ayarlayacağınız hakkında daha fazla ayrıntı için Foundry Yerel SDK başvurusunu okumalısınız.

Kataloğa göz atın ve bir model seçin

Foundry Yerel SDK'sı tüm kullanılabilir modelleri listeleyen bir model kataloğu sağlar. Bu adımda SDK'yı başlatır ve sohbet yardımcınız için bir model seçersiniz.

  • SDK'yi başlatmak ve bir model seçmek için içeriğini aşağıdaki kodla açın Program.cs ve değiştirin:

    CancellationToken ct = CancellationToken.None;
    
    var config = new Configuration
    {
        AppName = "foundry_local_samples",
        LogLevel = Microsoft.AI.Foundry.Local.LogLevel.Information
    };
    
    using var loggerFactory = LoggerFactory.Create(builder =>
    {
        builder.SetMinimumLevel(Microsoft.Extensions.Logging.LogLevel.Information);
    });
    var logger = loggerFactory.CreateLogger<Program>();
    
    // Initialize the singleton instance
    await FoundryLocalManager.CreateAsync(config, logger);
    var mgr = FoundryLocalManager.Instance;
    
    // Download and register all execution providers.
    var currentEp = "";
    await mgr.DownloadAndRegisterEpsAsync((epName, percent) =>
    {
        if (epName != currentEp)
        {
            if (currentEp != "") Console.WriteLine();
            currentEp = epName;
        }
        Console.Write($"\r  {epName.PadRight(30)}  {percent,6:F1}%");
    });
    if (currentEp != "") Console.WriteLine();
    
    // Select and load a model from the catalog
    var catalog = await mgr.GetCatalogAsync();
    var model = await catalog.GetModelAsync("qwen2.5-0.5b")
        ?? throw new Exception("Model not found");
    
    await model.DownloadAsync(progress =>
    {
        Console.Write($"\rDownloading model: {progress:F2}%");
        if (progress >= 100f) Console.WriteLine();
    });
    
    await model.LoadAsync();
    Console.WriteLine("Model loaded and ready.");
    
    // Get a chat client
    var chatClient = await model.GetChatClientAsync();
    

    Bu GetModelAsync yöntemi, katalogdaki belirli bir modelle eşleşen kısa ve kolay bir ad olan model takma adını kabul eder. DownloadAsync yöntemi, modelin ağırlıklarını yerel önbelleğinize getirir ve LoadAsync modeli çıkarım için hazır hale getirir.

Sistem istemi tanımlama

Sistem istemi, yardımcının kişiliğini ve davranışını ayarlar. Konuşma geçmişindeki ilk iletidir ve model konuşma boyunca buna başvurur.

Yardımcının nasıl yanıt vereceğini şekillendirmek için bir sistem istemi ekleyin:

// Start the conversation with a system prompt
var messages = new List<ChatMessage>
{
    new ChatMessage
    {
        Role = "system",
        Content = "You are a helpful, friendly assistant. Keep your responses " +
                  "concise and conversational. If you don't know something, say so."
    }
};

Tavsiye

Yardımcının davranışını değiştirmek için farklı sistem istemleriyle denemeler yapın. Örneğin, korsan, öğretmen veya etki alanı uzmanı olarak yanıt vermesini sağlayabilirsiniz.

Çok aşamalı konuşmayı uygula

Bir sohbet asistanının birden çok etkileşimde bağlamı koruması gerekir. Bunu, tüm iletilerin (sistem, kullanıcı ve yardımcı) listesini tutarak ve her istekle birlikte tam listeyi göndererek elde edebilirsiniz. Model bağlamsal olarak ilgili yanıtlar oluşturmak için bu geçmişi kullanır.

Şu şekilde bir konuşma döngüsü ekleyin:

  • Konsoldan kullanıcı girişini okur.
  • Kullanıcı iletisini geçmişe ekler.
  • Tüm geçmişi modele gönderir.
  • Bir sonraki adım için yardımcının yanıtını geçmişe ekler.
while (true)
{
    Console.Write("You: ");
    var userInput = Console.ReadLine();
    if (string.IsNullOrWhiteSpace(userInput) ||
        userInput.Equals("quit", StringComparison.OrdinalIgnoreCase) ||
        userInput.Equals("exit", StringComparison.OrdinalIgnoreCase))
    {
        break;
    }

    // Add the user's message to conversation history
    messages.Add(new ChatMessage { Role = "user", Content = userInput });

    // Stream the response token by token
    Console.Write("Assistant: ");
    var fullResponse = string.Empty;
    var streamingResponse = chatClient.CompleteChatStreamingAsync(messages, ct);
    await foreach (var chunk in streamingResponse)
    {
        var content = chunk.Choices[0].Message.Content;
        if (!string.IsNullOrEmpty(content))
        {
            Console.Write(content);
            Console.Out.Flush();
            fullResponse += content;
        }
    }
    Console.WriteLine("\n");

    // Add the complete response to conversation history
    messages.Add(new ChatMessage { Role = "assistant", Content = fullResponse });
}

için yapılan her çağrı CompleteChatAsync , ileti geçmişinin tamamını alır. Modelin önceki dönüşleri "anımsaması" budur; çağrılar arasında durum depolamaz.

Akış yanıtları ekleme

Akış, oluşturulduktan sonra her belirteci yazdırır ve bu da yardımcının daha hızlı yanıt vermesine neden olur. Çağrıyı CompleteChatAsync ile CompleteChatStreamingAsync çağrısı ile değiştirerek yanıtı belirteç bazında akışa alın.

Akış kullanmak için konuşma döngüsünü güncelleştirin:

// Stream the response token by token
Console.Write("Assistant: ");
var fullResponse = string.Empty;
var streamingResponse = chatClient.CompleteChatStreamingAsync(messages, ct);
await foreach (var chunk in streamingResponse)
{
    var content = chunk.Choices[0].Message.Content;
    if (!string.IsNullOrEmpty(content))
    {
        Console.Write(content);
        Console.Out.Flush();
        fullResponse += content;
    }
}
Console.WriteLine("\n");

Akış sürümü, akış tamamlandıktan sonra konuşma geçmişine eklenebilmesi için tam yanıtı biriktirir.

Kodu tamamlama

Program.cs öğesinin içeriğini aşağıdaki tam kod ile değiştirin:

using Microsoft.AI.Foundry.Local;
using Betalgo.Ranul.OpenAI.ObjectModels.RequestModels;
using Microsoft.Extensions.Logging;

CancellationToken ct = CancellationToken.None;

var config = new Configuration
{
    AppName = "foundry_local_samples",
    LogLevel = Microsoft.AI.Foundry.Local.LogLevel.Information
};

using var loggerFactory = LoggerFactory.Create(builder =>
{
    builder.SetMinimumLevel(Microsoft.Extensions.Logging.LogLevel.Information);
});
var logger = loggerFactory.CreateLogger<Program>();

// Initialize the singleton instance
await FoundryLocalManager.CreateAsync(config, logger);
var mgr = FoundryLocalManager.Instance;

// Download and register all execution providers.
var currentEp = "";
await mgr.DownloadAndRegisterEpsAsync((epName, percent) =>
{
    if (epName != currentEp)
    {
        if (currentEp != "") Console.WriteLine();
        currentEp = epName;
    }
    Console.Write($"\r  {epName.PadRight(30)}  {percent,6:F1}%");
});
if (currentEp != "") Console.WriteLine();

// Select and load a model from the catalog
var catalog = await mgr.GetCatalogAsync();
var model = await catalog.GetModelAsync("qwen2.5-0.5b")
    ?? throw new Exception("Model not found");

await model.DownloadAsync(progress =>
{
    Console.Write($"\rDownloading model: {progress:F2}%");
    if (progress >= 100f) Console.WriteLine();
});

await model.LoadAsync();
Console.WriteLine("Model loaded and ready.");

// Get a chat client
var chatClient = await model.GetChatClientAsync();

// Start the conversation with a system prompt
var messages = new List<ChatMessage>
{
    new ChatMessage
    {
        Role = "system",
        Content = "You are a helpful, friendly assistant. Keep your responses " +
                  "concise and conversational. If you don't know something, say so."
    }
};

Console.WriteLine("\nChat assistant ready! Type 'quit' to exit.\n");

while (true)
{
    Console.Write("You: ");
    var userInput = Console.ReadLine();
    if (string.IsNullOrWhiteSpace(userInput) ||
        userInput.Equals("quit", StringComparison.OrdinalIgnoreCase) ||
        userInput.Equals("exit", StringComparison.OrdinalIgnoreCase))
    {
        break;
    }

    // Add the user's message to conversation history
    messages.Add(new ChatMessage { Role = "user", Content = userInput });

    // Stream the response token by token
    Console.Write("Assistant: ");
    var fullResponse = string.Empty;
    var streamingResponse = chatClient.CompleteChatStreamingAsync(messages, ct);
    await foreach (var chunk in streamingResponse)
    {
        var content = chunk.Choices[0].Message.Content;
        if (!string.IsNullOrEmpty(content))
        {
            Console.Write(content);
            Console.Out.Flush();
            fullResponse += content;
        }
    }
    Console.WriteLine("\n");

    // Add the complete response to conversation history
    messages.Add(new ChatMessage { Role = "assistant", Content = fullResponse });
}

// Clean up - unload the model
await model.UnloadAsync();
Console.WriteLine("Model unloaded. Goodbye!");

Sohbet yardımcısını çalıştırın:

dotnet run

Şuna benzer bir çıkış görürsünüz:

Downloading model: 100.00%
Model loaded and ready.

Chat assistant ready! Type 'quit' to exit.

You: What is photosynthesis?
Assistant: Photosynthesis is the process plants use to convert sunlight, water, and carbon
dioxide into glucose and oxygen. It mainly happens in the leaves, inside structures
called chloroplasts.

You: Why is it important for other living things?
Assistant: It's essential because photosynthesis produces the oxygen that most living things
breathe. It also forms the base of the food chain — animals eat plants or eat other
animals that depend on plants for energy.

You: quit
Model unloaded. Goodbye!

Asistanın bağlamı önceki dönüşlerden nasıl hatırladığını fark edin. "Diğer canlılar için neden önemlidir?" sorusunu sorduğunuzda, hala fotosentezden bahsettiğinizi bilir.

Örnekler deposu

Bu makalenin örnek kodunun tamamına foundry-samples GitHub deposundan ulaşabilirsiniz. Depoyu klonlamak ve örneğe gitmek için şu adımları kullanın:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/javascript/foundry-local/tutorial-chat-assistant

Paketleri yükleme

Windows üzerinde geliştiriyor veya gönderiyorsanız Windows sekmesini seçin. Windows paketi Windows ML çalışma zamanıyla tümleşir; aynı API yüzey alanını daha geniş bir donanım hızlandırmasıyla sağlar.

npm install foundry-local-sdk-winml openai

Kataloğa göz atın ve bir model seçin

Foundry Yerel SDK'sı tüm kullanılabilir modelleri listeleyen bir model kataloğu sağlar. Bu adımda SDK'yı başlatır ve sohbet yardımcınız için bir model seçersiniz.

  1. adlı index.jsbir dosya oluşturun.

  2. SDK'yi başlatmak için aşağıdaki kodu ekleyin ve bir model seçin:

    // Initialize the Foundry Local SDK
    const manager = FoundryLocalManager.create({
        appName: 'foundry_local_samples',
        logLevel: 'info'
    });
    
    // Download and register all execution providers.
    let currentEp = '';
    await manager.downloadAndRegisterEps((epName, percent) => {
        if (epName !== currentEp) {
            if (currentEp !== '') process.stdout.write('\n');
            currentEp = epName;
        }
        process.stdout.write(`\r  ${epName.padEnd(30)}  ${percent.toFixed(1).padStart(5)}%`);
    });
    if (currentEp !== '') process.stdout.write('\n');
    
    // Select and load a model from the catalog
    const model = await manager.catalog.getModel('qwen2.5-0.5b');
    
    await model.download((progress) => {
        process.stdout.write(`\rDownloading model: ${progress.toFixed(2)}%`);
    });
    console.log('\nModel downloaded.');
    
    await model.load();
    console.log('Model loaded and ready.');
    
    // Create a chat client
    const chatClient = model.createChatClient();
    

    Bu getModel yöntemi, katalogdaki belirli bir modelle eşleşen kısa ve kolay bir ad olan model takma adını kabul eder. download yöntemi, modelin ağırlıklarını yerel önbelleğinize getirir ve load modeli çıkarım için hazır hale getirir.

Sistem istemi tanımlama

Sistem istemi, yardımcının kişiliğini ve davranışını ayarlar. Konuşma geçmişindeki ilk iletidir ve model konuşma boyunca buna başvurur.

Yardımcının nasıl yanıt vereceğini şekillendirmek için bir sistem istemi ekleyin:

// Start the conversation with a system prompt
const messages = [
    {
        role: 'system',
        content: 'You are a helpful, friendly assistant. Keep your responses ' +
                 'concise and conversational. If you don\'t know something, say so.'
    }
];

Tavsiye

Yardımcının davranışını değiştirmek için farklı sistem istemleriyle denemeler yapın. Örneğin, korsan, öğretmen veya etki alanı uzmanı olarak yanıt vermesini sağlayabilirsiniz.

Çok aşamalı konuşmayı uygula

Bir sohbet asistanının birden çok etkileşimde bağlamı koruması gerekir. Bunu, tüm iletilerin (sistem, kullanıcı ve yardımcı) listesini tutarak ve her istekle birlikte tam listeyi göndererek elde edebilirsiniz. Model bağlamsal olarak ilgili yanıtlar oluşturmak için bu geçmişi kullanır.

Şu şekilde bir konuşma döngüsü ekleyin:

  • Konsoldan kullanıcı girişini okur.
  • Kullanıcı iletisini geçmişe ekler.
  • Tüm geçmişi modele gönderir.
  • Bir sonraki adım için yardımcının yanıtını geçmişe ekler.
while (true) {
    const userInput = await askQuestion('You: ');
    if (userInput.trim().toLowerCase() === 'quit' ||
        userInput.trim().toLowerCase() === 'exit') {
        break;
    }

    // Add the user's message to conversation history
    messages.push({ role: 'user', content: userInput });

    // Stream the response token by token
    process.stdout.write('Assistant: ');
    let fullResponse = '';
    for await (const chunk of chatClient.completeStreamingChat(messages)) {
        const content = chunk.choices?.[0]?.delta?.content;
        if (content) {
            process.stdout.write(content);
            fullResponse += content;
        }
    }
    console.log('\n');

    // Add the complete response to conversation history
    messages.push({ role: 'assistant', content: fullResponse });
}

için yapılan her çağrı completeChat , ileti geçmişinin tamamını alır. Modelin önceki dönüşleri "anımsaması" budur; çağrılar arasında durum depolamaz.

Akış yanıtları ekleme

Akış, oluşturulduktan sonra her belirteci yazdırır ve bu da yardımcının daha hızlı yanıt vermesine neden olur. Çağrıyı completeChat ile completeStreamingChat çağrısı ile değiştirerek yanıtı belirteç bazında akışa alın.

Akış kullanmak için konuşma döngüsünü güncelleştirin:

// Stream the response token by token
process.stdout.write('Assistant: ');
let fullResponse = '';
for await (const chunk of chatClient.completeStreamingChat(messages)) {
    const content = chunk.choices?.[0]?.delta?.content;
    if (content) {
        process.stdout.write(content);
        fullResponse += content;
    }
}
console.log('\n');

Akış sürümü, akış tamamlandıktan sonra konuşma geçmişine eklenebilmesi için tam yanıtı biriktirir.

Kodu tamamlama

adlı index.js bir dosya oluşturun ve aşağıdaki tam kodu ekleyin:

import { FoundryLocalManager } from 'foundry-local-sdk';
import * as readline from 'readline';

// Initialize the Foundry Local SDK
const manager = FoundryLocalManager.create({
    appName: 'foundry_local_samples',
    logLevel: 'info'
});

// Download and register all execution providers.
let currentEp = '';
await manager.downloadAndRegisterEps((epName, percent) => {
    if (epName !== currentEp) {
        if (currentEp !== '') process.stdout.write('\n');
        currentEp = epName;
    }
    process.stdout.write(`\r  ${epName.padEnd(30)}  ${percent.toFixed(1).padStart(5)}%`);
});
if (currentEp !== '') process.stdout.write('\n');

// Select and load a model from the catalog
const model = await manager.catalog.getModel('qwen2.5-0.5b');

await model.download((progress) => {
    process.stdout.write(`\rDownloading model: ${progress.toFixed(2)}%`);
});
console.log('\nModel downloaded.');

await model.load();
console.log('Model loaded and ready.');

// Create a chat client
const chatClient = model.createChatClient();

// Start the conversation with a system prompt
const messages = [
    {
        role: 'system',
        content: 'You are a helpful, friendly assistant. Keep your responses ' +
                 'concise and conversational. If you don\'t know something, say so.'
    }
];

// Set up readline for console input
const rl = readline.createInterface({
    input: process.stdin,
    output: process.stdout
});

const askQuestion = (prompt) => new Promise((resolve) => {
    if (rl.closed) return resolve('quit');
    const onClose = () => resolve('quit');
    rl.once('close', onClose);
    try {
        rl.question(prompt, (answer) => {
            rl.off('close', onClose);
            resolve(answer);
        });
    } catch {
        rl.off('close', onClose);
        resolve('quit');
    }
});

console.log('\nChat assistant ready! Type \'quit\' to exit.\n');

while (true) {
    const userInput = await askQuestion('You: ');
    if (userInput.trim().toLowerCase() === 'quit' ||
        userInput.trim().toLowerCase() === 'exit') {
        break;
    }

    // Add the user's message to conversation history
    messages.push({ role: 'user', content: userInput });

    // Stream the response token by token
    process.stdout.write('Assistant: ');
    let fullResponse = '';
    for await (const chunk of chatClient.completeStreamingChat(messages)) {
        const content = chunk.choices?.[0]?.delta?.content;
        if (content) {
            process.stdout.write(content);
            fullResponse += content;
        }
    }
    console.log('\n');

    // Add the complete response to conversation history
    messages.push({ role: 'assistant', content: fullResponse });
}

// Clean up - unload the model
await model.unload();
console.log('Model unloaded. Goodbye!');
rl.close();

Sohbet yardımcısını çalıştırın:

node index.js

Şuna benzer bir çıkış görürsünüz:

Downloading model: 100.00%
Model downloaded.
Model loaded and ready.

Chat assistant ready! Type 'quit' to exit.

You: What is photosynthesis?
Assistant: Photosynthesis is the process plants use to convert sunlight, water, and carbon
dioxide into glucose and oxygen. It mainly happens in the leaves, inside structures
called chloroplasts.

You: Why is it important for other living things?
Assistant: It's essential because photosynthesis produces the oxygen that most living things
breathe. It also forms the base of the food chain — animals eat plants or eat other
animals that depend on plants for energy.

You: quit
Model unloaded. Goodbye!

Asistanın bağlamı önceki dönüşlerden nasıl hatırladığını fark edin. "Diğer canlılar için neden önemlidir?" sorusunu sorduğunuzda, hala fotosentezden bahsettiğinizi bilir.

Örnekler deposu

Bu makalenin örnek kodunun tamamına foundry-samples GitHub deposundan ulaşabilirsiniz. Depoyu klonlamak ve örneğe gitmek için şu adımları kullanın:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/python/foundry-local/tutorial-chat-assistant

Paketleri yükleme

Windows üzerinde geliştiriyor veya gönderiyorsanız Windows sekmesini seçin. Windows paketi Windows ML çalışma zamanıyla tümleşir; aynı API yüzey alanını daha geniş bir donanım hızlandırmasıyla sağlar.

pip install foundry-local-sdk-winml openai

Kataloğa göz atın ve bir model seçin

Foundry Yerel SDK'sı tüm kullanılabilir modelleri listeleyen bir model kataloğu sağlar. Bu adımda SDK'yı başlatır ve sohbet yardımcınız için bir model seçersiniz.

  1. adlı main.pybir dosya oluşturun.

  2. SDK'yi başlatmak için aşağıdaki kodu ekleyin ve bir model seçin:

    # Initialize the Foundry Local SDK
    config = Configuration(app_name="foundry_local_samples")
    FoundryLocalManager.initialize(config)
    manager = FoundryLocalManager.instance
    
    # Download and register all execution providers.
    current_ep = ""
    
    def ep_progress(ep_name: str, percent: float):
        nonlocal current_ep
        if ep_name != current_ep:
            if current_ep:
                print()
            current_ep = ep_name
        print(f"\r  {ep_name:<30}  {percent:5.1f}%", end="", flush=True)
    
    manager.download_and_register_eps(progress_callback=ep_progress)
    if current_ep:
        print()
    
    # Select and load a model from the catalog
    model = manager.catalog.get_model("qwen2.5-0.5b")
    model.download(
        lambda progress: print(
            f"\rDownloading model: {progress:.2f}%", end="", flush=True
        )
    )
    print()
    model.load()
    print("Model loaded and ready.")
    
    # Get a chat client
    client = model.get_chat_client()
    

    Bu get_model yöntemi, katalogdaki belirli bir modelle eşleşen kısa ve kolay bir ad olan model takma adını kabul eder. download yöntemi, modelin ağırlıklarını yerel önbelleğinize getirir ve load modeli çıkarım için hazır hale getirir.

Sistem istemi tanımlama

Sistem istemi, yardımcının kişiliğini ve davranışını ayarlar. Konuşma geçmişindeki ilk iletidir ve model konuşma boyunca buna başvurur.

Yardımcının nasıl yanıt vereceğini şekillendirmek için bir sistem istemi ekleyin:

# Start the conversation with a system prompt
messages = [
    {
        "role": "system",
        "content": "You are a helpful, friendly assistant. Keep your responses "
        "concise and conversational. If you don't know something, say so.",
    }
]

Tavsiye

Yardımcının davranışını değiştirmek için farklı sistem istemleriyle denemeler yapın. Örneğin, korsan, öğretmen veya etki alanı uzmanı olarak yanıt vermesini sağlayabilirsiniz.

Çok aşamalı konuşmayı uygula

Bir sohbet asistanının birden çok etkileşimde bağlamı koruması gerekir. Bunu, tüm iletilerin (sistem, kullanıcı ve yardımcı) listesini tutarak ve her istekle birlikte tam listeyi göndererek elde edebilirsiniz. Model bağlamsal olarak ilgili yanıtlar oluşturmak için bu geçmişi kullanır.

Şu şekilde bir konuşma döngüsü ekleyin:

  • Konsoldan kullanıcı girişini okur.
  • Kullanıcı iletisini geçmişe ekler.
  • Tüm geçmişi modele gönderir.
  • Bir sonraki adım için yardımcının yanıtını geçmişe ekler.
while True:
    user_input = input("You: ")
    if user_input.strip().lower() in ("quit", "exit"):
        break

    # Add the user's message to conversation history
    messages.append({"role": "user", "content": user_input})

    # Stream the response token by token
    print("Assistant: ", end="", flush=True)
    full_response = ""
    for chunk in client.complete_streaming_chat(messages):
        if not chunk.choices:
            continue
        content = chunk.choices[0].delta.content
        if content:
            print(content, end="", flush=True)
            full_response += content
    print("\n")

    # Add the complete response to conversation history
    messages.append({"role": "assistant", "content": full_response})

için yapılan her çağrı complete_chat , ileti geçmişinin tamamını alır. Modelin önceki dönüşleri "anımsaması" budur; çağrılar arasında durum depolamaz.

Akış yanıtları ekleme

Akış, oluşturulduktan sonra her belirteci yazdırır ve bu da yardımcının daha hızlı yanıt vermesine neden olur. Çağrıyı complete_chat ile complete_streaming_chat çağrısı ile değiştirerek yanıtı belirteç bazında akışa alın.

Akış kullanmak için konuşma döngüsünü güncelleştirin:

# Stream the response token by token
print("Assistant: ", end="", flush=True)
full_response = ""
for chunk in client.complete_streaming_chat(messages):
    if not chunk.choices:
        continue
    content = chunk.choices[0].delta.content
    if content:
        print(content, end="", flush=True)
        full_response += content
print("\n")

Akış sürümü, akış tamamlandıktan sonra konuşma geçmişine eklenebilmesi için tam yanıtı biriktirir.

Kodu tamamlama

adlı main.py bir dosya oluşturun ve aşağıdaki tam kodu ekleyin:

from foundry_local_sdk import Configuration, FoundryLocalManager



def main():
    # Initialize the Foundry Local SDK
    config = Configuration(app_name="foundry_local_samples")
    FoundryLocalManager.initialize(config)
    manager = FoundryLocalManager.instance

    # Download and register all execution providers.
    current_ep = ""

    def ep_progress(ep_name: str, percent: float):
        nonlocal current_ep
        if ep_name != current_ep:
            if current_ep:
                print()
            current_ep = ep_name
        print(f"\r  {ep_name:<30}  {percent:5.1f}%", end="", flush=True)

    manager.download_and_register_eps(progress_callback=ep_progress)
    if current_ep:
        print()

    # Select and load a model from the catalog
    model = manager.catalog.get_model("qwen2.5-0.5b")
    model.download(
        lambda progress: print(
            f"\rDownloading model: {progress:.2f}%", end="", flush=True
        )
    )
    print()
    model.load()
    print("Model loaded and ready.")

    # Get a chat client
    client = model.get_chat_client()

    # Start the conversation with a system prompt
    messages = [
        {
            "role": "system",
            "content": "You are a helpful, friendly assistant. Keep your responses "
            "concise and conversational. If you don't know something, say so.",
        }
    ]

    print("\nChat assistant ready! Type 'quit' to exit.\n")

    while True:
        user_input = input("You: ")
        if user_input.strip().lower() in ("quit", "exit"):
            break

        # Add the user's message to conversation history
        messages.append({"role": "user", "content": user_input})

        # Stream the response token by token
        print("Assistant: ", end="", flush=True)
        full_response = ""
        for chunk in client.complete_streaming_chat(messages):
            if not chunk.choices:
                continue
            content = chunk.choices[0].delta.content
            if content:
                print(content, end="", flush=True)
                full_response += content
        print("\n")

        # Add the complete response to conversation history
        messages.append({"role": "assistant", "content": full_response})

    # Clean up - unload the model
    model.unload()
    print("Model unloaded. Goodbye!")


if __name__ == "__main__":
    main()

Sohbet yardımcısını çalıştırın:

python main.py

Şuna benzer bir çıkış görürsünüz:

Downloading model: 100.00%
Model loaded and ready.

Chat assistant ready! Type 'quit' to exit.

You: What is photosynthesis?
Assistant: Photosynthesis is the process plants use to convert sunlight, water, and carbon
dioxide into glucose and oxygen. It mainly happens in the leaves, inside structures
called chloroplasts.

You: Why is it important for other living things?
Assistant: It's essential because photosynthesis produces the oxygen that most living things
breathe. It also forms the base of the food chain — animals eat plants or eat other
animals that depend on plants for energy.

You: quit
Model unloaded. Goodbye!

Asistanın bağlamı önceki dönüşlerden nasıl hatırladığını fark edin. "Diğer canlılar için neden önemlidir?" sorusunu sorduğunuzda, hala fotosentezden bahsettiğinizi bilir.

Örnekler deposu

Bu makalenin örnek kodunun tamamına foundry-samples GitHub deposundan ulaşabilirsiniz. Depoyu klonlamak ve örneğe gitmek için şu adımları kullanın:

git clone https://github.com/microsoft-foundry/foundry-samples.git
cd foundry-samples/samples/rust/foundry-local/tutorial-chat-assistant

Paketleri yükleme

Windows üzerinde geliştiriyor veya gönderiyorsanız Windows sekmesini seçin. Windows paketi Windows ML çalışma zamanıyla tümleşir; aynı API yüzey alanını daha geniş bir donanım hızlandırmasıyla sağlar.

cargo add foundry-local-sdk --features winml
cargo add tokio --features full
cargo add tokio-stream anyhow

Kataloğa göz atın ve bir model seçin

Foundry Yerel SDK'sı tüm kullanılabilir modelleri listeleyen bir model kataloğu sağlar. Bu adımda SDK'yı başlatır ve sohbet yardımcınız için bir model seçersiniz.

  • SDK'yi başlatmak ve bir model seçmek için içeriğini aşağıdaki kodla açın src/main.rs ve değiştirin:

    // Initialize the Foundry Local SDK
    let manager = FoundryLocalManager::create(FoundryLocalConfig::new("foundry_local_samples"))?;
    
    // Download and register all execution providers.
    manager
        .download_and_register_eps_with_progress(None, {
            let mut current_ep = String::new();
            move |ep_name: &str, percent: f64| {
                if ep_name != current_ep {
                    if !current_ep.is_empty() {
                        println!();
                    }
                    current_ep = ep_name.to_string();
                }
                print!("\r  {:<30}  {:5.1}%", ep_name, percent);
                io::stdout().flush().ok();
            }
        })
        .await?;
    println!();
    
    // Select and load a model from the catalog
    let model = manager.catalog().get_model("qwen2.5-0.5b").await?;
    
    if !model.is_cached().await? {
        println!("Downloading model...");
        model
            .download(Some(|progress: f64| {
                print!("\r  {progress:.1}%");
                io::stdout().flush().ok();
            }))
            .await?;
        println!();
    }
    
    model.load().await?;
    println!("Model loaded and ready.");
    
    // Create a chat client
    let client = model.create_chat_client().temperature(0.7).max_tokens(512);
    

    Bu get_model yöntemi, katalogdaki belirli bir modelle eşleşen kısa ve kolay bir ad olan model takma adını kabul eder. download yöntemi, modelin ağırlıklarını yerel önbelleğinize getirir ve load modeli çıkarım için hazır hale getirir.

Sistem istemi tanımlama

Sistem istemi, yardımcının kişiliğini ve davranışını ayarlar. Konuşma geçmişindeki ilk iletidir ve model konuşma boyunca buna başvurur.

Yardımcının nasıl yanıt vereceğini şekillendirmek için bir sistem istemi ekleyin:

// Start the conversation with a system prompt
let mut messages: Vec<ChatCompletionRequestMessage> = vec![
    ChatCompletionRequestSystemMessage::from(
        "You are a helpful, friendly assistant. Keep your responses \
         concise and conversational. If you don't know something, say so.",
    )
    .into(),
];

Tavsiye

Yardımcının davranışını değiştirmek için farklı sistem istemleriyle denemeler yapın. Örneğin, korsan, öğretmen veya etki alanı uzmanı olarak yanıt vermesini sağlayabilirsiniz.

Çok aşamalı konuşmayı uygula

Bir sohbet asistanının birden çok etkileşimde bağlamı koruması gerekir. Bunu, tüm iletilerin (sistem, kullanıcı ve yardımcı) vektörünü tutarak ve her istekle birlikte tam listeyi göndererek elde edebilirsiniz. Model bağlamsal olarak ilgili yanıtlar oluşturmak için bu geçmişi kullanır.

Şu şekilde bir konuşma döngüsü ekleyin:

  • Konsoldan kullanıcı girişini okur.
  • Kullanıcı iletisini geçmişe ekler.
  • Tüm geçmişi modele gönderir.
  • Bir sonraki adım için yardımcının yanıtını geçmişe ekler.
loop {
    print!("You: ");
    io::stdout().flush()?;

    let mut input = String::new();
    stdin.lock().read_line(&mut input)?;
    let input = input.trim();

    if input.eq_ignore_ascii_case("quit") || input.eq_ignore_ascii_case("exit") {
        break;
    }

    // Add the user's message to conversation history
    messages.push(ChatCompletionRequestUserMessage::from(input).into());

    // Stream the response token by token
    print!("Assistant: ");
    io::stdout().flush()?;
    let mut full_response = String::new();
    let mut stream = client.complete_streaming_chat(&messages, None).await?;
    while let Some(chunk) = stream.next().await {
        let chunk = chunk?;
        if let Some(choice) = chunk.choices.first() {
            if let Some(ref content) = choice.delta.content {
                print!("{content}");
                io::stdout().flush()?;
                full_response.push_str(content);
            }
        }
    }
    println!("\n");

    // Add the complete response to conversation history
    let assistant_msg: ChatCompletionRequestMessage = serde_json::from_value(
        serde_json::json!({"role": "assistant", "content": full_response}),
    )?;
    messages.push(assistant_msg);
}

için yapılan her çağrı complete_chat , ileti geçmişinin tamamını alır. Modelin önceki dönüşleri "anımsaması" budur; çağrılar arasında durum depolamaz.

Akış yanıtları ekleme

Akış, oluşturulduktan sonra her belirteci yazdırır ve bu da yardımcının daha hızlı yanıt vermesine neden olur. Çağrıyı complete_chat ile complete_streaming_chat çağrısı ile değiştirerek yanıtı belirteç bazında akışa alın.

Akış kullanmak için konuşma döngüsünü güncelleştirin:

// Stream the response token by token
print!("Assistant: ");
io::stdout().flush()?;
let mut full_response = String::new();
let mut stream = client.complete_streaming_chat(&messages, None).await?;
while let Some(chunk) = stream.next().await {
    let chunk = chunk?;
    if let Some(choice) = chunk.choices.first() {
        if let Some(ref content) = choice.delta.content {
            print!("{content}");
            io::stdout().flush()?;
            full_response.push_str(content);
        }
    }
}
println!("\n");

Akış sürümü, akış tamamlandıktan sonra konuşma geçmişine eklenebilmesi için tam yanıtı biriktirir.

Kodu tamamlama

src/main.rs öğesinin içeriğini aşağıdaki tam kod ile değiştirin:

use foundry_local_sdk::{
    ChatCompletionRequestMessage,
    ChatCompletionRequestSystemMessage, ChatCompletionRequestUserMessage,
    FoundryLocalConfig, FoundryLocalManager,
};
use std::io::{self, BufRead, Write};
use tokio_stream::StreamExt;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
    // Initialize the Foundry Local SDK
    let manager = FoundryLocalManager::create(FoundryLocalConfig::new("foundry_local_samples"))?;

    // Download and register all execution providers.
    manager
        .download_and_register_eps_with_progress(None, {
            let mut current_ep = String::new();
            move |ep_name: &str, percent: f64| {
                if ep_name != current_ep {
                    if !current_ep.is_empty() {
                        println!();
                    }
                    current_ep = ep_name.to_string();
                }
                print!("\r  {:<30}  {:5.1}%", ep_name, percent);
                io::stdout().flush().ok();
            }
        })
        .await?;
    println!();

    // Select and load a model from the catalog
    let model = manager.catalog().get_model("qwen2.5-0.5b").await?;

    if !model.is_cached().await? {
        println!("Downloading model...");
        model
            .download(Some(|progress: f64| {
                print!("\r  {progress:.1}%");
                io::stdout().flush().ok();
            }))
            .await?;
        println!();
    }

    model.load().await?;
    println!("Model loaded and ready.");

    // Create a chat client
    let client = model.create_chat_client().temperature(0.7).max_tokens(512);

    // Start the conversation with a system prompt
    let mut messages: Vec<ChatCompletionRequestMessage> = vec![
        ChatCompletionRequestSystemMessage::from(
            "You are a helpful, friendly assistant. Keep your responses \
             concise and conversational. If you don't know something, say so.",
        )
        .into(),
    ];

    println!("\nChat assistant ready! Type 'quit' to exit.\n");

    let stdin = io::stdin();
    loop {
        print!("You: ");
        io::stdout().flush()?;

        let mut input = String::new();
        stdin.lock().read_line(&mut input)?;
        let input = input.trim();

        if input.eq_ignore_ascii_case("quit") || input.eq_ignore_ascii_case("exit") {
            break;
        }

        // Add the user's message to conversation history
        messages.push(ChatCompletionRequestUserMessage::from(input).into());

        // Stream the response token by token
        print!("Assistant: ");
        io::stdout().flush()?;
        let mut full_response = String::new();
        let mut stream = client.complete_streaming_chat(&messages, None).await?;
        while let Some(chunk) = stream.next().await {
            let chunk = chunk?;
            if let Some(choice) = chunk.choices.first() {
                if let Some(ref content) = choice.delta.content {
                    print!("{content}");
                    io::stdout().flush()?;
                    full_response.push_str(content);
                }
            }
        }
        println!("\n");

        // Add the complete response to conversation history
        let assistant_msg: ChatCompletionRequestMessage = serde_json::from_value(
            serde_json::json!({"role": "assistant", "content": full_response}),
        )?;
        messages.push(assistant_msg);
    }

    // Clean up - unload the model
    model.unload().await?;
    println!("Model unloaded. Goodbye!");

    Ok(())
}

Sohbet yardımcısını çalıştırın:

cargo run

Şuna benzer bir çıkış görürsünüz:

Downloading model: 100.00%
Model loaded and ready.

Chat assistant ready! Type 'quit' to exit.

You: What is photosynthesis?
Assistant: Photosynthesis is the process plants use to convert sunlight, water, and carbon
dioxide into glucose and oxygen. It mainly happens in the leaves, inside structures
called chloroplasts.

You: Why is it important for other living things?
Assistant: It's essential because photosynthesis produces the oxygen that most living things
breathe. It also forms the base of the food chain — animals eat plants or eat other
animals that depend on plants for energy.

You: quit
Model unloaded. Goodbye!

Asistanın bağlamı önceki dönüşlerden nasıl hatırladığını fark edin. "Diğer canlılar için neden önemlidir?" sorusunu sorduğunuzda, hala fotosentezden bahsettiğinizi bilir.

Kaynakları temizle

Modeli kaldırmanızın ardından model ağırlıkları yerel önbelleğinizde kalır. Bu, uygulamayı bir sonraki çalıştırmanızda indirme adımının atlandığı ve modelin daha hızlı yüklendiği anlamına gelir. Disk alanını geri kazanmak istemediğiniz sürece fazladan temizleme gerekmez.