Batch inference

Important

This feature is in Public Preview.

These notebooks run large-scale batch inference on AI Runtime serverless GPUs. They cover a range of modalities, such as LLM text generation and speech-to-text, using frameworks like Ray Data, vLLM, and Transformers.

Tutorial Description
Qwen2.5-32B batch inference with Ray Data and vLLM Run multilingual batch inference with eight persistent vLLM replicas of Qwen2.5-32B-Instruct on 8 H100 GPUs and save the results to Unity Catalog.
Batch speech-to-text with Whisper Transcribe audio in batches with OpenAI Whisper large-v3-turbo on serverless GPUs, using libraries pre-installed in AI Runtime.