•18 min read

AI giọng nói thời gian thực: Chuyển đổi giọng nói Whisper trực tiếp bằng WebRTC & WebSockets trong Rust

AI giọng nói thời gian thực: Chuyển đổi giọng nói Whisper trực tiếp bằng WebRTC & WebSockets trong Rust

Hướng dẫn này trình bày chi tiết việc xây dựng một đường ống chuyển đổi giọng nói thành văn bản theo thời gian thực dưới 200ms. Chúng ta sẽ đề cập đến việc thu nhận âm thanh từ client qua WebRTC và WebSockets, phân đoạn âm thanh, suy luận Whisper bằng cách sử dụng các ràng buộc C của Rust, Silero VAD và truyền phát các token chuyển đổi giọng nói thành văn bản một phần.

Audio Briefing
0:00 / 0:00

Tổng quan kiến trúc

Hệ thống bao gồm giao diện WebRTC/WebSocket phía client và backend dựa trên Rust. Client thu âm thanh, mã hóa và truyền đến máy chủ. Máy chủ giải mã, xử lý và chuyển đổi âm thanh bằng mô hình Whisper đã được lượng tử hóa, tận dụng các đặc tính hiệu suất của Rust và các ràng buộc C cho whisper.cpp. Phát hiện hoạt động giọng nói (VAD) được tích hợp để tối ưu hóa suy luận.

Thu nhận và truyền âm thanh phía client

Việc thu nhận âm thanh phía client được xử lý thông qua MediaDevices.getUserMedia() và API MediaRecorder của WebRTC, hoặc trực tiếp qua AudioContext để kiểm soát xử lý chi tiết hơn. Đối với truyền phát thời gian thực, độ trễ thấp, WebSockets được ưu tiên để gửi các bộ đệm âm thanh thô. WebRTC cung cấp khả năng ngang hàng và xử lý phương tiện tích hợp, nhưng đối với dịch vụ chuyển đổi giọng nói thành văn bản tập trung vào máy chủ, kết nối WebSocket cho dữ liệu âm thanh thường đơn giản hơn để quản lý và mở rộng.

Chúng ta sẽ sử dụng AudioContext để thu dữ liệu PCM thô, lấy mẫu lại ở 16kHz và gửi qua WebSocket.

// client/src/audioStreamer.ts
export class AudioStreamer {
  private audioContext: AudioContext | null = null;
  private mediaStream: MediaStream | null = null;
  private audioInput: MediaStreamAudioSourceNode | null = null;
  private processor: AudioWorkletNode | ScriptProcessorNode | null = null;
  private ws: WebSocket | null = null;
  private readonly sampleRate = 16000; // Target sample rate for Whisper
  private readonly bufferSize = 4096; // Audio buffer size for processing

  constructor(private websocketUrl: string) {}

  public async start(): Promise<void> {
    if (this.audioContext) {
      console.warn("AudioStreamer already started.");
      return;
    }

    this.audioContext = new (window.AudioContext || (window as any).webkitAudioContext)({
      sampleRate: this.sampleRate,
    });

    try {
      this.mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
      this.audioInput = this.audioContext.createMediaStreamSource(this.mediaStream);

      // Use AudioWorklet for better performance and off-main-thread processing
      // Fallback to ScriptProcessorNode if AudioWorklet is not available
      if (this.audioContext.audioWorklet) {
        await this.audioContext.audioWorklet.addModule('/audio-processor.js');
        this.processor = new AudioWorkletNode(this.audioContext, 'audio-processor');
      } else {
        console.warn("AudioWorklet not supported, falling back to ScriptProcessorNode.");
        this.processor = this.audioContext.createScriptProcessor(this.bufferSize, 1, 1);
        (this.processor as ScriptProcessorNode).onaudioprocess = this.handleAudioProcess.bind(this);
      }

      this.audioInput.connect(this.processor);
      this.processor.connect(this.audioContext.destination); // Connect to destination to keep it alive

      this.ws = new WebSocket(this.websocketUrl);
      this.ws.binaryType = 'arraybuffer';

      this.ws.onopen = () => console.log('WebSocket connected.');
      this.ws.onclose = () => console.log('WebSocket disconnected.');
      this.ws.onerror = (error) => console.error('WebSocket error:', error);

      if (this.processor instanceof AudioWorkletNode) {
        this.processor.port.onmessage = (event) => {
          if (this.ws?.readyState === WebSocket.OPEN) {
            this.ws.send(event.data); // Send raw PCM float32 data
          }
        };
      }

    } catch (error) {
      console.error('Error starting audio stream:', error);
      this.stop();
      throw error;
    }
  }

  private handleAudioProcess(event: AudioProcessingEvent): void {
    if (this.ws?.readyState === WebSocket.OPEN) {
      const inputBuffer = event.inputBuffer.getChannelData(0);
      this.ws.send(inputBuffer.buffer); // Send raw PCM float32 data
    }
  }

  public stop(): void {
    if (this.processor) {
      this.processor.disconnect();
      this.processor = null;
    }
    if (this.audioInput) {
      this.audioInput.disconnect();
      this.audioInput = null;
    }
    if (this.mediaStream) {
      this.mediaStream.getTracks().forEach(track => track.stop());
      this.mediaStream = null;
    }
    if (this.audioContext) {
      this.audioContext.close();
      this.audioContext = null;
    }
    if (this.ws) {
      this.ws.close();
      this.ws = null;
    }
    console.log('AudioStreamer stopped.');
  }
}

// client/public/audio-processor.js (AudioWorklet module)
// This file needs to be served by your web server
class AudioProcessor extends AudioWorkletProcessor {
  process(inputs, outputs, parameters) {
    const input = inputs[0];
    if (input.length > 0) {
      const channelData = input[0]; // Get the first channel (mono)
      this.port.postMessage(channelData.buffer, [channelData.buffer]); // Transfer ownership
    }
    return true; // Keep the processor alive
  }
}
registerProcessor('audio-processor', AudioProcessor);

Backend Rust: WebSockets, xử lý âm thanh và suy luận

Backend Rust sẽ sử dụng tokio cho các hoạt động bất đồng bộ, warp hoặc axum cho máy chủ WebSocket và hound để mã hóa WAV (nếu cần để gỡ lỗi/lưu trữ). Đối với Whisper, chúng ta sẽ sử dụng whisper-rs (ràng buộc Rust cho whisper.cpp) và silero-vad cho VAD.

Các phụ thuộc

# Cargo.toml
[dependencies]
tokio = { version = "1", features = ["full"] }
warp = "0.3" # Or axum = { version = "0.6", features = ["ws"] }
futures-util = "0.3"
bytes = "1"
log = "0.4"
env_logger = "0.10"
# For whisper.cpp bindings
whisper-rs = { version = "0.1.1", features = ["full"] } # Ensure whisper.cpp is built with GGML_CUDA=1 if using GPU
# For VAD
silero-vad = "0.1.0" # Or a custom VAD implementation
# For audio processing
symphonia = { version = "0.5", features = ["all"] } # For potential audio decoding/resampling if not 16kHz PCM
# For serialization/deserialization
serde = { version = "1", features = ["derive"] }
serde_json = "1"

Cấu trúc Backend

Logic cốt lõi bao gồm:

  1. Máy chủ WebSocket: Chấp nhận kết nối client và xử lý các khung âm thanh đến.
  2. Quản lý bộ đệm âm thanh: Tích lũy các khung âm thanh đến vào một bộ đệm lớn hơn phù hợp cho VAD và Whisper.
  3. VAD: Phát hiện các phân đoạn giọng nói để kích hoạt suy luận Whisper.
  4. Suy luận Whisper: Chuyển đổi giọng nói đã phát hiện.
  5. Truyền phát kết quả: Gửi kết quả chuyển đổi giọng nói một phần và cuối cùng trở lại client.
// src/main.rs
use tokio::sync::mpsc;
use tokio_stream::wrappers::ReceiverStream;
use futures_util::{StreamExt, SinkExt};
use warp::ws::{Message, WebSocket};
use warp::Filter;
use std::sync::{Arc, Mutex};
use std::collections::VecDeque;
use std::time::{Instant, Duration};

// --- Whisper and VAD related imports ---
use whisper_rs::{FullParams, SamplingStrategy, WhisperContext, WhisperContextParameters};
use silero_vad::{Vad, VadNode};

const SAMPLE_RATE: u32 = 16000;
const WHISPER_MODEL_PATH: &str = "path/to/ggml-medium.en.bin"; // Path to your quantized Whisper model
const VAD_MODEL_PATH: &str = "path/to/silero_vad.onnx"; // Path to your Silero VAD ONNX model

// Audio buffer for a single client
struct ClientAudioBuffer {
    buffer: VecDeque<f32>,
    last_audio_activity: Instant,
    is_speaking: bool,
}

impl ClientAudioBuffer {
    fn new() -> Self {
        ClientAudioBuffer {
            buffer: VecDeque::new(),
            last_audio_activity: Instant::now(),
            is_speaking: false,
        }
    }

    fn push_audio(&mut self, audio_data: &[f32]) {
        self.buffer.extend(audio_data);
        self.last_audio_activity = Instant::now();
    }

    fn get_audio_chunk(&mut self, duration_ms: u64) -> Option<Vec<f32>> {
        let samples_needed = (SAMPLE_RATE as f32 * duration_ms as f32 / 1000.0) as usize;
        if self.buffer.len() >= samples_needed {
            let chunk: Vec<f32> = self.buffer.drain(0..samples_needed).collect();
            Some(chunk)
        } else {
            None
        }
    }

    fn clear_buffer(&mut self) {
        self.buffer.clear();
    }
}

#[tokio::main]
async fn main() {
    env_logger::init();

    // Load Whisper model once
    let whisper_context = Arc::new(
        WhisperContext::new_with_params(
            WHISPER_MODEL_PATH,
            WhisperContextParameters::default(),
        )
        .expect("Failed to load Whisper model"),
    );
    log::info!("Whisper model loaded: {}", WHISPER_MODEL_PATH);

    // Load VAD model once
    let vad_model = Arc::new(
        Vad::builder()
            .set_model_path(VAD_MODEL_PATH)
            .set_sample_rate(SAMPLE_RATE)
            .build()
            .expect("Failed to load Silero VAD model"),
    );
    log::info!("Silero VAD model loaded: {}", VAD_MODEL_PATH);


    let whisper_context_filter = warp::any().map(move || Arc::clone(&whisper_context));
    let vad_model_filter = warp::any().map(move || Arc::clone(&vad_model));

    let ws_route = warp::path("ws")
        .and(warp::ws())
        .and(whisper_context_filter)
        .and(vad_model_filter)
        .map(|ws: warp::ws::Ws, whisper_ctx: Arc<WhisperContext>, vad_model: Arc<Vad>| {
            ws.on_upgrade(move |websocket| handle_websocket(websocket, whisper_ctx, vad_model))
        });

    let routes = ws_route.with(warp::log("websocket_server"));

    log::info!("Server started on 127.0.0.1:8080");
    warp::serve(routes).run(([127, 0, 0, 1], 8080)).await;
}

async fn handle_websocket(
    websocket: WebSocket,
    whisper_ctx: Arc<WhisperContext>,
    vad_model: Arc<Vad>,
) {
    let (mut client_ws_tx, mut client_ws_rx) = websocket.split();

    let (audio_tx, audio_rx) = mpsc::channel::<Vec<f32>>(100); // Channel for raw audio chunks
    let (transcription_tx, mut transcription_rx) = mpsc::channel::<String>(10); // Channel for transcription results

    // Spawn a task to send transcription results back to the client
    tokio::spawn(async move {
        while let Some(transcription) = transcription_rx.recv().await {
            if let Err(e) = client_ws_tx.send(Message::text(transcription)).await {
                log::error!("Failed to send transcription to client: {}", e);
                break;
            }
        }
        log::info!("Transcription sender task terminated.");
    });

    // Spawn a task to process audio and run VAD/Whisper
    let whisper_ctx_clone = Arc::clone(&whisper_ctx);
    let vad_model_clone = Arc::clone(&vad_model);
    tokio::spawn(async move {
        let mut client_audio_buffer = ClientAudioBuffer::new();
        let mut vad_node = VadNode::new(vad_model_clone, SAMPLE_RATE, 512); // VAD processing frame size

        let mut current_speech_buffer: Vec<f32> = Vec::new();
        let mut last_vad_activity = Instant::now();
        let vad_timeout = Duration::from_secs(2); // How long to wait after speech ends before transcribing

        let mut whisper_session = whisper_ctx_clone.create_state().expect("Failed to create Whisper state");

        while let Some(audio_chunk) = audio_rx.recv().await {
            client_audio_buffer.push_audio(&audio_chunk);

            // Process audio in smaller VAD-friendly chunks
            while let Some(vad_chunk) = client_audio_buffer.get_audio_chunk(30) { // 30ms chunks for VAD
                let speech_prob = vad_node.process(&vad_chunk).expect("VAD processing failed");

                if speech_prob > 0.5 { // Threshold for speech detection
                    current_speech_buffer.extend(vad_chunk);
                    last_vad_activity = Instant::now();
                    client_audio_buffer.is_speaking = true;
                } else {
                    // If not speaking, but we were recently, check for timeout
                    if client_audio_buffer.is_speaking && last_vad_activity.elapsed() > vad_timeout {
                        // Speech has ended, transcribe the accumulated buffer
                        if !current_speech_buffer.is_empty() {
                            log::info!("Speech ended, transcribing {} samples.", current_speech_buffer.len());
                            let transcription = run_whisper_inference(
                                &whisper_ctx_clone,
                                &mut whisper_session,
                                &current_speech_buffer,
                            );
                            if let Err(e) = transcription_tx.send(transcription).await {
                                log::error!("Failed to send final transcription: {}", e);
                            }
                            current_speech_buffer.clear();
                        }
                        client_audio_buffer.is_speaking = false;
                    } else if client_audio_buffer.is_speaking {
                        // Still within timeout, keep accumulating non-speech for context
                        current_speech_buffer.extend(vad_chunk);
                    }
                }
            }

            // Periodically transcribe partial results if speaking
            if client_audio_buffer.is_speaking && current_speech_buffer.len() > (SAMPLE_RATE as usize * 1) { // Transcribe every 1 second of speech
                let partial_transcription = run_whisper_inference(
                    &whisper_ctx_clone,
                    &mut whisper_session,
                    &current_speech_buffer,
                );
                if let Err(e) = transcription_tx.send(format!("[partial] {}", partial_transcription)).await {
                    log::error!("Failed to send partial transcription: {}", e);
                }
            }
        }
        // Handle any remaining speech in buffer when audio stream ends
        if !current_speech_buffer.is_empty() {
            log::info!("Stream ended, transcribing remaining {} samples.", current_speech_buffer.len());
            let transcription = run_whisper_inference(
                &whisper_ctx_clone,
                &mut whisper_session,
                &current_speech_buffer,
            );
            if let Err(e) = transcription_tx.send(transcription).await {
                log::error!("Failed to send final transcription on stream end: {}", e);
            }
        }
        log::info!("Audio processor task terminated.");
    });

    // Receive audio data from client
    while let Some(result) = client_ws_rx.next().await {
        match result {
            Ok(msg) => {
                if msg.is_binary() {
                    let audio_bytes = msg.as_bytes();
                    // Assuming client sends f32 raw PCM
                    let audio_data: Vec<f32> = audio_bytes
                        .chunks_exact(4)
                        .map(|chunk| f32::from_le_bytes(chunk.try_into().unwrap()))
                        .collect();

                    if let Err(e) = audio_tx.send(audio_data).await {
                        log::error!("Failed to send audio chunk to processor: {}", e);
                        break;
                    }
                } else if msg.is_text() {
                    log::debug!("Received text message from client: {}", msg.to_str().unwrap_or_default());
                    // Handle control messages if any
                }
            }
            Err(e) => {
                log::error!("WebSocket receive error: {}", e);
                break;
            }
        }
    }
    log::info!("Client WebSocket disconnected.");
}

fn run_whisper_inference(
    ctx: &WhisperContext,
    state: &mut whisper_rs::WhisperState,
    audio_data: &[f32],
) -> String {
    let mut params = FullParams::new(SamplingStrategy::Greedy { best_of: 1 });
    params.set_print_progress(false);
    params.set_print_special(false);
    params.set_print_realtime(false);
    params.set_print_timestamps(false);
    params.set_language(Some("en"));
    params.set_n_threads(4); // Adjust based on CPU cores

    // Run the inference
    state.full(params, audio_data).expect("Failed to run Whisper inference");

    // Iterate over the segments and collect the text
    let mut result = String::new();
    let num_segments = state.full_n_segments().expect("Failed to get number of segments");
    for i in 0..num_segments {
        let text = state.full_get_segment_text(i).expect("Failed to get segment text");
        result.push_str(&text);
    }
    result.trim().to_string()
}

Các cân nhắc và đánh đổi về hiệu suất

Tính năng/Chỉ sốKênh dữ liệu WebRTCWebSocket (PCM thô)whisper.cpp (Đã lượng tử hóa)Silero VAD
Độ trễThấp (P2P)Thấp (Client-Server)Trung bình (phụ thuộc vào kích thước mô hình/CPU/GPU)Rất thấp
Thông lượngCaoCaoN/A (thời gian suy luận)N/A (thời gian suy luận)
Độ phức tạpCao (ICE, SDP, NAT)Trung bình (Server-Client)Trung bình (ràng buộc C, quản lý mô hình)Thấp (thời gian chạy ONNX)
Chất lượng âm thanhĐã thương lượng (Opus)Thô (có thể cấu hình)Đầu vào: 16kHz PCMĐầu vào: 16kHz PCM
Sử dụng CPUThấp (mã hóa phía client)Thấp (mã hóa phía client)Cao (suy luận CPU/GPU)Thấp
Chi phí mạngCao hơn (giao thức)Thấp hơn (dữ liệu thô)N/AN/A
Khả năng mở rộngKhó hơn (P2P)Dễ hơn (cân bằng tải)Mở rộng theo chiều dọc (GPU)Mở rộng theo chiều ngang (nhiều phiên bản hơn)
Triển khaiPhức tạpĐơn giản hơnYêu cầu tệp mô hìnhYêu cầu mô hình ONNX

Phân tích độ trễ:

  • Thu nhận & Mã hóa âm thanh phía client: ~10-30ms (bộ đệm trình duyệt, xử lý AudioContext).
  • Độ trễ mạng (Client đến Server): ~10-100ms (phụ thuộc vào điều kiện mạng).
  • Đệm âm thanh máy chủ: ~30-100ms (để tích lũy đủ âm thanh cho VAD/Whisper).
  • Xử lý VAD: ~5-10ms mỗi đoạn.
  • Suy luận Whisper: ~50-500ms (phụ thuộc vào kích thước mô hình, phần cứng, độ dài đoạn âm thanh). Đối với dưới 200ms, đây là đường dẫn quan trọng. Sử dụng ggml-tiny.en hoặc ggml-base.en trên CPU hoặc GPU có khả năng là điều cần thiết.
  • Độ trễ mạng (Server đến Client): ~10-100ms.
  • Kết xuất client: ~10ms.

Để đạt được độ trễ đầu cuối dưới 200ms, cần phân đoạn mạnh mẽ, VAD nhanh và suy luận Whisper được tối ưu hóa cao (ví dụ: ggml-tiny.en trên CPU hoặc GPU mạnh mẽ).

Các vấn đề và khắc phục sự cố trong sản xuất

  1. Lỗi tải mô hình Whisper:
    • Triệu chứng: Failed to load Whisper model hoặc whisper_init_from_file: failed to open
    • Nguyên nhân: Đường dẫn không chính xác đến tệp mô hình ggml-*.bin, quyền truy cập tệp hoặc mô hình bị hỏng.
    • Khắc phục: Kiểm tra lại WHISPER_MODEL_PATH. Đảm bảo tiến trình Rust có quyền đọc. Tải lại mô hình nếu bị hỏng. Đảm bảo whisper-rs được xây dựng với phiên bản whisper.cpp tương thích.
  2. Lỗi tải mô hình VAD:
    • Triệu chứng: Failed to load Silero VAD model
    • Nguyên nhân: Đường dẫn không chính xác đến silero_vad.onnx, thiếu các phụ thuộc thời gian chạy ONNX hoặc phiên bản mô hình ONNX không tương thích.
    • Khắc phục: Xác minh VAD_MODEL_PATH. Đảm bảo onnxruntime được cài đặt và liên kết chính xác nếu silero-vad dựa vào nó (nó thường đi kèm một phiên bản tối thiểu).
  3. Lấy mẫu lại âm thanh/Không khớp định dạng:
    • Triệu chứng: Chuyển đổi giọng nói bị méo, không có chuyển đổi giọng nói hoặc lỗi whisper_full: invalid audio length.
    • Nguyên nhân: Client gửi âm thanh ở tốc độ lấy mẫu hoặc định dạng khác (ví dụ: 48kHz, int16) so với mong đợi (16kHz, f32).
    • Khắc phục: Đảm bảo AudioContext của client được cấu hình cho 16kHz. Xác minh chuyển đổi f32 trên máy chủ. Nếu client gửi int16, hãy chuyển đổi sang f32 trên máy chủ.
      // Example: converting i16 to f32
      fn convert_i16_to_f32(audio_data_i16: &[i16]) -> Vec<f32> {
          audio_data_i16.iter().map(|&s| s as f32 / 32768.0).collect()
      }
      
  4. Độ trễ cao / Chuyển đổi giọng nói chậm:
    • Triệu chứng: Chuyển đổi giọng nói xuất hiện bị trễ đáng kể.
    • Nguyên nhân: Các đoạn âm thanh lớn được gửi đến Whisper, CPU/GPU chậm, mô hình Whisper lớn (ggml-large), không đủ luồng cho Whisper hoặc VAD không kích hoạt suy luận đủ nhanh.
    • Khắc phục:
      • Sử dụng các mô hình Whisper nhỏ hơn (ggml-tiny.en, ggml-base.en).
      • Tăng params.set_n_threads() cho Whisper (lên đến số lõi vật lý).
      • Tối ưu hóa các tham số VAD (vad_timeout, kích thước get_audio_chunk) để kích hoạt suy luận nhanh hơn.
      • Đảm bảo whisper.cpp được biên dịch với hỗ trợ GPU (ví dụ: GGML_CUDA=1) nếu có GPU và whisper-rs được cấu hình để sử dụng nó.
  5. Mất kết nối WebSocket:
    • Triệu chứng: Nhật ký client hoặc máy chủ hiển thị các sự kiện đóng WebSocket thường xuyên.
    • Nguyên nhân: Mạng không ổn định, máy chủ quá tải, lỗi không được xử lý trong trình xử lý WebSocket hoặc AudioContext phía client bị thu gom rác nếu không được kết nối với destination.
    • Khắc phục: Triển khai logic kết nối lại phía client. Đảm bảo xử lý lỗi phía máy chủ mạnh mẽ. Giữ AudioContext được kết nối với destination hoặc GainNode được kết nối với destination để ngăn nó bị loại bỏ. Triển khai nhịp tim WebSocket (ping/pong) để phát hiện các kết nối chết.
  6. Rò rỉ bộ nhớ:
    • Triệu chứng: Mức sử dụng bộ nhớ máy chủ tăng đều đặn theo thời gian.
    • Nguyên nhân: Bộ đệm âm thanh không được xóa, WhisperState của Whisper không được quản lý đúng cách hoặc VecDeque tăng vô hạn.
    • Khắc phục: Đảm bảo ClientAudioBuffer.clear_buffer() được gọi khi thích hợp. whisper-rs WhisperState nên được sử dụng lại cho mỗi phiên. Giám sát kích thước VecDeque và triển khai giới hạn nếu cần.

Các câu hỏi thường gặp

  1. Tại sao sử dụng WebSockets thay vì WebRTC để truyền phát âm thanh? Mặc dù WebRTC cung cấp khả năng ngang hàng và xử lý phương tiện tích hợp, nhưng đối với dịch vụ chuyển đổi giọng nói thành văn bản tập trung vào máy chủ, WebSockets thường đơn giản hóa kiến trúc. Tín hiệu của WebRTC, đàm phán ICE và xuyên NAT làm tăng đáng kể độ phức tạp. WebSockets cung cấp một kênh song công hoàn toàn, độ trễ thấp phù hợp để gửi các bộ đệm âm thanh thô trực tiếp đến một máy chủ trung tâm để xử lý, điều này dễ dàng mở rộng theo chiều ngang hơn.

  2. Làm cách nào để đạt được độ trễ đầu cuối dưới 200ms? Điều này rất khó khăn. Các chiến lược chính bao gồm:

    • Phía client: Thu và gửi các đoạn âm thanh nhỏ (ví dụ: 30-50ms) ngay lập tức.
    • Mạng: Giảm thiểu độ trễ mạng (ví dụ: triển khai máy chủ gần người dùng về mặt địa lý).
    • Phía máy chủ:
      • Sử dụng mô hình Whisper được tối ưu hóa cao, đã được lượng tử hóa (ví dụ: ggml-tiny.en hoặc ggml-base.en).
      • Tận dụng tăng tốc GPU cho suy luận Whisper nếu có.
      • Sử dụng VAD mạnh mẽ để chỉ chuyển đổi giọng nói các phân đoạn giọng nói, giảm thiểu các cuộc gọi Whisper.
      • Xử lý âm thanh trong các đoạn nhỏ, có kích thước cố định cho VAD và sau đó tích lũy cho Whisper.
      • Tận dụng hiệu suất của Rust và các ràng buộc C của whisper.cpp để giảm thiểu chi phí.
      • Truyền phát kết quả một phần ngay khi chúng có sẵn.
  3. Tôi có thể sử dụng mô hình Whisper lớn hơn để có độ chính xác tốt hơn không? Có, nhưng phải trả giá bằng độ trễ tăng lên. Các mô hình lớn hơn như ggml-medium.en hoặc ggml-large.en cung cấp độ chính xác cao hơn nhưng yêu cầu tài nguyên tính toán và thời gian suy luận nhiều hơn đáng kể. Đối với các ứng dụng thời gian thực với yêu cầu độ trễ nghiêm ngặt, một mô hình nhỏ hơn thường là một sự thỏa hiệp cần thiết. Cân nhắc sử dụng một mô hình nhỏ hơn cho kết quả một phần theo thời gian thực và một mô hình lớn hơn cho chuyển đổi giọng nói cuối cùng, đã được xử lý hậu kỳ nếu độ chính xác là tối quan trọng.

  4. VAD cải thiện đường ống chuyển đổi giọng nói thành văn bản như thế nào? Phát hiện hoạt động giọng nói (VAD) rất quan trọng đối với hiệu suất thời gian thực và hiệu quả tài nguyên. Nó xác định các phân đoạn giọng nói trong luồng âm thanh, cho phép mô hình Whisper chỉ xử lý âm thanh liên quan. Điều này làm giảm số lượng cuộc gọi suy luận Whisper, tiết kiệm chu kỳ CPU/GPU và ngăn chặn việc chuyển đổi giọng nói im lặng hoặc tiếng ồn xung quanh, dẫn đến kết quả nhanh hơn và rõ ràng hơn. Nó cũng giúp phân đoạn lời nói liên tục thành các câu nói có ý nghĩa.

  5. Điều gì sẽ xảy ra nếu âm thanh client của tôi không phải là 16kHz PCM? Mô hình Whisper mong đợi âm thanh f32 PCM đơn âm 16kHz. Nếu client của bạn gửi âm thanh ở định dạng khác (ví dụ: 48kHz, âm thanh nổi, int16), bạn phải lấy mẫu lại và chuyển đổi nó ở phía client hoặc máy chủ. Thực hiện điều này ở phía client (sử dụng thuộc tính sampleRate của AudioContext) sẽ giảm tải công việc cho máy chủ. Nếu được thực hiện trên máy chủ, các thư viện như symphonia hoặc rubato có thể xử lý việc lấy mẫu lại và chuyển đổi định dạng một cách hiệu quả.

Share this article:

Stay Updated

Get the latest posts delivered straight to your inbox.

Free Developer Utilities

Free In-Browser Developer Tools

Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.

Explore Tools
Advertisement
Tìm hiểu Lifetimes trong Rust
rust

Tìm hiểu Lifetimes trong Rust

Nắm vững lifetimes và borrow checker của Rust: hiểu variance của tham chiếu, elision lifetime ẩn danh so với có tên, và tránh các xung đột trình biên dịch phức tạp.

Read more