AI giọng nói thời gian thực: Chuyển đổi giọng nói Whisper trực tiếp bằng WebRTC & WebSockets trong Rust

Mục lục bài viết(6 mục)
Hướng dẫn này trình bày chi tiết việc xây dựng một đường ống chuyển đổi giọng nói thành văn bản theo thời gian thực dưới 200ms. Chúng ta sẽ đề cập đến việc thu nhận âm thanh từ client qua WebRTC và WebSockets, phân đoạn âm thanh, suy luận Whisper bằng cách sử dụng các ràng buộc C của Rust, Silero VAD và truyền phát các token chuyển đổi giọng nói thành văn bản một phần.
Tổng quan kiến trúc
Hệ thống bao gồm giao diện WebRTC/WebSocket phía client và backend dựa trên Rust. Client thu âm thanh, mã hóa và truyền đến máy chủ. Máy chủ giải mã, xử lý và chuyển đổi âm thanh bằng mô hình Whisper đã được lượng tử hóa, tận dụng các đặc tính hiệu suất của Rust và các ràng buộc C cho whisper.cpp. Phát hiện hoạt động giọng nói (VAD) được tích hợp để tối ưu hóa suy luận.
Thu nhận và truyền âm thanh phía client
Việc thu nhận âm thanh phía client được xử lý thông qua MediaDevices.getUserMedia() và API MediaRecorder của WebRTC, hoặc trực tiếp qua AudioContext để kiểm soát xử lý chi tiết hơn. Đối với truyền phát thời gian thực, độ trễ thấp, WebSockets được ưu tiên để gửi các bộ đệm âm thanh thô. WebRTC cung cấp khả năng ngang hàng và xử lý phương tiện tích hợp, nhưng đối với dịch vụ chuyển đổi giọng nói thành văn bản tập trung vào máy chủ, kết nối WebSocket cho dữ liệu âm thanh thường đơn giản hơn để quản lý và mở rộng.
Chúng ta sẽ sử dụng AudioContext để thu dữ liệu PCM thô, lấy mẫu lại ở 16kHz và gửi qua WebSocket.
// client/src/audioStreamer.ts
export class AudioStreamer {
private audioContext: AudioContext | null = null;
private mediaStream: MediaStream | null = null;
private audioInput: MediaStreamAudioSourceNode | null = null;
private processor: AudioWorkletNode | ScriptProcessorNode | null = null;
private ws: WebSocket | null = null;
private readonly sampleRate = 16000; // Target sample rate for Whisper
private readonly bufferSize = 4096; // Audio buffer size for processing
constructor(private websocketUrl: string) {}
public async start(): Promise<void> {
if (this.audioContext) {
console.warn("AudioStreamer already started.");
return;
}
this.audioContext = new (window.AudioContext || (window as any).webkitAudioContext)({
sampleRate: this.sampleRate,
});
try {
this.mediaStream = await navigator.mediaDevices.getUserMedia({ audio: true });
this.audioInput = this.audioContext.createMediaStreamSource(this.mediaStream);
// Use AudioWorklet for better performance and off-main-thread processing
// Fallback to ScriptProcessorNode if AudioWorklet is not available
if (this.audioContext.audioWorklet) {
await this.audioContext.audioWorklet.addModule('/audio-processor.js');
this.processor = new AudioWorkletNode(this.audioContext, 'audio-processor');
} else {
console.warn("AudioWorklet not supported, falling back to ScriptProcessorNode.");
this.processor = this.audioContext.createScriptProcessor(this.bufferSize, 1, 1);
(this.processor as ScriptProcessorNode).onaudioprocess = this.handleAudioProcess.bind(this);
}
this.audioInput.connect(this.processor);
this.processor.connect(this.audioContext.destination); // Connect to destination to keep it alive
this.ws = new WebSocket(this.websocketUrl);
this.ws.binaryType = 'arraybuffer';
this.ws.onopen = () => console.log('WebSocket connected.');
this.ws.onclose = () => console.log('WebSocket disconnected.');
this.ws.onerror = (error) => console.error('WebSocket error:', error);
if (this.processor instanceof AudioWorkletNode) {
this.processor.port.onmessage = (event) => {
if (this.ws?.readyState === WebSocket.OPEN) {
this.ws.send(event.data); // Send raw PCM float32 data
}
};
}
} catch (error) {
console.error('Error starting audio stream:', error);
this.stop();
throw error;
}
}
private handleAudioProcess(event: AudioProcessingEvent): void {
if (this.ws?.readyState === WebSocket.OPEN) {
const inputBuffer = event.inputBuffer.getChannelData(0);
this.ws.send(inputBuffer.buffer); // Send raw PCM float32 data
}
}
public stop(): void {
if (this.processor) {
this.processor.disconnect();
this.processor = null;
}
if (this.audioInput) {
this.audioInput.disconnect();
this.audioInput = null;
}
if (this.mediaStream) {
this.mediaStream.getTracks().forEach(track => track.stop());
this.mediaStream = null;
}
if (this.audioContext) {
this.audioContext.close();
this.audioContext = null;
}
if (this.ws) {
this.ws.close();
this.ws = null;
}
console.log('AudioStreamer stopped.');
}
}
// client/public/audio-processor.js (AudioWorklet module)
// This file needs to be served by your web server
class AudioProcessor extends AudioWorkletProcessor {
process(inputs, outputs, parameters) {
const input = inputs[0];
if (input.length > 0) {
const channelData = input[0]; // Get the first channel (mono)
this.port.postMessage(channelData.buffer, [channelData.buffer]); // Transfer ownership
}
return true; // Keep the processor alive
}
}
registerProcessor('audio-processor', AudioProcessor);
Backend Rust: WebSockets, xử lý âm thanh và suy luận
Backend Rust sẽ sử dụng tokio cho các hoạt động bất đồng bộ, warp hoặc axum cho máy chủ WebSocket và hound để mã hóa WAV (nếu cần để gỡ lỗi/lưu trữ). Đối với Whisper, chúng ta sẽ sử dụng whisper-rs (ràng buộc Rust cho whisper.cpp) và silero-vad cho VAD.
Các phụ thuộc
# Cargo.toml
[dependencies]
tokio = { version = "1", features = ["full"] }
warp = "0.3" # Or axum = { version = "0.6", features = ["ws"] }
futures-util = "0.3"
bytes = "1"
log = "0.4"
env_logger = "0.10"
# For whisper.cpp bindings
whisper-rs = { version = "0.1.1", features = ["full"] } # Ensure whisper.cpp is built with GGML_CUDA=1 if using GPU
# For VAD
silero-vad = "0.1.0" # Or a custom VAD implementation
# For audio processing
symphonia = { version = "0.5", features = ["all"] } # For potential audio decoding/resampling if not 16kHz PCM
# For serialization/deserialization
serde = { version = "1", features = ["derive"] }
serde_json = "1"
Cấu trúc Backend
Logic cốt lõi bao gồm:
- Máy chủ WebSocket: Chấp nhận kết nối client và xử lý các khung âm thanh đến.
- Quản lý bộ đệm âm thanh: Tích lũy các khung âm thanh đến vào một bộ đệm lớn hơn phù hợp cho VAD và Whisper.
- VAD: Phát hiện các phân đoạn giọng nói để kích hoạt suy luận Whisper.
- Suy luận Whisper: Chuyển đổi giọng nói đã phát hiện.
- Truyền phát kết quả: Gửi kết quả chuyển đổi giọng nói một phần và cuối cùng trở lại client.
// src/main.rs
use tokio::sync::mpsc;
use tokio_stream::wrappers::ReceiverStream;
use futures_util::{StreamExt, SinkExt};
use warp::ws::{Message, WebSocket};
use warp::Filter;
use std::sync::{Arc, Mutex};
use std::collections::VecDeque;
use std::time::{Instant, Duration};
// --- Whisper and VAD related imports ---
use whisper_rs::{FullParams, SamplingStrategy, WhisperContext, WhisperContextParameters};
use silero_vad::{Vad, VadNode};
const SAMPLE_RATE: u32 = 16000;
const WHISPER_MODEL_PATH: &str = "path/to/ggml-medium.en.bin"; // Path to your quantized Whisper model
const VAD_MODEL_PATH: &str = "path/to/silero_vad.onnx"; // Path to your Silero VAD ONNX model
// Audio buffer for a single client
struct ClientAudioBuffer {
buffer: VecDeque<f32>,
last_audio_activity: Instant,
is_speaking: bool,
}
impl ClientAudioBuffer {
fn new() -> Self {
ClientAudioBuffer {
buffer: VecDeque::new(),
last_audio_activity: Instant::now(),
is_speaking: false,
}
}
fn push_audio(&mut self, audio_data: &[f32]) {
self.buffer.extend(audio_data);
self.last_audio_activity = Instant::now();
}
fn get_audio_chunk(&mut self, duration_ms: u64) -> Option<Vec<f32>> {
let samples_needed = (SAMPLE_RATE as f32 * duration_ms as f32 / 1000.0) as usize;
if self.buffer.len() >= samples_needed {
let chunk: Vec<f32> = self.buffer.drain(0..samples_needed).collect();
Some(chunk)
} else {
None
}
}
fn clear_buffer(&mut self) {
self.buffer.clear();
}
}
#[tokio::main]
async fn main() {
env_logger::init();
// Load Whisper model once
let whisper_context = Arc::new(
WhisperContext::new_with_params(
WHISPER_MODEL_PATH,
WhisperContextParameters::default(),
)
.expect("Failed to load Whisper model"),
);
log::info!("Whisper model loaded: {}", WHISPER_MODEL_PATH);
// Load VAD model once
let vad_model = Arc::new(
Vad::builder()
.set_model_path(VAD_MODEL_PATH)
.set_sample_rate(SAMPLE_RATE)
.build()
.expect("Failed to load Silero VAD model"),
);
log::info!("Silero VAD model loaded: {}", VAD_MODEL_PATH);
let whisper_context_filter = warp::any().map(move || Arc::clone(&whisper_context));
let vad_model_filter = warp::any().map(move || Arc::clone(&vad_model));
let ws_route = warp::path("ws")
.and(warp::ws())
.and(whisper_context_filter)
.and(vad_model_filter)
.map(|ws: warp::ws::Ws, whisper_ctx: Arc<WhisperContext>, vad_model: Arc<Vad>| {
ws.on_upgrade(move |websocket| handle_websocket(websocket, whisper_ctx, vad_model))
});
let routes = ws_route.with(warp::log("websocket_server"));
log::info!("Server started on 127.0.0.1:8080");
warp::serve(routes).run(([127, 0, 0, 1], 8080)).await;
}
async fn handle_websocket(
websocket: WebSocket,
whisper_ctx: Arc<WhisperContext>,
vad_model: Arc<Vad>,
) {
let (mut client_ws_tx, mut client_ws_rx) = websocket.split();
let (audio_tx, audio_rx) = mpsc::channel::<Vec<f32>>(100); // Channel for raw audio chunks
let (transcription_tx, mut transcription_rx) = mpsc::channel::<String>(10); // Channel for transcription results
// Spawn a task to send transcription results back to the client
tokio::spawn(async move {
while let Some(transcription) = transcription_rx.recv().await {
if let Err(e) = client_ws_tx.send(Message::text(transcription)).await {
log::error!("Failed to send transcription to client: {}", e);
break;
}
}
log::info!("Transcription sender task terminated.");
});
// Spawn a task to process audio and run VAD/Whisper
let whisper_ctx_clone = Arc::clone(&whisper_ctx);
let vad_model_clone = Arc::clone(&vad_model);
tokio::spawn(async move {
let mut client_audio_buffer = ClientAudioBuffer::new();
let mut vad_node = VadNode::new(vad_model_clone, SAMPLE_RATE, 512); // VAD processing frame size
let mut current_speech_buffer: Vec<f32> = Vec::new();
let mut last_vad_activity = Instant::now();
let vad_timeout = Duration::from_secs(2); // How long to wait after speech ends before transcribing
let mut whisper_session = whisper_ctx_clone.create_state().expect("Failed to create Whisper state");
while let Some(audio_chunk) = audio_rx.recv().await {
client_audio_buffer.push_audio(&audio_chunk);
// Process audio in smaller VAD-friendly chunks
while let Some(vad_chunk) = client_audio_buffer.get_audio_chunk(30) { // 30ms chunks for VAD
let speech_prob = vad_node.process(&vad_chunk).expect("VAD processing failed");
if speech_prob > 0.5 { // Threshold for speech detection
current_speech_buffer.extend(vad_chunk);
last_vad_activity = Instant::now();
client_audio_buffer.is_speaking = true;
} else {
// If not speaking, but we were recently, check for timeout
if client_audio_buffer.is_speaking && last_vad_activity.elapsed() > vad_timeout {
// Speech has ended, transcribe the accumulated buffer
if !current_speech_buffer.is_empty() {
log::info!("Speech ended, transcribing {} samples.", current_speech_buffer.len());
let transcription = run_whisper_inference(
&whisper_ctx_clone,
&mut whisper_session,
¤t_speech_buffer,
);
if let Err(e) = transcription_tx.send(transcription).await {
log::error!("Failed to send final transcription: {}", e);
}
current_speech_buffer.clear();
}
client_audio_buffer.is_speaking = false;
} else if client_audio_buffer.is_speaking {
// Still within timeout, keep accumulating non-speech for context
current_speech_buffer.extend(vad_chunk);
}
}
}
// Periodically transcribe partial results if speaking
if client_audio_buffer.is_speaking && current_speech_buffer.len() > (SAMPLE_RATE as usize * 1) { // Transcribe every 1 second of speech
let partial_transcription = run_whisper_inference(
&whisper_ctx_clone,
&mut whisper_session,
¤t_speech_buffer,
);
if let Err(e) = transcription_tx.send(format!("[partial] {}", partial_transcription)).await {
log::error!("Failed to send partial transcription: {}", e);
}
}
}
// Handle any remaining speech in buffer when audio stream ends
if !current_speech_buffer.is_empty() {
log::info!("Stream ended, transcribing remaining {} samples.", current_speech_buffer.len());
let transcription = run_whisper_inference(
&whisper_ctx_clone,
&mut whisper_session,
¤t_speech_buffer,
);
if let Err(e) = transcription_tx.send(transcription).await {
log::error!("Failed to send final transcription on stream end: {}", e);
}
}
log::info!("Audio processor task terminated.");
});
// Receive audio data from client
while let Some(result) = client_ws_rx.next().await {
match result {
Ok(msg) => {
if msg.is_binary() {
let audio_bytes = msg.as_bytes();
// Assuming client sends f32 raw PCM
let audio_data: Vec<f32> = audio_bytes
.chunks_exact(4)
.map(|chunk| f32::from_le_bytes(chunk.try_into().unwrap()))
.collect();
if let Err(e) = audio_tx.send(audio_data).await {
log::error!("Failed to send audio chunk to processor: {}", e);
break;
}
} else if msg.is_text() {
log::debug!("Received text message from client: {}", msg.to_str().unwrap_or_default());
// Handle control messages if any
}
}
Err(e) => {
log::error!("WebSocket receive error: {}", e);
break;
}
}
}
log::info!("Client WebSocket disconnected.");
}
fn run_whisper_inference(
ctx: &WhisperContext,
state: &mut whisper_rs::WhisperState,
audio_data: &[f32],
) -> String {
let mut params = FullParams::new(SamplingStrategy::Greedy { best_of: 1 });
params.set_print_progress(false);
params.set_print_special(false);
params.set_print_realtime(false);
params.set_print_timestamps(false);
params.set_language(Some("en"));
params.set_n_threads(4); // Adjust based on CPU cores
// Run the inference
state.full(params, audio_data).expect("Failed to run Whisper inference");
// Iterate over the segments and collect the text
let mut result = String::new();
let num_segments = state.full_n_segments().expect("Failed to get number of segments");
for i in 0..num_segments {
let text = state.full_get_segment_text(i).expect("Failed to get segment text");
result.push_str(&text);
}
result.trim().to_string()
}
Các cân nhắc và đánh đổi về hiệu suất
| Tính năng/Chỉ số | Kênh dữ liệu WebRTC | WebSocket (PCM thô) | whisper.cpp (Đã lượng tử hóa) | Silero VAD |
|---|---|---|---|---|
| Độ trễ | Thấp (P2P) | Thấp (Client-Server) | Trung bình (phụ thuộc vào kích thước mô hình/CPU/GPU) | Rất thấp |
| Thông lượng | Cao | Cao | N/A (thời gian suy luận) | N/A (thời gian suy luận) |
| Độ phức tạp | Cao (ICE, SDP, NAT) | Trung bình (Server-Client) | Trung bình (ràng buộc C, quản lý mô hình) | Thấp (thời gian chạy ONNX) |
| Chất lượng âm thanh | Đã thương lượng (Opus) | Thô (có thể cấu hình) | Đầu vào: 16kHz PCM | Đầu vào: 16kHz PCM |
| Sử dụng CPU | Thấp (mã hóa phía client) | Thấp (mã hóa phía client) | Cao (suy luận CPU/GPU) | Thấp |
| Chi phí mạng | Cao hơn (giao thức) | Thấp hơn (dữ liệu thô) | N/A | N/A |
| Khả năng mở rộng | Khó hơn (P2P) | Dễ hơn (cân bằng tải) | Mở rộng theo chiều dọc (GPU) | Mở rộng theo chiều ngang (nhiều phiên bản hơn) |
| Triển khai | Phức tạp | Đơn giản hơn | Yêu cầu tệp mô hình | Yêu cầu mô hình ONNX |
Phân tích độ trễ:
- Thu nhận & Mã hóa âm thanh phía client: ~10-30ms (bộ đệm trình duyệt, xử lý
AudioContext). - Độ trễ mạng (Client đến Server): ~10-100ms (phụ thuộc vào điều kiện mạng).
- Đệm âm thanh máy chủ: ~30-100ms (để tích lũy đủ âm thanh cho VAD/Whisper).
- Xử lý VAD: ~5-10ms mỗi đoạn.
- Suy luận Whisper: ~50-500ms (phụ thuộc vào kích thước mô hình, phần cứng, độ dài đoạn âm thanh). Đối với dưới 200ms, đây là đường dẫn quan trọng. Sử dụng
ggml-tiny.enhoặcggml-base.entrên CPU hoặc GPU có khả năng là điều cần thiết. - Độ trễ mạng (Server đến Client): ~10-100ms.
- Kết xuất client: ~10ms.
Để đạt được độ trễ đầu cuối dưới 200ms, cần phân đoạn mạnh mẽ, VAD nhanh và suy luận Whisper được tối ưu hóa cao (ví dụ: ggml-tiny.en trên CPU hoặc GPU mạnh mẽ).
Các vấn đề và khắc phục sự cố trong sản xuất
- Lỗi tải mô hình Whisper:
- Triệu chứng:
Failed to load Whisper modelhoặcwhisper_init_from_file: failed to open - Nguyên nhân: Đường dẫn không chính xác đến tệp mô hình
ggml-*.bin, quyền truy cập tệp hoặc mô hình bị hỏng. - Khắc phục: Kiểm tra lại
WHISPER_MODEL_PATH. Đảm bảo tiến trình Rust có quyền đọc. Tải lại mô hình nếu bị hỏng. Đảm bảowhisper-rsđược xây dựng với phiên bảnwhisper.cpptương thích.
- Triệu chứng:
- Lỗi tải mô hình VAD:
- Triệu chứng:
Failed to load Silero VAD model - Nguyên nhân: Đường dẫn không chính xác đến
silero_vad.onnx, thiếu các phụ thuộc thời gian chạy ONNX hoặc phiên bản mô hình ONNX không tương thích. - Khắc phục: Xác minh
VAD_MODEL_PATH. Đảm bảoonnxruntimeđược cài đặt và liên kết chính xác nếusilero-vaddựa vào nó (nó thường đi kèm một phiên bản tối thiểu).
- Triệu chứng:
- Lấy mẫu lại âm thanh/Không khớp định dạng:
- Triệu chứng: Chuyển đổi giọng nói bị méo, không có chuyển đổi giọng nói hoặc lỗi
whisper_full: invalid audio length. - Nguyên nhân: Client gửi âm thanh ở tốc độ lấy mẫu hoặc định dạng khác (ví dụ: 48kHz, int16) so với mong đợi (16kHz, f32).
- Khắc phục: Đảm bảo
AudioContextcủa client được cấu hình cho 16kHz. Xác minh chuyển đổif32trên máy chủ. Nếu client gửiint16, hãy chuyển đổi sangf32trên máy chủ.rust// Example: converting i16 to f32 fn convert_i16_to_f32(audio_data_i16: &[i16]) -> Vec<f32> { audio_data_i16.iter().map(|&s| s as f32 / 32768.0).collect() }
- Triệu chứng: Chuyển đổi giọng nói bị méo, không có chuyển đổi giọng nói hoặc lỗi
- Độ trễ cao / Chuyển đổi giọng nói chậm:
- Triệu chứng: Chuyển đổi giọng nói xuất hiện bị trễ đáng kể.
- Nguyên nhân: Các đoạn âm thanh lớn được gửi đến Whisper, CPU/GPU chậm, mô hình Whisper lớn (
ggml-large), không đủ luồng cho Whisper hoặc VAD không kích hoạt suy luận đủ nhanh. - Khắc phục:
- Sử dụng các mô hình Whisper nhỏ hơn (
ggml-tiny.en,ggml-base.en). - Tăng
params.set_n_threads()cho Whisper (lên đến số lõi vật lý). - Tối ưu hóa các tham số VAD (
vad_timeout, kích thướcget_audio_chunk) để kích hoạt suy luận nhanh hơn. - Đảm bảo
whisper.cppđược biên dịch với hỗ trợ GPU (ví dụ:GGML_CUDA=1) nếu có GPU vàwhisper-rsđược cấu hình để sử dụng nó.
- Sử dụng các mô hình Whisper nhỏ hơn (
- Mất kết nối WebSocket:
- Triệu chứng: Nhật ký client hoặc máy chủ hiển thị các sự kiện đóng WebSocket thường xuyên.
- Nguyên nhân: Mạng không ổn định, máy chủ quá tải, lỗi không được xử lý trong trình xử lý WebSocket hoặc
AudioContextphía client bị thu gom rác nếu không được kết nối vớidestination. - Khắc phục: Triển khai logic kết nối lại phía client. Đảm bảo xử lý lỗi phía máy chủ mạnh mẽ. Giữ
AudioContextđược kết nối vớidestinationhoặcGainNodeđược kết nối vớidestinationđể ngăn nó bị loại bỏ. Triển khai nhịp tim WebSocket (ping/pong) để phát hiện các kết nối chết.
- Rò rỉ bộ nhớ:
- Triệu chứng: Mức sử dụng bộ nhớ máy chủ tăng đều đặn theo thời gian.
- Nguyên nhân: Bộ đệm âm thanh không được xóa,
WhisperStatecủa Whisper không được quản lý đúng cách hoặcVecDequetăng vô hạn. - Khắc phục: Đảm bảo
ClientAudioBuffer.clear_buffer()được gọi khi thích hợp.whisper-rsWhisperStatenên được sử dụng lại cho mỗi phiên. Giám sát kích thướcVecDequevà triển khai giới hạn nếu cần.
Các câu hỏi thường gặp
-
Tại sao sử dụng WebSockets thay vì WebRTC để truyền phát âm thanh? Mặc dù WebRTC cung cấp khả năng ngang hàng và xử lý phương tiện tích hợp, nhưng đối với dịch vụ chuyển đổi giọng nói thành văn bản tập trung vào máy chủ, WebSockets thường đơn giản hóa kiến trúc. Tín hiệu của WebRTC, đàm phán ICE và xuyên NAT làm tăng đáng kể độ phức tạp. WebSockets cung cấp một kênh song công hoàn toàn, độ trễ thấp phù hợp để gửi các bộ đệm âm thanh thô trực tiếp đến một máy chủ trung tâm để xử lý, điều này dễ dàng mở rộng theo chiều ngang hơn.
-
Làm cách nào để đạt được độ trễ đầu cuối dưới 200ms? Điều này rất khó khăn. Các chiến lược chính bao gồm:
- Phía client: Thu và gửi các đoạn âm thanh nhỏ (ví dụ: 30-50ms) ngay lập tức.
- Mạng: Giảm thiểu độ trễ mạng (ví dụ: triển khai máy chủ gần người dùng về mặt địa lý).
- Phía máy chủ:
- Sử dụng mô hình Whisper được tối ưu hóa cao, đã được lượng tử hóa (ví dụ:
ggml-tiny.enhoặcggml-base.en). - Tận dụng tăng tốc GPU cho suy luận Whisper nếu có.
- Sử dụng VAD mạnh mẽ để chỉ chuyển đổi giọng nói các phân đoạn giọng nói, giảm thiểu các cuộc gọi Whisper.
- Xử lý âm thanh trong các đoạn nhỏ, có kích thước cố định cho VAD và sau đó tích lũy cho Whisper.
- Tận dụng hiệu suất của Rust và các ràng buộc C của
whisper.cppđể giảm thiểu chi phí. - Truyền phát kết quả một phần ngay khi chúng có sẵn.
- Sử dụng mô hình Whisper được tối ưu hóa cao, đã được lượng tử hóa (ví dụ:
-
Tôi có thể sử dụng mô hình Whisper lớn hơn để có độ chính xác tốt hơn không? Có, nhưng phải trả giá bằng độ trễ tăng lên. Các mô hình lớn hơn như
ggml-medium.enhoặcggml-large.encung cấp độ chính xác cao hơn nhưng yêu cầu tài nguyên tính toán và thời gian suy luận nhiều hơn đáng kể. Đối với các ứng dụng thời gian thực với yêu cầu độ trễ nghiêm ngặt, một mô hình nhỏ hơn thường là một sự thỏa hiệp cần thiết. Cân nhắc sử dụng một mô hình nhỏ hơn cho kết quả một phần theo thời gian thực và một mô hình lớn hơn cho chuyển đổi giọng nói cuối cùng, đã được xử lý hậu kỳ nếu độ chính xác là tối quan trọng. -
VAD cải thiện đường ống chuyển đổi giọng nói thành văn bản như thế nào? Phát hiện hoạt động giọng nói (VAD) rất quan trọng đối với hiệu suất thời gian thực và hiệu quả tài nguyên. Nó xác định các phân đoạn giọng nói trong luồng âm thanh, cho phép mô hình Whisper chỉ xử lý âm thanh liên quan. Điều này làm giảm số lượng cuộc gọi suy luận Whisper, tiết kiệm chu kỳ CPU/GPU và ngăn chặn việc chuyển đổi giọng nói im lặng hoặc tiếng ồn xung quanh, dẫn đến kết quả nhanh hơn và rõ ràng hơn. Nó cũng giúp phân đoạn lời nói liên tục thành các câu nói có ý nghĩa.
-
Điều gì sẽ xảy ra nếu âm thanh client của tôi không phải là 16kHz PCM? Mô hình Whisper mong đợi âm thanh
f32PCM đơn âm 16kHz. Nếu client của bạn gửi âm thanh ở định dạng khác (ví dụ: 48kHz, âm thanh nổi,int16), bạn phải lấy mẫu lại và chuyển đổi nó ở phía client hoặc máy chủ. Thực hiện điều này ở phía client (sử dụng thuộc tínhsampleRatecủaAudioContext) sẽ giảm tải công việc cho máy chủ. Nếu được thực hiện trên máy chủ, các thư viện nhưsymphoniahoặcrubatocó thể xử lý việc lấy mẫu lại và chuyển đổi định dạng một cách hiệu quả.
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

I/O Linux thông lượng cao trong Rust: io_uring, Tokio & Zero-Copy Networking
Hướng dẫn toàn diện về I/O Linux thông lượng cao trong Rust: io_uring, Tokio & zero-copy networking với kiến trúc và ví dụ code cấp độ sản xuất.
Read more
Bằng chứng không kiến thức trong Rust & Circom: Xác minh SnarkJS & Hướng dẫn sản xuất
Hướng dẫn toàn diện về bằng chứng không kiến thức trong Rust & Circom: xác minh SnarkJS & hướng dẫn sản xuất với kiến trúc cấp độ sản xuất và các ví dụ mã.
Read more
Tìm hiểu Lifetimes trong Rust
Nắm vững lifetimes và borrow checker của Rust: hiểu variance của tham chiếu, elision lifetime ẩn danh so với có tên, và tránh các xung đột trình biên dịch phức tạp.
Read more