Gradium has launched two cutting-edge models for real-time speech translation: stt-translate, which converts speech to text, and s2s-translate, which translates spoken audio directly into another language. These models support five languages—English, French, German, Spanish, and Portuguese—and facilitate 20 different language pair translations, streamlining the process into a single step. Gradium asserts that their models outperform competitors like gpt-realtime-translate and gemini-3.5-live-translate in terms of accuracy and latency. The stt-translate model integrates transcription and translation into one seamless operation, eliminating the need for an intermediary transcript. This design choice enhances speed and efficiency. The s2s-translate model builds on this by providing synthesized audio output alongside translated transcripts, all through a single WebSocket connection. Gradium measures translation quality using BLEU and MetricX metrics, with their models showing superior performance in both categories. The average latency for s2s-translate is approximately 3.0 seconds, making it competitive in the market. Use cases for these models include live dubbing, multilingual customer support, and real-time meeting translations, showcasing their versatility and effectiveness in various applications.
Gradium Introduces Advanced Real-Time Speech Translation Models
Gradium has unveiled two innovative real-time speech translation models, stt-translate and s2s-translate, claiming superior accuracy and lower latency compared to existing solutions.
