Tencent AI Unveils Covo-Audio: A 7B Parameter Speech Language Model for Real-Time Audio Interactions

Tencent AI Lab has introduced Covo-Audio, a 7B-parameter model that integrates speech processing and language understanding for real-time audio conversations.

3 min readTechnology

Tencent AI Lab has launched Covo-Audio, a large audio language model featuring 7 billion parameters. This innovative model is engineered to merge speech processing with language comprehension, enabling it to handle continuous audio inputs and produce audio outputs within a single framework. The architecture comprises four key components: an Audio Encoder utilizing Whisper-large-v3 for noise resilience, an Audio Adapter that down-samples audio for compatibility, a backbone based on Qwen2.5-7B-Base for processing interleaved audio and text, and a Speech Tokenizer and Decoder that reconstructs high-quality audio. A significant advancement is the Hierarchical Tri-modal Interleaving strategy, which aligns continuous acoustic features with discrete speech tokens and natural language text. This method enhances the model's ability to maintain semantic coherence. Additionally, the Intelligence Speaker Decoupling strategy allows for flexible voice customization with minimal text-to-speech data. The Covo-Audio-Chat-FD variant supports full-duplex communication, managing conversational dynamics with specific tokens. Performance evaluations indicate that Covo-Audio excels in various benchmarks, achieving top scores in music understanding and conversational tasks, demonstrating its efficiency and effectiveness in audio reasoning and dialogue.

Technology