Google AI, in collaboration with various researchers, has introduced WAXAL, a groundbreaking multilingual speech dataset specifically designed for African languages. This dataset encompasses 24 languages and is divided into two distinct components: Automatic Speech Recognition (ASR) and Text-to-Speech (TTS). The ASR segment is derived from natural speech recordings, while the TTS segment features high-quality audio from single speakers. The separation of these components acknowledges the differing requirements for ASR and TTS systems. The ASR data is collected through image-prompted speech, allowing speakers to describe images in their native languages, resulting in more authentic language use. This method captures a wide range of linguistic variations, although it complicates transcription. In contrast, the TTS data is meticulously recorded under controlled conditions, ensuring high fidelity and consistency across recordings. This innovative approach addresses the existing data scarcity for low-resource African languages, paving the way for improved speech technologies.
Google AI Unveils WAXAL: A Comprehensive Multilingual Speech Dataset for African Languages
WAXAL aims to enhance speech technology for African languages by providing a diverse dataset for ASR and TTS systems.
