IBM has launched two new models in its Granite Speech series: Granite Speech 4.1 2B and Granite Speech 4.1 2B-NAR, both featuring around 2 billion parameters. These models are accessible on Hugging Face under an Apache 2.0 license. They address a common challenge faced by enterprise AI teams: balancing computational demands with accuracy in automatic speech recognition (ASR). The Granite Speech 4.1 2B model is tailored for multilingual ASR and bidirectional speech translation, supporting languages such as English, French, German, Spanish, Portuguese, and Japanese. In contrast, the 2B-NAR variant focuses solely on ASR, optimizing for low-latency applications but excluding Japanese support. Additionally, a third model, Granite Speech 4.1 2B-Plus, has been released, which includes features for speaker attribution and word-level timestamps. Both models utilize a three-part architecture comprising a speech encoder, a modality adapter, and a language model, with significant differences in their decoding processes. The autoregressive model generates text sequentially, while the non-autoregressive model employs a bidirectional approach for faster inference. The training data varied, with the standard model utilizing 174,000 hours of audio, while the NAR model was trained on 130,000 hours. This release highlights IBM's commitment to advancing speech technology.
IBM Unveils Granite Speech 4.1 2B Models: Advanced ASR Solutions
IBM has introduced two innovative speech recognition models, Granite Speech 4.1 2B and Granite Speech 4.1 2B-NAR, designed to enhance automatic speech recognition and translation capabilities.
