Microsoft Unveils Harrier-OSS-v1: A Cutting-Edge Multilingual Embedding Model Family

Microsoft has launched Harrier-OSS-v1, a new suite of multilingual embedding models that set new benchmarks in the field of semantic representation.

3 min readTechnology

Microsoft has introduced Harrier-OSS-v1, a collection of three multilingual text embedding models aimed at delivering superior semantic representations across multiple languages. This release features models with varying scales: 270 million parameters, 0.6 billion parameters, and 27 billion parameters. The Harrier-OSS-v1 models have achieved state-of-the-art performance on the Multilingual MTEB (Massive Text Embedding Benchmark) v2, marking a significant advancement in open-source retrieval technology. Unlike traditional bidirectional encoder models like BERT, Harrier-OSS-v1 employs decoder-only architectures, which process context differently. Each token in these models can only reference preceding tokens, utilizing last-token pooling to create a single vector representation of the input. A notable characteristic of these models is their ability to handle long-context inputs, with a context window of 32,768 tokens, allowing for the embedding of larger documents without losing semantic integrity. Additionally, the models are instruction-tuned, requiring task-specific instructions for optimal performance. The smaller models benefit from knowledge distillation, enhancing their quality despite fewer parameters. The Harrier-OSS-v1 family demonstrates exceptional capabilities in multilingual tasks, making it a valuable tool for global applications.

Technology