Meta Unveils TRIBE v2: An Advanced Brain Encoding Model for fMRI Response Prediction

Meta's FAIR team has launched TRIBE v2, a tri-modal model that predicts fMRI responses by integrating video, audio, and text stimuli, enhancing our understanding of brain function.

5 min readTechnology

The field of neuroscience has traditionally focused on isolating cognitive functions to specific brain areas, leading to a fragmented understanding of how the brain processes multisensory information. To address this challenge, Meta's FAIR team has developed TRIBE v2, a tri-modal foundation model that aligns AI representations with human brain activity, enabling high-resolution predictions of fMRI responses across various stimuli.

TRIBE v2 employs a sophisticated architecture that includes three frozen foundation models for feature extraction: LLaMA 3.2 for text, V-JEPA2 for video, and Wav2Vec-BERT 2.0 for audio. These models process inputs to create a multi-modal time series that is analyzed by a Transformer encoder. This innovative approach allows for subject-specific predictions of brain activity, projecting outputs onto cortical and subcortical structures.

The model was trained on extensive fMRI datasets, demonstrating a significant increase in predictive accuracy as data volume grows. Notably, TRIBE v2 excels in zero-shot generalization, accurately predicting responses for new subjects without additional training. It also facilitates in-silico experimentation, replicating key functional areas in the brain through virtual simulations, thereby advancing our understanding of neural networks.

Technology