Apoorv Vyas

Apoorv Vyas

Staff Research Scientist  ·  Meta FAIR
Bay Area, California

I train large multimodal models that hear and see — audio, video, and speech at the scale where representation learning and generation start to meet.

At Meta I proposed and led PE-AV, a family of audio–video–text aligned encoders, and worked on the audio side of Movie Gen, Audiobox, Voicebox, SAM Audio, and MMS. Before that I did my PhD at EPFL and Idiap with Hervé Bourlard and François Fleuret, where linear and clustered attention came out of trying to make speech recognition cheaper.

Research

Multimodal representation learning

Joint embeddings across audio, video, and text, and the data engines that make contrastive training work at 100M-video scale.

Large-scale generative training

Speech and sound generation from text, video, or voice prompts — architecture studies, scaling de-risking, and data curation.

Efficient Transformers

Attention that stays affordable on long sequences: linear attention as an RNN, and clustered attention over queries.

News

2026

PE-AV accepted to CVPR 2026; SAM Audio accepted to ICML 2026.

Dec 2025

Released PE-AV — audio–video–text encoders, with code and open weights.

Dec 2025

Our team launched SAM Audio, extending Segment Anything to audio separation.

Oct 2024

Movie Gen released, including the video-conditioned audio model.

Selected publications

Full list on Google Scholar

Scaling Speech Technology to 1,000+ Languages

Vineel Pratap, Andros Tjandra, Bowen Shi, Paden Tomasello, Arun Babu, Sayani Kundu, Ali Elkahky, Zhaoheng Ni, Apoorv Vyas, Maryam Fazel-Zarandi, Alexei Baevski, Yossi Adi, Xiaohui Zhang, Wei-Ning Hsu, Alexis Conneau, Michael Auli

JMLR 2024
Earlier work — speech recognition and sensing, 2014–2021

Commercial Block Detection in Broadcast News Videos

Apoorv Vyas, Raghvendra Kannao, Vishal Bhargava, Prithwijit Guha

ICVGIP 2014

Also: Pkwrap: a PyTorch Package for LF-MMI Training of Acoustic Models (Madikeri, Tong, Zuluaga-Gomez, Vyas, Motlicek, Bourlard) — paper · code.

Experience

Meta — FAIR (Fundamental AI Research)

Nov 2022 — Present · Menlo Park, CA

Staff Research Scientist (2026–) · Senior Research Scientist (2024–26) · Research Scientist (2022–24)

  • Multimodal representation learning. Proposed and led PE-AV, a family of audio–video–text aligned encoders trained with contrastive learning scaled on caption quality, data types, loss pairs, and model size. Ran the early ablations, scaled the data pipeline to 100M videos, and directed workstreams as the effort grew from 2 to 6 researchers.
  • Multimodal post-training. Integrated audio and speech recognition into a joint text–perception model; designed and ran the ablations behind the integration approach and drove the work to retain model quality.
  • Large-scale audio generation. Movie Gen Audio, Audiobox, and Voicebox — architecture studies, scaling de-risking, data curation, and the joint speech and sound-effect training behind unified generation.

Idiap Research Institute / EPFL

Jul 2018 — Aug 2022 · Switzerland

Doctoral Researcher — Thesis: Efficient Transformer-Based Speech Recognition

  • Compute efficiency. Linear attention, a kernelized formulation expressing autoregressive Transformers as RNNs at linear cost, and clustered attention, which approximates self-attention via input clustering.
  • Data efficiency. Uncertainty-aware semi-supervised learning via dropout, and cross-domain and cross-lingual analysis of SSL pretraining.

Intel Labs

Apr 2015 — May 2018 · India

Research / Systems Engineer

  • Ensemble method for out-of-distribution detection in deep networks.
  • Low-power supervised semantic hashing for approximate image retrieval.

Education & service

Education

PhD, Electrical Engineering EPFL · 2018–2022

B.Tech, Electronics & Electrical Engineering IIT Guwahati · 2010–2014

Patents & service

9 granted US patents 5 at Intel · 4 at Meta

Reviewer NeurIPS · ICLR · ICML · CVPR · ECCV · AAAI · TNNLS