Multimodal representation learning
Joint embeddings across audio, video, and text, and the data engines that make contrastive training work at 100M-video scale.
I train large multimodal models that hear and see — audio, video, and speech at the scale where representation learning and generation start to meet.
At Meta I proposed and led PE-AV, a family of audio–video–text aligned encoders, and worked on the audio side of Movie Gen, Audiobox, Voicebox, SAM Audio, and MMS. Before that I did my PhD at EPFL and Idiap with Hervé Bourlard and François Fleuret, where linear and clustered attention came out of trying to make speech recognition cheaper.
Joint embeddings across audio, video, and text, and the data engines that make contrastive training work at 100M-video scale.
Speech and sound generation from text, video, or voice prompts — architecture studies, scaling de-risking, and data curation.
Attention that stays affordable on long sequences: linear attention as an RNN, and clustered attention over queries.
PE-AV accepted to CVPR 2026; SAM Audio accepted to ICML 2026.
Released PE-AV — audio–video–text encoders, with code and open weights.
Our team launched SAM Audio, extending Segment Anything to audio separation.
Movie Gen released, including the video-conditioned audio model.
Staff Research Scientist (2026–) · Senior Research Scientist (2024–26) · Research Scientist (2022–24)
Doctoral Researcher — Thesis: Efficient Transformer-Based Speech Recognition
Research / Systems Engineer
PhD, Electrical Engineering EPFL · 2018–2022
B.Tech, Electronics & Electrical Engineering IIT Guwahati · 2010–2014
9 granted US patents 5 at Intel · 4 at Meta
Reviewer NeurIPS · ICLR · ICML · CVPR · ECCV · AAAI · TNNLS