About
I am a Research Scientist Lead at Luma AI, working on multimodal foundation models.
I received my Ph.D. in Computer Vision from the University of Michigan, advised by Prof. Andrew Owens. Before that, I earned my Master's degree at Michigan and my Bachelor's degree from Shanghai Jiao Tong University.
During my Ph.D., I interned with the Codec Avatars Lab at Meta and the SODA Group at Adobe Research. I also collaborated with Prof. David Fouhey and Prof. Alex Wong.
News
- I am currently a Tech Lead at Luma AI for video-audio generation efforts.
- I joined Luma AI as a full-time Research Scientist!
- I defended my Ph.D. and graduated!
- Three papers were accepted to CVPR 2025! See you in Nashville!
- Check out our new work MultiFoley for creative Foley sound generation!
Research
My research focuses on multimodal learning across vision, audio, language, and touch, unified models that bind these modalities into a shared representation, and diffusion models for multimodal generation.
Work Experience
Industry · Multimodal AI
Luma AI
Redwood City, CA
Research internship · Generative audio
Adobe Research
San Francisco, CA
Research mentors
Featured project
MultiFoleyCVPR 2025
Video-guided sound generation with text, audio, and video conditioning.
Research internship · Spatial audio&video
Meta · Codec Avatars Lab
Pittsburgh, PA
Research mentors
Featured project
Real Acoustic FieldsCVPR 2024 · Highlight
A multimodal dataset capturing real-world room acoustics.
Publication
* Equal contribution
Connecting Sight and Sound through Space, Time and Language
Ziyang Chen
Ph.D. Dissertation, University of Michigan, 2025
dissertation
Video-Guided Foley Sound Generation with Multimodal Controls
Ziyang Chen,
Prem Seetharaman,
Bryan Russell,
Oriol Nieto,
David Bourgin,
Andrew Owens,
Justin Salamon
CVPR 2025
project page
·
paper
·
bibtex
GPS as a Control Signal for Image Generation
Chao Feng,
Ziyang Chen,
Aleksander Holynski,
Alexei A. Efros,
Andrew Owens
CVPR 2025
project page
·
paper
·
code
·
bibtex
Supervising Sound Localization Using In-the-wild Ego-motion
Anna Min,
Ziyang Chen,
Hang Zhao,
Andrew Owens
CVPR 2025 (Highlight)
paper
Images that Sound: Composing Images and Sounds on a Single Canvas
Ziyang Chen,
Daniel Geng,
Andrew Owens
NeurIPS 2024
CVPR 2024 AI Art Gallery
project page
·
paper
·
code
·
bibtex
Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
Ziyang Chen,
Israel D. Gebru,
Christian Richardt,
Anurag Kumar,
William Laney,
Andrew Owens,
Alexander Richard
CVPR 2024 (Highlight -- 2.8% accept rate)
project page
·
paper
·
dataset
·
bibtex
Binding Touch to
Everything: Learning Unified Multimodal Tactile Representations
Fengyu Yang*,
Chao Feng*,
Ziyang Chen*,
...... ,
Andrew Owens,
Alex Wong
CVPR 2024
project page
·
paper
·
code
·
bibtex
Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation
Ziyang Chen, Shengyi Qian, Andrew Owens
ICCV 2023
project page
·
paper
·
code
·
bibtex
Conditional Generation of Audio from Video via Foley Analogies
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, Andrew Owens
CVPR 2023
project page
·
paper
·
code
·
bibtex
Self-Supervised Video Forensics by Audio-Visual Anomaly Detection
Chao Feng, Ziyang Chen, Andrew Owens
CVPR 2023 (Highlight -- 2.5% accept rate)
project page
·
paper
·
code
·
bibtex
Sound Localization by Self-Supervised Time Delay Estimation
Ziyang Chen, David F. Fouhey, Andrew Owens
ECCV 2022
project page
·
paper
·
slides
·
code
·
bibtex
Mix and Localize: Localizing Sound Sources in Mixtures
Xixi Hu*, Ziyang Chen*, Andrew Owens
CVPR 2022
project page
·
paper
·
slides
·
code
·
bibtex
Structure from Silence: Learning Scene Structure from Ambient Sound
Ziyang Chen*, Xixi Hu*, Andrew Owens
CoRL 2021 (Oral)
project page
·
paper
·
slides
·
code
·
bibtex
Profession
Workshop Organizer
Reviewer
- Conferences CVPR (2023–2026), ICCV (2023, 2025), ECCV (2024), NeurIPS (2024), SIGGRAPH (2024), WACV (2023)
- Journals IJCV
Teaching Assistant
- VE312 · Digital Integrated Circuits
- VE311 · Electronic Circuits











