Portrait of Ziyang Chen

About

I am a Research Scientist Lead at Luma AI, working on multimodal foundation models.

I received my Ph.D. in Computer Vision from the University of Michigan, advised by Prof. Andrew Owens. Before that, I earned my Master's degree at Michigan and my Bachelor's degree from Shanghai Jiao Tong University.

During my Ph.D., I interned with the Codec Avatars Lab at Meta and the SODA Group at Adobe Research. I also collaborated with Prof. David Fouhey and Prof. Alex Wong.

News

  • I am currently a Tech Lead at Luma AI for video-audio generation efforts.
  • I joined Luma AI as a full-time Research Scientist!
  • I defended my Ph.D. and graduated!
  • Three papers were accepted to CVPR 2025! See you in Nashville!
  • Check out our new work MultiFoley for creative Foley sound generation!

Research

Illustration of multimodal learning across different data types
Multimodal Learning
Diagram representing a unified multimodal model
Unified Models
Illustration of the diffusion model generation process
Diffusion Models

My research focuses on multimodal learning across vision, audio, language, and touch, unified models that bind these modalities into a shared representation, and diffusion models for multimodal generation.

Work Experience

Industry · Multimodal AI

Luma AI

Redwood City, CA

Current
Research Scientist Lead Jan. 2026 — Present
Research Scientist Jun. 2025 — Jan. 2026

Selected work

  • Ray3.14Fast video generation through distillation
  • Uni-1Unified multimodal reasoning and pixel generation
  • Generative mediaSpeech-to-video and joint audio-video generation

Publication

* Equal contribution

Connecting Sight and Sound through Space, Time and Language
Ziyang Chen
Ph.D. Dissertation, University of Michigan, 2025
dissertation

MultiFoley video-guided sound generation overview

Video-Guided Foley Sound Generation with Multimodal Controls
Ziyang Chen, Prem Seetharaman, Bryan Russell, Oriol Nieto, David Bourgin, Andrew Owens, Justin Salamon
CVPR 2025
project page  ·  paper  ·  bibtex

GPS-controlled image generation overview

GPS as a Control Signal for Image Generation
Chao Feng, Ziyang Chen, Aleksander Holynski, Alexei A. Efros, Andrew Owens
CVPR 2025
project page  ·  paper  ·  code  ·  bibtex

Egomotion-supervised sound localization overview

Supervising Sound Localization Using In-the-wild Ego-motion
Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens
CVPR 2025 (Highlight)
paper

Images that Sound composition examples

Images that Sound: Composing Images and Sounds on a Single Canvas
Ziyang Chen, Daniel Geng, Andrew Owens
NeurIPS 2024
CVPR 2024 AI Art Gallery
project page  ·  paper  ·  code  ·  bibtex

Real Acoustic Fields capture setup

Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark
Ziyang Chen, Israel D. Gebru, Christian Richardt, Anurag Kumar, William Laney, Andrew Owens, Alexander Richard
CVPR 2024 (Highlight -- 2.8% accept rate)
project page  ·  paper  ·  dataset  ·  bibtex

UniTouch multimodal tactile representation overview

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
Fengyu Yang*, Chao Feng*, Ziyang Chen*, ...... , Andrew Owens, Alex Wong
CVPR 2024
project page  ·  paper  ·  code  ·  bibtex

Sound Localization from Motion overview

Sound Localization from Motion: Jointly Learning Sound Direction and Camera Rotation
Ziyang Chen, Shengyi Qian, Andrew Owens
ICCV 2023
project page  ·  paper  ·  code  ·  bibtex

Conditional Foley generation examples

Conditional Generation of Audio from Video via Foley Analogies
Yuexi Du, Ziyang Chen, Justin Salamon, Bryan Russell, Andrew Owens
CVPR 2023
project page  ·  paper  ·  code  ·  bibtex

Audio-visual anomaly detection overview

Self-Supervised Video Forensics by Audio-Visual Anomaly Detection
Chao Feng, Ziyang Chen, Andrew Owens
CVPR 2023 (Highlight -- 2.5% accept rate)
project page  ·  paper  ·  code  ·  bibtex

Stereo sound localization method overview

Sound Localization by Self-Supervised Time Delay Estimation
Ziyang Chen, David F. Fouhey, Andrew Owens
ECCV 2022
project page  ·  paper  ·  slides  ·  code  ·  bibtex

Mix and Localize method overview

Mix and Localize: Localizing Sound Sources in Mixtures
Xixi Hu*, Ziyang Chen*, Andrew Owens
CVPR 2022
project page  ·  paper  ·  slides  ·  code  ·  bibtex

Structure from Silence method overview

Structure from Silence: Learning Scene Structure from Ambient Sound
Ziyang Chen*, Xixi Hu*, Andrew Owens
CoRL 2021 (Oral)
project page  ·  paper  ·  slides  ·  code  ·  bibtex

Profession

Workshop Organizer

Reviewer

  • Conferences CVPR (2023–2026), ICCV (2023, 2025), ECCV (2024), NeurIPS (2024), SIGGRAPH (2024), WACV (2023)
  • Journals IJCV

Teaching Assistant

  • VE312 · Digital Integrated Circuits Fall 2018 · UM–SJTU Joint Institute · With Yaping Dan
  • VE311 · Electronic Circuits Summer 2019 · UM–SJTU Joint Institute · With Chang-Ching Tu