Fetching the paper…
Reading the bibliography…
Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise.
Procédés et systèmes d’enregistrement et de reproduction sonores en trois dimensions
Daniel Courville and Ambisonic Studio · 1994
Earlier work this paper cites.
Integration of spatial sound in immersive virtual environments an experimental study on effects of spatial sound on presence
Sandra Poeschl, Konstantin Wall, and Nicola Doering · 2013
Earlier work this paper cites.
The Foley grail: The art of performing sound for film, games, and animation
Vanessa Theme Ament · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visually indicated sounds
Andrew Owens, Phillip Isola, Josh McDermott, Antonio Torralba, Edward H Adelson, and William T Freeman · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel P. W. Ellis, Jort F. Gemmeke, Aren Jansen, R. Channing Moore, Manoj Plakal, Devin Platt, Rif A. Saurous, Bryan Seybold, Malcolm Slaney, Ron J. Weiss, and Kevin Wilson · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2017
Earlier work this paper cites.
What do different evaluation metrics tell us about saliency models?
Zoya Bylinskii, Tilke Judd, Aude Oliva, Antonio Torralba, and Frédo Durand · 2018
Earlier work this paper cites.
Cube padding for weakly-supervised saliency prediction in 360 videos
Hsien-Tzu Cheng, Chun-Hung Chao, Jin-Dong Dong, Hao-Kai Wen, Tyng-Luh Liu, and Min Sun · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360 video
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg · 2018
Earlier work this paper cites.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
Earlier work this paper cites.
Towards generating ambisonics using audio-visual cue for virtual reality
Aakanksha Rana, Cagri Ozcinar, and Aljosa Smolic · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
Fr \ \backslash ’echet audio distance: A reference-free metric for evaluating music enhancement algorithms
Dominik Roblek, Kevin Kilgour, Matt Sharifi, and Mauricio Zuluaga · 2019
Earlier work this paper cites.
User experience of stereo and spatial audio in 360° live music videos
Jukka Holm, Kaisa Väänänen, and Anas Battah · 2020
Earlier work this paper cites.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Nvasconcelos · 2020
Earlier work this paper cites.
Semantic object prediction and spatial sound super-resolution with binaural sounds
Arun Balajee Vasudevan, Dengxin Dai, and Luc Van Gool · 2020
Earlier work this paper cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Taming visually guided sound generation
Vladimir Iashin and Esa Rahtu · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi · 2021
Cited alongside, same era.
Ambisonics: A practical 3D audio theory for recording, studio production, sound reinforcement, and virtual reality
I hear your true colors: Image guided audio generation
Roy Sheffer and Yossi Adi · 2023
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Franz Zotter and Matthias Frank · 2021
Cited alongside, same era.
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman · 2022
Cited alongside, same era.
Spatial audio in 360° videos: does it influence visual attention?
Amit Hirway, Yuansong Qiao, and Niall Murray · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2022
Cited alongside, same era.
Efficient training of audio transformers with patchout
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer · 2022
Cited alongside, same era.
Panoramic vision transformer for saliency detection in 360 ∘ \circ videos
Heeseung Yun, Sehun Lee, and Gunhee Kim · 2022
Cited alongside, same era.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Later among the works it cites.
Evaluating visual attention and qoe for 360° videos with non-spatial and spatial audio
Amit Hirway, Yuansong Qiao, and Niall Murray · 2024
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar · 2024
Later among the works it cites.
Voiceldm: Text-to-speech with environmental context
Yeonghyeon Lee, Inmo Yeon, Juhan Nam, and Joon Son Chung · 2024
Later among the works it cites.
Enhancing spatial audio generation with source separation and channel panning loss
Wootaek Lim and Juhan Nam · 2024
Later among the works it cites.
Visually guided binaural audio generation with cross-modal consistency
Miao Liu, Jing Wang, Xinyuan Qian, and Xiang Xie · 2024
Later among the works it cites.
Masked generative video-to-audio transformers with enhanced synchronicity
Santiago Pascual, Chunghsin Yeh, Ioannis Tsiamas, and Joan Serrà · 2024
Later among the works it cites.
Starss23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel A Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, et al · 2024
Later among the works it cites.
V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, and Weidong Cai · 2024
Later among the works it cites.
Codec-superb: An in-depth analysis of sound codec models
Haibin Wu, Ho-Lam Chung, Yi-Cheng Lin, Yuan-Kuei Wu, Xuanjun Chen, Yu-Chi Pai, Hsiu-Hsuan Wang, Kai-Wei Chang, Alexander H Liu, and Hung-yi Lee · 2024
Later among the works it cites.
Dualspeech: Enhancing speaker-fidelity and text-intelligibility through dual classifier-free guidance
Jinhyeok Yang, Junhyeok Lee, Hyeong-Seok Choi, Seunghoon Ji, Hyeongju Kim, and Juheon Lee · 2024
Later among the works it cites.
Masked audio generation using a single non-autoregressive transformer
Alon Ziv, Itai Gat, Gael Le Lan, Tal Remez, Felix Kreuk, Jade Copet, Alexandre Défossez, Gabriel Synnaeve, and Yossi Adi · 2024
Later among the works it cites.
Immersediffusion: A generative spatial audio latent diffusion model
Mojtaba Heydari, Mehrez Souden, Bruno Conejo, and Joshua Atkins · 2025
Closest in time.
Cotracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht · 2025
Closest in time.
Diff-sage: End-to-end spatial audio generation using diffusion models
Saksham Singh Kushwaha, Jianbo Ma, Mark RP Thomas, Yapeng Tian, and Avery Bruni · 2025
Closest in time.