Fetching the paper…
Reading the bibliography…
Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content.
Periphony: With-height sound reproduction
Michael A Gerzon. 1973 · 1973
Earlier work this paper cites.
Local sound field reproduction using digital signal processing
Ole Kirkeby, Philip A Nelson, Felipe Orduna-Bustamante, and Hareo Hamada. 1996 · 1996
Earlier work this paper cites.
Towards optimal soundfield representation. In Audio Engineering Society Convention 106 . Audio Engineering Society
Glenn Dickins and Rodney Kennedy. 1999 · 1999
Earlier work this paper cites.
Expressive sonification of footstep sounds
Roberto Bresin, Anna de Witt, Stefano Papetti, Marco Civolani, and Federico Fontana. 2010 · 2010
Earlier work this paper cites.
Multichannel stereophonic sound system with and without accompanying picture
BS Series. 2010 · 2010
Earlier work this paper cites.
Spatial sound for computer games and virtual reality
David Murphy and Flaithrí Neff. 2011 · 2011
Earlier work this paper cites.
Spatial audio
Francis Rumsey. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context. In Computer vision–ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13 . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba. 2016 · 2016
Earlier work this paper cites.
Temporal multimodal learning in audiovisual speech recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 3574–3582
Di Hu, Xuelong Li, et al · 2016
Earlier work this paper cites.
Ambient sound provides supervision for visual learning. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14 . Springer, 801–816
Andrew Owens, Jiajun Wu, Josh H McDermott, William T Freeman, and Antonio Torralba. 2016 · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio. In Arxiv
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alexander Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Earlier work this paper cites.
Look, listen and learn. In Proceedings of the IEEE international conference on computer vision . 609–617
Relja Arandjelovic and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
See, hear, and read: Deep aligned representations
Yusuf Aytar, Carl Vondrick, and Antonio Torralba. 2017 · 2017
Earlier work this paper cites.
Lip reading sentences in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6447–6456
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Surround by sound: A review of spatial audio recording and reproduction
Wen Zhang, Parasanga N Samarasinghe, Hanchi Chen, and Thushara D Abhayapala. 2017 · 2017
Earlier work this paper cites.
Objects that sound. In Proceedings of the European conference on computer vision (ECCV) . 435–451
Relja Arandjelovic and Andrew Zisserman. 2018 · 2018
Earlier work this paper cites.
The importance of spatial audio in modern games and virtual environments. In 2018 IEEE Games, Entertainment, Media Conference (GEM) . IEEE, 1–9
James Broderick, Jim Duggan, and Sam Redfern. 2018 · 2018
Earlier work this paper cites.
Scene-aware audio for 360 videos
Dingzeyu Li, Timothy R Langlois, and Changxi Zheng. 2018 · 2018
Earlier work this paper cites.
Self-Supervised Generation of Spatial Audio for 360 deg \deg Video. In Advances in Neural Information Processing Systems , Vol. 31
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang. 2018 · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features. In Proceedings of the European conference on computer vision (ECCV) . 631–648
Andrew Owens and Alexei A Efros. 2018 · 2018
Earlier work this paper cites.
Pyroomacoustics: A python package for audio room simulation and array processing algorithms. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 351–355
Robin Scheibler, Eric Bezzam, and Ivan Dokmanić. 2018 · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 4358–4366
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon. 2018 · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos. In Proceedings of the European conference on computer vision (ECCV) . 247–263
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018 · 2018
Earlier work this paper cites.
The sound of pixels. In Proceedings of the European conference on computer vision (ECCV) . 570–586
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018 · 2018
Earlier work this paper cites.
Visual to sound: Generating natural sound for videos in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3550–3558
Yipin Zhou, Zhaowen Wang, Chen Fang, Trung Bui, and Tamara L Berg. 2018 · 2018
Earlier work this paper cites.
Self-supervised moving vehicle tracking with stereo sound. In Proceedings of the IEEE/CVF international conference on computer vision . 7053–7062
Chuang Gan, Hang Zhao, Peihao Chen, David Cox, and Antonio Torralba. 2019 · 2019
Cited alongside, same era.
2.5 d visual sound. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 324–333
Ruohan Gao and Kristen Grauman. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Self-supervised audio-visual co-segmentation. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2357–2361
Andrew Rouditchenko, Hang Zhao, Chuang Gan, Josh McDermott, and Antonio Torralba. 2019 · 2019
Cited alongside, same era.
Dual attention matching for audio-visual event localization. In Proceedings of the IEEE/CVF international conference on computer vision . 6292–6300
Noise2music: Text-conditioned music generation with diffusion models
Qingqing Huang, Daniel S Park, Tao Wang, Timo I Denk, Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang, Jiahui Yu, Christian Frank, et al · 2023
Later among the works it cites.
Segment Anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 4015–4026
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick. 2023 · 2023
Later among the works it cites.
AudioGen: Textually Guided Audio Generation. In The Eleventh International Conference on Learning Representations
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2023 · 2023
Later among the works it cites.
Magic3D: High-Resolution Text-to-3D Content Creation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 300–309
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang. 2019 · 2019
Cited alongside, same era.
Talking face generation by adversarially disentangled audio-visual representation. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 9299–9306
Hang Zhou, Yu Liu, Ziwei Liu, Ping Luo, and Xiaogang Wang. 2019 · 2019
Cited alongside, same era.
Foley music: Learning to generate music from videos. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16 . Springer, 758–775
Chuang Gan, Deng Huang, Peihao Chen, Joshua B Tenenbaum, and Antonio Torralba. 2020 · 2020
Cited alongside, same era.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020 · 2020
Cited alongside, same era.
Discriminative sounding objects localization via self-supervised audiovisual matching
Di Hu, Rui Qian, Minyue Jiang, Xiao Tan, Shilei Wen, Errui Ding, Weiyao Lin, and Dejing Dou. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Unified multisensory perception: Weakly-supervised audio-visual video parsing. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part III 16 . Springer, 436–454
Yapeng Tian, Dingzeyu Li, and Chenliang Xu. 2020 · 2020
Cited alongside, same era.
Spatial Sound-History, Principle, Progress and Challenge
Bosun Xie. 2020 · 2020
Cited alongside, same era.
Align your gaussians: Text-to-4d with dynamic 3d gaussians and composed diffusion models
Huan Ling, Seung Wook Kim, Antonio Torralba, Sanja Fidler, and Karsten Kreis. 2023 · 2023
Later among the works it cites.
AudioLDM: text-to-audio generation with latent diffusion models. In Proceedings of the 40th International Conference on Machine Learning . 21450–21474
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley. 2023 · 2023
Later among the works it cites.
Videofusion: Decomposed diffusion models for high-quality video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10209–10218
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan. 2023 · 2023
Later among the works it cites.
DreamFusion: Text-to-3D using 2D Diffusion. In The Eleventh International Conference on Learning Representations
Ben Poole, Ajay Jain, Jonathan T. Barron, and Ben Mildenhall. 2023 · 2023
Later among the works it cites.
AV-RIR: Audio-Visual Room Impulse Response Estimation
Anton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya, and Dinesh Manocha. 2023 · 2023
Later among the works it cites.
I hear your true colors: Image guided audio generation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
Roy Sheffer and Yossi Adi. 2023 · 2023
Later among the works it cites.
CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
Zineng Tang, Ziyi Yang, Mahmoud Khademi, Yang Liu, Chenguang Zhu, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
A Neural Implicit Representation for the Image Stack: Depth, All in Focus, and High Dynamic Range
Chao Wang, Ana Serrano, Xingang Pan, Krzysztof Wolski, Bin Chen, Hans-Peter Seidel, Christian Theobalt, Karol Myszkowski, and Thomas Leimkühler. 2023 · 2023
Later among the works it cites.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu. 2023 · 2023
Later among the works it cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Homes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. 2024 · 2024
Closest in time.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2024 · 2024
Closest in time.
High-fidelity audio compression with improved rvqgan
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar. 2024 · 2024
Closest in time.
Efficient neural music generation
Max WY Lam, Qiao Tian, Tang Li, Zongyu Yin, Siyuan Feng, Ming Tu, Yuliang Ji, Rui Xia, Mingbo Ma, Xuchen Song, et al · 2024
Closest in time.
Av-nerf: Learning neural fields for real-world audio-visual scene synthesis
Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar, and Chenliang Xu. 2024 · 2024
Closest in time.
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. 2024 · 2024
Closest in time.
Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models
Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao. 2024 · 2024
Closest in time.
Naturalspeech: End-to-end text-to-speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al · 2024
Closest in time.
Any-to-any generation via composable diffusion
Zineng Tang, Ziyi Yang, Chenguang Zhu, Michael Zeng, and Mohit Bansal. 2024 · 2024
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. 2024 · 2024
Closest in time.
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
Jinlong Xue, Yayue Deng, Yingming Gao, and Ya Li. 2024 · 2024
Closest in time.
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data. In CVPR
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao. 2024 · 2024
Closest in time.
Lavss: Location-guided audio-visual spatial audio separation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 5508–5519
Yuxin Ye, Wenming Yang, and Yapeng Tian. 2024 · 2024
Closest in time.
Masked Audio Generation using a Single Non-Autoregressive Transformer
Alon Ziv, Itai Gat, Gael Le Lan, Tal Remez, Felix Kreuk, Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. 2024 · 2024
Closest in time.
Exploiting audio-visual consistency with partial supervision for spatial audio generation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 2056–2063
Yan-Bo Lin and Yu-Chiang Frank Wang. 2021 · 2063
Closest in time.