Fetching the paper…
Reading the bibliography…
Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthesis -- and a first-of-its-kind NeRF-based approach for multimodal learning.
Ray tracing volume densities
James T Kajiya and Brian P Von Herzen · 1984
Earlier work this paper cites.
Optical models for direct volume rendering
Nelson Max · 1995
Earlier work this paper cites.
The cipic hrtf database
V.R. Algazi, R.O. Duda, D.M. Thompson, and C. Avendano · 2001
Earlier work this paper cites.
Classifying images of materials: Achieving viewpoint and illumination independence
Manik Varma and Andrew Zisserman · 2002
Earlier work this paper cites.
Advanced audio coding (aac)
International Organization for Standardization · 2006
Earlier work this paper cites.
Mathematics of the discrete Fourier transform (DFT): with audio applications
Julius Orion Smith · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Xiph opus
Xiph.Org Foundation · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Acoustic classification and optimization for multi-modal rendering of real-world scenes
Carl Schissler, Christian Loftin, and Dinesh Manocha · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Generative modeling of audible shapes for object perception
Zhoutong Zhang, Jiajun Wu, Qiujia Li, Zhengjia Huang, James Traer, Josh H McDermott, Joshua B Tenenbaum, and William T Freeman · 2017
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Scene-aware audio for 360° videos
Dingzeyu Li, Timothy R. Langlois, and Changxi Zheng · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360° video
Pedro Morgado, Nuno Vasconcelos, Timothy R. Langlois, and Oliver Wang · 2018
Cited alongside, same era.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
Scene representation networks: Continuous 3d-structure-aware neural scene representations
Vincent Sitzmann, Michael Zollhöfer, and Gordon Wetzstein · 2019
Cited alongside, same era.
Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Soundspaces: Audio-visual navigation in 3d environments
Ir-gan: Room impulse response generator for far-field speech recognition
Anton Ratnarajah, Zhenyu Tang, and Dinesh Manocha · 2021
Later among the works it cites.
Neural synthesis of binaural speech from mono audio
Alexander Richard, Dejan Markovic, Israel D Gebru, Steven Krenn, Gladstone Butler, Fernando de la Torre, and Yaser Sheikh · 2021
Later among the works it cites.
Habitat 2.0: Training home assistants to rearrange their habitat
Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, Aaron Gokaslan, Vladimir Vondrus, Sameer Dharur, Franziska Meier, Wojciech Galuba, Angel Chang, Zsolt Kira, Vladlen Koltun, Jitendra Malik, Manolis Savva, and Dhruv Batra · 2021
Later among the works it cites.
Image2reverb: Cross-modal reverb impulse response synthesis
Nikhil Singh, Jeff Mentch, Jerry Ng, Matthew Beveridge, and Iddo Drori · 2021
Later among the works it cites.
NeRF − − -- : Neural radiance fields without known camera parameters
Zirui Wang, Shangzhe Wu, Weidi Xie, Min Chen, and Victor Adrian Prisacariu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Cited alongside, same era.
Differentiable volumetric rendering: Learning implicit 3d representations without 3d supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger · 2020
Cited alongside, same era.
Multiview neural surface reconstruction by disentangling geometry and appearance
Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman · 2020
Cited alongside, same era.
Scene-aware audio rendering via deep acoustic analysis
Zhenyu Tang, Nicholas J Bryan, Dingzeyu Li, Timothy R Langlois, and Dinesh Manocha · 2020
Cited alongside, same era.
Semantic audio-visual navigation
Changan Chen, Ziad Al-Halah, and Kristen Grauman · 2021
Cited alongside, same era.
Ad-nerf: Audio driven neural radiance fields for talking head synthesis
Yudong Guo, Keyu Chen, Sen Liang, Yong-Jin Liu, Hujun Bao, and Juyong Zhang · 2021
Cited alongside, same era.
Later among the works it cites.
Barf: Bundle-adjusting neural radiance fields
Chen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, and Simon Lucey · 2021
Later among the works it cites.
Localizing visual sounds the easy way
Shentong Mo and Pedro Morgado · 2022
Later among the works it cites.
Audio-visual segmentation
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong · 2022
Later among the works it cites.
Learning neural acoustic fields
Andrew Luo, Yilun Du, Michael J Tarr, Joshua B Tenenbaum, Antonio Torralba, and Chuang Gan · 2022
Later among the works it cites.
Inras: Implicit neural representation for audio scenes
Kun Su, Mingfei Chen, and Eli Shlizerman · 2022
Later among the works it cites.
Gwa: A large high-quality acoustic dataset for audio processing
Zhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, and Dinesh Manocha · 2022
Later among the works it cites.
Neural 3d video synthesis from multi-view video
Tianye Li, Mira Slavcheva, Michael Zollhöfer, Simon Green, Christoph Lassner, Changil Kim, Tanner Schmidt, Steven Lovegrove, Michael Goesele, Richard A. Newcombe, and Zhaoyang Lv · 2022
Later among the works it cites.
Deep impulse responses: Estimating and parameterizing filters with deep networks
Alexander Richard, Peter Dodds, and Vamsi Krishna Ithapu · 2022
Later among the works it cites.
Few-shot audio-visual learning of environment acoustics
Sagnik Majumder, Changan Chen, Ziad Al-Halah, and Kristen Grauman · 2022
Later among the works it cites.
Instant neural graphics primitives with a multiresolution hash encoding
Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller · 2022
Later among the works it cites.
Nerfstudio: A modular framework for neural radiance field development
Matthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li, Brent Yi, Justin Kerr, Terrance Wang, Alexander Kristoffersen, Jake Austin, Kamyar Salahi, Abhik Ahuja, David McAllister, and Angjoo Kanazawa · 2023
Closest in time.
Adobe audition
Adobe Inc · 2023
Closest in time.