Fetching the paper…
Reading the bibliography…
We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment.
New method of measuring reverberation time
Manfred R. Schroeder · 1965
Earlier work this paper cites.
Image method for efficiently simulating small‐room acoustics
Jont B. Allen and David A. Berkley · 1979
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Comparison of different impulse response measurement techniques
Guy-Bart Stan, Jean-Jacques Embrechts, and Dominique Archambeau · 2002
Earlier work this paper cites.
A beam tracing method for interactive architectural acoustics
Thomas Funkhouser, Nicolas Tsingos, Ingrid Carlbom, Gary Elko, Mohan Sondhi, James E West, Gopal Pingali, Patrick Min, and Addy Ngan · 2004
Earlier work this paper cites.
Acoustic modeling using the digital waveguide mesh
D.T. Murphy, Antti Kelloniemi, Jack Mullen, and Simon Shelley · 2007
Earlier work this paper cites.
Impulse response measurement techniques and their applicability in the real world
Martin Holters, Tobias Corbach, and Udo Zölzer · 2009
Earlier work this paper cites.
Modeling of complex geometries and boundary conditions in finite difference/finite volume time domain room acoustics simulation
Stefan Bilbao · 2013
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Overview of geometrical room acoustic modeling techniques
Lauri Savioja and U Peter Svensson · 2015
Earlier work this paper cites.
Interactive sound propagation with bidirectional path tracing
Chunxiao Cao, Zhong Ren, Carl Schissler, Dinesh Manocha, and Kun Zhou · 2016
Earlier work this paper cites.
Estimation of room acoustic parameters: The ACE challenge
James Eaton, Nikolay Gaubitch, Allistair Moore, and Patrick Naylor · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond Y.K. Lau, Zhen Wang, and Stephen Paul Smolley · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
More than 50 years of artificial reverberation
Vesa Välimäki, Julian Parker, Lauri Savioja, Julius O. Smith, and Jonathan Abel · 2016
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Earlier work this paper cites.
Audio-visual speech enhancement using multimodal deep convolutional neural networks
Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Yu Tsao, Hsiu-Wen Chang, and Hsin-Min Wang · 2017
Earlier work this paper cites.
Blind estimation of the reverberation fingerprint of unknown acoustic environments
Prateek Murgai, Mark Rau, and Jean-Marc Jot · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Earlier work this paper cites.
Blind reverberation time estimation using a convolutional neural network
Hannes Gamper and Ivan J Tashev · 2018
Earlier work this paper cites.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Earlier work this paper cites.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Earlier work this paper cites.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2018
Cited alongside, same era.
Gibson Env: real-world perception for embodied agents
Fei Xia, Amir R. Zamir, Zhi-Yang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Cited alongside, same era.
Joint estimation of reverberation time and early-to-late reverberation ratio from single-channel speech signals
Feifei Xiong, Stefan Goetze, Birger Kollmeier, and Bernd T Meyer · 2018
Cited alongside, same era.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Cited alongside, same era.
A neural vocoder with hierarchical generation of amplitude and phase spectra for statistical parametric speech synthesis
Yang Ai and Zhen-Hua Ling · 2019
Cited alongside, same era.
Phase-aware speech enhancement with deep complex u-net
Single-channel blind direct-to-reverberation ratio estimation using masking
Wolfgang Mack, Shuwen Deng, and Emanuël AP Habets · 2020
Later among the works it cites.
An overview of deep-learning-based audio-visual speech enhancement and separation
Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng Yu, Dong Yu, and Jesper Jensen · 2020
Later among the works it cites.
Learning representations from audio-visual spatial alignment
Pedro Morgado, Yi Li, and Nuno Vasconcelos · 2020
Later among the works it cites.
Audio-visual speech enhancement using conditional variational auto-encoders
Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-PIneda, Laurent Girin, and Radu Horaud · 2020
Later among the works it cites.
Blind arbitrary reverb matching
Andy Sarroff and Roth Michaels · 2020
Later among the works it cites.
Simulation-based auralization of room acoustics
Lauri Savioja and Ning Xiang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hyeong-Seok Choi, Jang-Hyun Kim, Jaesung Huh, Adrian Kim, Jung-Woo Ha, and Kyogu Lee · 2019
Cited alongside, same era.
2.5d visual sound
Ruohan Gao and Kristen Grauman · 2019
Cited alongside, same era.
Real-time estimation of reverberation time for selection of suitable binaural room impulse responses
Florian Klein, Annika Neidhardt, and Marius Seipel · 2019
Cited alongside, same era.
Estimation of late reverberation characteristics from a single two-dimensional environmental image using convolutional neural networks
Homare Kon and Hideki Koike · 2019
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brebisson, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Mosnet: Deep learning based objective assessment for voice conversion
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang · 2019
Cited alongside, same era.
The Replica dataset: A digital replica of indoor spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al · 2019
Cited alongside, same era.
Later among the works it cites.
Acoustic matching by embedding impulse responses
Jiaqi Su, Zeyu Jin, and Adam Finkelstein · 2020
Later among the works it cites.
Dereverberation using joint estimation of dry speech signal and acoustic system
Sanna Wager, Keunwoo Choi, and Simon Durand · 2020
Later among the works it cites.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Linagzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Semantic audio-visual navigation
Changan Chen, Ziad Al-Halah, and Kristen Grauman · 2021
Later among the works it cites.
Learning to set waypoints for audio-visual navigation
Changan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao, Santhosh Kumar Ramakrishnan, and Kristen Grauman · 2021
Later among the works it cites.
Learning audio-visual dereverberation
Changan Chen, Wei Sun, David Harwath, and Kristen Grauman · 2021
Later among the works it cites.
VisualVoice: Audio-visual speech separation with cross-modal consistency
Ruohan Gao and Kristen Grauman · 2021
Later among the works it cites.
Anticipative Video Transformer
Rohit Girdhar and Kristen Grauman · 2021
Later among the works it cites.
Reverb conversion of mixed vocal tracks using an end-to-end convolutional deep neural network
Junghyun Koo, Seungryeol Paik, and Kyogu Lee · 2021
Later among the works it cites.
Space-time crop & attend: Improving cross-modal video representation learning
Mandela Patrick, Yuki M Asano, Bernie Huang, Ishan Misra, Florian Metze, Joao Henriques, and Andrea Vedaldi · 2021
Later among the works it cites.
Image2reverb: Cross-modal reverb impulse response synthesis
Nikhil Singh, Jeff Mentch, Jerry Ng, Matthew Beveridge, and Iddo Drori · 2021
Later among the works it cites.
Filtered noise shaping for time domain room impulse response estimation from reverberant speech
Christian Steinmetz, Vamsi Krishna Ithapu, and Paul Calamia · 2021
Later among the works it cites.
The right to talk: An audio-visual transformer approach
Thanh-Dat Truong, Chi Nhan Duong, The De Vu, Hoang Anh Pham, Bhiksha Raj, Ngan Le, and Khoa Luu · 2021
Later among the works it cites.
Creation of auditory augmented reality using a position-dynamic binaural synthesis system—technical components, psychoacoustic needs, and perceptual evaluation
Stephan Werner, Florian Klein, Annika Neidhardt, Ulrike Sloma, Christian Schneiderwind, and Karlheinz Brandenburg · 2021
Later among the works it cites.
Few-shot audio-visual learning of environment acoustics
Sagnik Majumder, Changan Chen, Ziad Al-Halah, and Kristen Grauman · 2022
Closest in time.
Sound adversarial audio-visual navigation
Yinfeng Yu, Wenbing Huang, Fuchun Sun, Changan Chen, Yikai Wang, and Xiaohong Liu · 2022
Closest in time.