Fetching the paper…
Reading the bibliography…
Binaural audio plays a significant role in constructing immersive augmented and virtual realities.
Role of spectral cues in median plane localization
Futoshi Asano, Yoiti Suzuki, and Toshio Sone · 1990
Earlier work this paper cites.
The dominant role of low-frequency interaural time differences in sound localization
Frederic L Wightman and Doris J Kistler · 1992
Earlier work this paper cites.
Hrft measurements of a kemar dummy-head microphone, 1994
Bill Gardner, Keith Martin, et al · 1994
Earlier work this paper cites.
Introduction to head-related transfer functions (hrtfs): Representations of hrtfs in time, frequency, and space
Corey I Cheng and Gregory H Wakefield · 1999
Earlier work this paper cites.
Creating interactive virtual acoustic environments
Lauri Savioja, Jyri Huopaniemi, Tapio Lokki, and Ritta Väänänen · 1999
Earlier work this paper cites.
3-d sound for virtual reality and multimedia
Durand R Begault and Leonard J Trejo · 2000
Earlier work this paper cites.
Rendering localized spatial audio in a virtual auditory space
Dmitry N. Zotkin, Ramani Duraiswami, and Larry S. Davis · 2004
Earlier work this paper cites.
Bayesian regularization and nonnegative deconvolution for room impulse response estimation
Yuanqing Lin and Daniel D Lee · 2006
Earlier work this paper cites.
Natural sound rendering for headphones: Integration of signal processing techniques
Kaushik Sunder, Jianjun He, Ee-Leng Tan, and Woon-Seng Gan · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Room impulse response interpolation using a sparse spatio-temporal representation of the sound field
Niccolo Antonello, Enzo De Sena, Marc Moonen, Patrick A Naylor, and Toon Van Waterschoot · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Building and evaluation of a real room impulse response dataset
Igor Szöke, Miroslav Skácel, Ladislav Mošner, Jakub Paliesek, and Jan Černockỳ · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Measurement of head-related transfer functions: A review
Song Li and Jürgen Peissig · 2020
Earlier work this paper cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Earlier work this paper cites.
Wavegrad: Estimating gradients for waveform generation
N. Chen, Y. Zhang, H. Zen, R.J. Weiss, M. Norouzi, and W. Chan · 2021
Cited alongside, same era.
Ilvr: Conditioning method for denoising diffusion probabilistic models
J. Choi, S. Kim, Y. Jeong, Y. Gwon, and S. Yoon · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
P. Dhariwal and A. Nichol · 2021
Cited alongside, same era.
Implicit hrtf modeling using temporal convolutional networks
Israel D Gebru, Dejan Marković, Alexander Richard, Steven Krenn, Gladstone A Butler, Fernando De la Torre, and Yaser Sheikh · 2021
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro · 2021
Cited alongside, same era.
A study on speech enhancement based on diffusion probabilistic model
Y. Lu, Y. Tsao, and S. Watanabe · 2021
S3rp: Self-supervised super-resolution and prediction for advection-diffusion process
C. Wang, K. Yeo, X. Jin, A. Codas, L. J. Klein, and B. Elmegreen · 2021
Later among the works it cites.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Later among the works it cites.
Infergrad: Improving diffusion models for vocoder by considering inference in training
Z. Chen, X. Tan, K. Wang, S. Pan, D. Mandic, L. He, and S. Zhao · 2022
Closest in time.
Vector quantized diffusion model for text-to-image synthesis
S. Gu, D. Chen, J. Bao, F. Wen, B. Zhang, D. Chen, L. Yuan, and B. Guo · 2022
Closest in time.
Cascaded diffusion models for high fidelity image generation
J. Ho, C. Saharia, W. Chan, Fleet D. J, M. Norouzi, and T. Salimans · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Improved denoising diffusion probabilistic models
A. Nichol and P. Dhariwal · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
A. Nichol, P. Dhariwal, A. Ramesh, P. Shyam, P. Mishkin, B. McGrew, I. Sutskever, and M. Chen · 2021
Cited alongside, same era.
Diffusevae: Efficient, controllable and high-fidelity generation from low-dimensional latents
K. Pandey, A. Mukherjee, P. Rai, and A. Kumar · 2021
Cited alongside, same era.
Grad-tts: A diffusion probabilistic model for text-to-speech
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M.A. Kudinov · 2021
Cited alongside, same era.
Neural synthesis of binaural speech from mono audio
Alexander Richard, Dejan Markovic, Israel D Gebru, Steven Krenn, Gladstone Alexander Butler, Fernando Torre, and Yaser Sheikh · 2021
Cited alongside, same era.
Image super-resolution via iterative refinement
C. Saharia, J. Ho, W. Chan, T. Salimans, D. J. Fleet, and M. Norouzi · 2021
Cited alongside, same era.
Diffusionclip: Text-guided diffusion models for robust image manipulation
G. Kim, T. Kwon, and J. C. Ye · 2022
Closest in time.
Bilateral denoising diffusion models
M.W.Y. Lam, J. Wang, R. Huang, D. Su, and D. Yu · 2022
Closest in time.
Priorgrad: Improving conditional denoising diffusion models with data-driven adaptive prior
S. Lee, H. Kim, C. Shin, X. Tan, C. Liu, Q. Meng, T. Qin, W. Chen, S. Yoon, and T.Y. Liu · 2022
Closest in time.
Conditional diffusion probabilistic model for speech enhancement
Y. Lu, Z. Wang, S. Watanabe, A. Richard, C. Yu, Y. Tsao, and S. Zhao · 2022
Closest in time.
Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal attention
Kranti Kumar Parida, Siddharth Srivastava, and Gaurav Sharma · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
Deep impulse responses: Estimating and parameterizing filters with deep networks
Alexander Richard, Peter Dodds, and Vamsi Krishna Ithapu · 2022
Closest in time.
Naturalspeech: End-to-end text to speech synthesis with human-level quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, et al · 2022
Closest in time.
Speech enhancement with score-based generative models in the complex stft domain
S. Welker, J. Richter, and T. Gerkmann · 2022
Closest in time.
Tackling the generative learning trilemma with denoising diffusion gans
Z. Xiao, K. Kreis, and A. Vahdat · 2022
Closest in time.