Fetching the paper…
Reading the bibliography…
Recently, diffusion models have achieved great success in mono-channel audio generation.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 1910
Earlier work this paper cites.
Panns: Large-scale pretrained audio neural networks for audio pattern recognition, 2020
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley · 1912
Earlier work this paper cites.
The generalized correlation method for estimation of time delay
Charles Knapp and Glifford Carter · 1976
Earlier work this paper cites.
The mpeg-4 video standard verification model
Thomas Sikora · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Pulse-code modulation–an overview
Stanley P Lipshitz and John Vanderkooy · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
A probabilistic model for robust localization based on a binaural auditory front-end
Tobias May, Steven Van De Par, and Armin Kohlrausch · 2010
Earlier work this paper cites.
Amazon mechanical turk: A research tool for organizations and information systems scholars
Kevin Crowston · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Karol J. Piczak · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation, 2016
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Siqi Liu, Zhenhai Zhu, Ning Ye, Sergio Guadarrama, and Kevin Murphy · 2017
Earlier work this paper cites.
Fr \ \backslash ’echet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2018
Earlier work this paper cites.
Self-supervised generation of spatial audio for 360 video, 2018
Pedro Morgado, Nuno Vasconcelos, Timothy Langlois, and Oliver Wang · 2018
Earlier work this paper cites.
Pyroomacoustics: A python package for audio room simulation and array processing algorithms
Robin Scheibler, Eric Bezzam, and Ivan Dokmanić · 2018
Earlier work this paper cites.
Audio-visual event localization in unconstrained videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu · 2018
Earlier work this paper cites.
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
2.5 d visual sound
Ruohan Gao and Kristen Grauman · 2019
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Cortical mechanisms of spatial hearing
Kiki van der Heijden, Josef P Rauschecker, Beatrice de Gelder, and Elia Formisano · 2019
Earlier work this paper cites.
Audio caption: Listen and tell
Mengyue Wu, Heinrich Dinkel, and Kai Yu · 2019
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Earlier work this paper cites.
A minimal personalization of dynamic binaural synthesis with mixed structural modeling and scattering delay networks
Michele Geronazzo, Jason Yves Tissieres, and Stefania Serafin · 2020
Earlier work this paper cites.
A sequence matching network for polyphonic sound event localization and detection
Thi Ngoc Tho Nguyen, Douglas L Jones, and Woon-Seng Gan · 2020
Earlier work this paper cites.
Sound event localization based on sound intensity vector refined by dnn-based denoising and source separation
Masahiro Yasuda, Yuma Koizumi, Shoichiro Saito, Hisashi Uematsu, and Keisuke Imoto · 2020
Cited alongside, same era.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Cited alongside, same era.
Binaural reproduction based on bilateral ambisonics and ear-aligned hrtfs
Zamir Ben-Hur, David Lou Alon, Ravish Mehra, and Boaz Rafaely · 2021
Cited alongside, same era.
An improved event-independent network for polyphonic sound event localization and detection
Yin Cao, Turab Iqbal, Qiuqiang Kong, Fengyan An, Wenwu Wang, and Mark D Plumbley · 2021
Cited alongside, same era.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal · 2021
Cited alongside, same era.
Binaural sound source distance estimation and localization for a moving listener
Daniel Aleksander Krause, Guillermo García-Barrios, Archontis Politis, and Annamaria Mesaros · 2023
Later among the works it cites.
Controllable music production with diffusion models and guidance gradients
Mark Levy, Bruno Di Giorgi, Floris Weers, Angelos Katharopoulos, and Tom Nickson · 2023
Later among the works it cites.
Audioldm: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Later among the works it cites.
Localizing concurrent sound sources with binaural microphones: A simulation study
Jakeh Orr, William Ebel, and Yan Gai · 2023
Later among the works it cites.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
gpurir: A python library for room impulse response simulation with gpu acceleration
David Diaz-Guerra, Antonio Miguel, and Jose R Beltran · 2021
Cited alongside, same era.
Geometry-aware multi-task learning for binaural audio generation from video
Rishabh Garg, Ruohan Gao, and Kristen Grauman · 2021
Cited alongside, same era.
Implicit hrtf modeling using temporal convolutional networks
Israel D Gebru, Dejan Marković, Alexander Richard, Steven Krenn, Gladstone A Butler, Fernando De la Torre, and Yaser Sheikh · 2021
Cited alongside, same era.
Time delay estimation for speaker localization using cnn-based parametrized gcc-phat features
Daniele Salvati, Carlo Drioli, Gian Luca Foresti, et al · 2021
Cited alongside, same era.
Goal-driven, neurobiological-inspired convolutional neural network models of human spatial hearing
Kiki van der Heijden and Siamak Mehrkanoon · 2021
Cited alongside, same era.
Visually informed binaural audio generation without binaural audios
Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, and Dahua Lin · 2021
Cited alongside, same era.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Cited alongside, same era.
Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation
Leigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie, and Tat-Seng Chua · 2023
Later among the works it cites.
I hear your true colors: Image guided audio generation
Roy Sheffer and Yossi Adi · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, et al · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation
Dongchao Yang, Jinchuan Tian, Xu Tan, Rongjie Huang, Songxiang Liu, Xuankai Chang, Jiatong Shi, Sheng Zhao, Jiang Bian, Xixin Wu, et al · 2023
Later among the works it cites.
Amphion: An open-source audio, music and speech generation toolkit
Xueyao Zhang, Liumeng Xue, Yuancheng Wang, Yicheng Gu, Xi Chen, Zihao Fang, Haopeng Chen, Lexiao Zou, Chaoren Wang, Jun Han, et al · 2023
Later among the works it cites.
Virtual reality technology
Grigore C Burdea and Philippe Coiffet · 2024
Closest in time.
Controllable generation with text-to-image diffusion models: A survey
Pu Cao, Feng Zhou, Qing Song, and Lu Yang · 2024
Closest in time.
Simple and controllable music generation
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez · 2024
Closest in time.
See-2-sound: Zero-shot spatial environment-to-spatial sound
Rishit Dagli, Shivesh Prakash, Robert Wu, and Houman Khosravani · 2024
Closest in time.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Closest in time.
Compa: Addressing the gap in compositional reasoning in audio-language models
Sreyan Ghosh, Ashish Seth, Sonal Kumar, Utkarsh Tyagi, Chandra Kiran Reddy Evuru, Ramaneswaran S, S Sakshi, Oriol Nieto, Ramani Duraiswami, and Dinesh Manocha · 2024
Closest in time.
A survey of sound source localization and detection methods and their applications
Gabriel Jekateryńczuk and Zbigniew Piotrowski · 2024
Closest in time.
Understanding diffusion objectives as the elbo with simple data augmentation
Diederik Kingma and Ruiqi Gao · 2024
Closest in time.
Cyclic learning for binaural audio generation and localization
Zhaojian Li, Bin Zhao, and Yuan Yuan · 2024
Closest in time.
Visually guided binaural audio generation with cross-modal consistency
Miao Liu, Jing Wang, Xinyuan Qian, and Xiang Xie · 2024
Closest in time.
Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi · 2024
Closest in time.
Follow-your-click: Open-domain regional image animation via short prompts
Yue Ma, Yingqing He, Hongfa Wang, Andong Wang, Chenyang Qi, Chengfei Cai, Xiu Li, Zhifeng Li, Heung-Yeung Shum, Wei Liu, et al · 2024
Closest in time.
Tango 2: Aligning diffusion-based text-to-audio generations through direct preference optimization
Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei-Ning Hsu, Rada Mihalcea, and Soujanya Poria · 2024
Closest in time.
Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, and Wenwu Wang · 2024
Closest in time.
Auditory cortex-inspired spectral attention modulation for binaural sound localization in hrtf mismatch
Waradon Phokhinanan, Nicolas Obin, and Sylvain Argentieri · 2024
Closest in time.
Starss23: An audio-visual dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel A Krause, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, et al · 2024
Closest in time.
V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Heng Wang, Jianbo Ma, Santiago Pascual, Richard Cartwright, and Weidong Cai · 2024
Closest in time.
Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners
Yazhou Xing, Yingqing He, Zeyue Tian, Xintao Wang, and Qifeng Chen · 2024
Closest in time.
A detailed audio-text data simulation pipeline using single-event sounds, 2024
Xuenan Xu, Xiaohang Xu, Zeyu Xie, Pingyue Zhang, Mengyue Wu, and Kai Yu · 2024
Closest in time.