Fetching the paper…
Reading the bibliography…
Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel coarse-to-fine audio reconstruction method.
Mechanisms and streams for processing of “what” and “where” in auditory cortex
Josef P Rauschecker and Biao Tian · 2000
Earlier work this paper cites.
Subdivisions of auditory cortex and processing streams in primates
Jon H Kaas and Troy A Hackett · 2000
Earlier work this paper cites.
Musical genre classification of audio signals
George Tzanetakis and Perry Cook · 2002
Earlier work this paper cites.
The neuroanatomical and functional organization of speech perception
Sophie K Scott and Ingrid S Johnsrude · 2003
Earlier work this paper cites.
The cortical organization of speech processing
Gregory Hickok and David Poeppel · 2007
Earlier work this paper cites.
Maps and streams in the auditory cortex: nonhuman primates illuminate human speech processing
Josef P Rauschecker and Sophie K Scott · 2009
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al · 2011
Earlier work this paper cites.
Reconstructing speech from human auditory cortex
Brian N Pasley, Stephen V David, Nima Mesgarani, Adeen Flinker, Shihab A Shamma, Nathan E Crone, Robert T Knight, and Edward F Chang · 2012
Earlier work this paper cites.
Speech reconstruction from human auditory cortex with deep neural networks
Minda Yang, Sameer A Sheth, Catherine A Schevon, Guy M McKhann II, and Nima Mesgarani · 2015
Earlier work this paper cites.
Brains on beats
Umut Güçlü, Jordy Thielen, Michael Hanke, and Marcel Van Gerven · 2016
Earlier work this paper cites.
Reconstructing the spectrotemporal modulations of real-life sounds from fmri response patterns
Roberta Santoro, Michelle Moerel, Federico De Martino, Giancarlo Valente, Kamil Ugurbil, Essa Yacoub, and Elia Formisano · 2017
Earlier work this paper cites.
The hierarchical cortical organization of human speech processing
Wendy A de Heer, Alexander G Huth, Thomas L Griffiths, Jack L Gallant, and Frédéric E Theunissen · 2017
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy
Alexander JE Kell, Daniel LK Yamins, Erica N Shook, Sam V Norman-Haignere, and Josh H McDermott · 2018
Earlier work this paper cites.
Reconstructing intelligible speech from the human auditory cortex
Akbari Hassan, Bahar Khalighinejad, Jose L Herrero, Ashesh D Mehta, and Nima Mesgarani · 2018
Cited alongside, same era.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2018
Cited alongside, same era.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Evidence of a predictive coding hierarchy in the human brain listening to speech
Charlotte Caucheteux, Alexandre Gramfort, and Jean-Rémi King · 2023
Later among the works it cites.
Jong-Yun Park, Mitsuaki Tsukamoto, Misato Tanaka, and Yukiyasu Kamitani · 2023
Later among the works it cites.
Music can be reconstructed from human auditory cortex activity using nonlinear decoding models
Ludovic Bellier, Anaïs Llorens, Déborah Marciano, Aysegul Gunduz, Gerwin Schalk, Peter Brunner, and Robert T Knight · 2023
Later among the works it cites.
Neural decoding of music from the eeg
Ian Daly · 2023
Later among the works it cites.
Brain2music: Reconstructing music from human brain activity
Timo I Denk, Yu Takagi, Takuya Matsuyama, Andrea Agostinelli, Tomoya Nakai, Christian Frank, and Shinji Nishimoto · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Panns: Large-scale pretrained audio neural networks for audio pattern recognition
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley · 2020
Cited alongside, same era.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Cited alongside, same era.
Taming visually guided sound generation
Vladimir Iashin and Esa Rahtu · 2021
Cited alongside, same era.
Toward a realistic model of speech processing in the brain with self-supervised learning
Juliette Millet, Charlotte Caucheteux, Yves Boubenec, Alexandre Gramfort, Ewan Dunbar, Christophe Pallier, Jean-Remi King, et al · 2022
Cited alongside, same era.
Self-supervised models of audio effectively explain human cortical responses to speech
Aditya R Vaidya, Shailee Jain, and Alexander G Huth · 2022
Cited alongside, same era.
Masked autoencoders that listen
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Later among the works it cites.
Synthesizing speech from ecog with a combination of transformer-based encoder and neural vocoder
Kai Shigemi, Shuji Komeiji, Takumi Mitsuhashi, Yasushi Iimura, Hiroharu Suzuki, Hidenori Sugano, Koichi Shinoda, Kohei Yatabe, and Toshihisa Tanaka · 2023
Later among the works it cites.
Braintalker: Low-resource brain-to-speech synthesis with transfer learning using wav2vec 2.0
Miseul Kim, Zhenyu Piao, Jihyun Lee, and Hong-Goo Kang · 2023
Later among the works it cites.
Dissecting neural computations in the human auditory pathway using deep neural networks for speech
Yuanning Li, Gopala K Anumanchipalli, Abdelrahman Mohamed, Peili Chen, Laurel H Carney, Junfeng Lu, Jinsong Wu, and Edward F Chang · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov · 2023
Later among the works it cites.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Haohe Liu, Qiao Tian, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D Plumbley · 2023
Later among the works it cites.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al · 2023
Later among the works it cites.
Audioldm: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Later among the works it cites.
Diffvoice: Text-to-speech with latent diffusion
Zhijun Liu, Yiwei Guo, and Kai Yu · 2023
Later among the works it cites.
A natural language fmri dataset for voxelwise encoding models
Amanda LeBel, Lauren Wagner, Shailee Jain, Aneesh Adhikari-Desai, Bhavin Gupta, Allyson Morgenthal, Jerry Tang, Lixiang Xu, and Alexander G Huth · 2023
Later among the works it cites.
A neural speech decoding framework leveraging deep learning and speech synthesis
Xupeng Chen, Ran Wang, Amirhossein Khalilian-Gourtani, Leyao Yu, Patricia Dugan, Daniel Friedman, Werner Doyle, Orrin Devinsky, Yao Wang, and Adeen Flinker · 2024
Closest in time.