Fetching the paper…
Reading the bibliography…
Our ability to comprehend speech remains, to date, unrivaled by deep learning models.
A scale for the measurement of the psychological magnitude pitch
Stevens, S. S., Volkmann, J., and Newman, E. B · 1937
Earlier work this paper cites.
The relation of pitch to frequency: A revised scale
Stevens, S. S. and Volkmann, J · 1940
Earlier work this paper cites.
Fast readout of object identity from macaque inferior temporal cortex
Hung, C. P., Kreiman, G., Poggio, T., and DiCarlo, J. J · 2005
Earlier work this paper cites.
Decoding the visual and subjective contents of the human brain
Kamitani, Y. and Tong, F · 2005
Earlier work this paper cites.
A population-average, landmark-and surface-based (pals) atlas of human cerebral cortex
Van Essen, D. C · 2005
Earlier work this paper cites.
The motor theory of speech perception reviewed
Galantucci, B., Fowler, C. A., and Turvey, M. T · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A., Fernández, S., Gomez, F., and Schmidhuber, J · 2006
Earlier work this paper cites.
A component based noise correction method (compcor) for bold and perfusion based fmri
Behzadi, Y., Restom, K., Liau, J., and Liu, T. T · 2007
Earlier work this paper cites.
The cortical organization of speech processing
Hickok, G. and Poeppel, D · 2007
Earlier work this paper cites.
Discovering the false discovery rate
Benjamini, Y · 2010
Earlier work this paper cites.
Automatic parcellation of human cortical gyri and sulci using standard anatomical nomenclature
Destrieux, C., Fischl, B., Dale, A., and Halgren, E · 2010
Earlier work this paper cites.
The unique role of the visual word form area in reading
Dehaene, S. and Cohen, L · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al · 2011
Earlier work this paper cites.
Freesurfer
Fischl, B · 2012
Earlier work this paper cites.
Integration over multiple timescales in primary auditory cortex
David, S. V. and Shamma, S. A · 2013
Earlier work this paper cites.
Representational geometry: integrating cognition, computation, and the brain
Kriegeskorte, N. and Kievit, R. A · 2013
Earlier work this paper cites.
Phonetic feature encoding in human superior temporal gyrus
Mesgarani, N., Cheung, C., Johnson, K., and Chang, E. F · 2014
Cited alongside, same era.
Performance-optimized hierarchical models predict neural responses in higher visual cortex
Yamins, D. L., Hong, H., Cadieu, C. F., Solomon, E. A., Seibert, D., and DiCarlo, J. J · 2014
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., Chen, G., et al · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
Collobert, R., Puhrsch, C., and Synnaeve, G · 2016
Cited alongside, same era.
Toward an integration of deep learning and neuroscience
Marblestone, A. H., Wayne, G., and Kording, K. P · 2016
Cited alongside, same era.
fmriprep: a robust preprocessing pipeline for functional mri
Esteban, O., Markiewicz, C. J., Blair, R. W., Moodie, C. A., Isik, A. I., Erramuzpe, A., Kent, J. D., Goncalves, M., DuPre, E., Snyder, M., et al · 2019
Later among the works it cites.
Audio tagging with noisy labels and minimal supervision
Fonseca, E., Plakal, M., Font, F., Ellis, D. P., and Serra, X · 2019
Later among the works it cites.
Cascaded tuning to amplitude modulation for natural sound recognition
Koumura, T., Terashima, H., and Furukawa, S · 2019
Later among the works it cites.
Unidirectional monosynaptic connections from auditory areas to the primary visual cortex in the marmoset monkey
Majka, P., Rosa, M. G., Bai, S., Chan, J. M., Huo, B.-X., Jermakow, N., Lin, M. K., Takahashi, Y. S., Wolkowicz, I. H., Worthy, K. H., et al · 2019
Later among the works it cites.
librosa/librosa: 0.6.3, 2019
McFee, B., McVicar, M., Balke, S., Lostanlen, V., Thomé, C., Raffel, C., Lee, D., Kyungyun Lee, Nieto, O., Zalkow, F., Ellis, D., Battenberg, E., Yamamoto, R., Moore, J., Ziyao Wei, Bittner, R., Keunwoo Choi, Nullmightybofo, Friesch, P., Fabian-Robert Stöter, , Thassilo, Vollrath, M., Siddhartha Kumar Golu, Nehz, Waloschek, S., , Seth, Naktinis, R., Repetto, D., Hawthorne, C., and CJ Carr · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Using goal-driven deep learning models to understand sensory cortex
Yamins, D. L. and DiCarlo, J. J · 2016
Cited alongside, same era.
Freesound datasets: a platform for the creation of open audio datasets
Fonseca, E., Pons Puig, J., Favory, X., Font Corbera, F., Bogdanov, D., Ferraro, A., Oramas, S., Porter, A., and Serra, X · 2017
Cited alongside, same era.
Contextual modulation of primary visual cortex by auditory signals
Petro, L., Paton, A., and Muckli, L · 2017
Cited alongside, same era.
The spoken wikipedia corpus collection: Harvesting, alignment and an application to hyperlistening
Baumann, T., Köhn, A., and Hennig, F · 2018
Cited alongside, same era.
Out-of-the-box universal Romanization tool uroman
Hermjakob, U., May, J., and Knight, K · 2018
Cited alongside, same era.
A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy
Kell, A. J., Yamins, D. L., Shook, E. N., Norman-Haignere, S. V., and McDermott, J. H · 2018
Cited alongside, same era.
Encoding and decoding neuronal dynamics: Methodological framework to uncover the algorithms of cognition
King, J.-R., Gwilliams, L., Holdgraf, C., Sassenhagen, J., Barachant, A., Engemann, D., Larson, E., and Gramfort, A · 2018
Cited alongside, same era.
Later among the works it cites.
A 204-subject multimodal neuroimaging dataset to study language processing
Schoffelen, J.-M., Oostenveld, R., Lam, N. H., Uddén, J., Hultén, A., and Hagoort, P · 2019
Later among the works it cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, H., Mohamed, A., and Auli, M · 2020
Later among the works it cites.
Brain-optimized extraction of complex sound features that drive continuous auditory perception
Berezutskaya, J., Freudenburg, Z. V., Güçlü, U., van Gerven, M. A., and Ramsey, N. F · 2020
Later among the works it cites.
Language processing in brains and deep neural networks: computational convergence and its limits
Caucheteux, C. and King, J.-R · 2020
Later among the works it cites.
Estimating and interpreting nonlinear receptive field of sensory neural responses with deep neural network models
Keshishian, M., Akbari, H., Khalighinejad, B., Herrero, J. L., Mehta, A. D., and Mesgarani, N · 2020
Later among the works it cites.
Cortical response to naturalistic stimuli is largely predictable with deep neural networks
Khosla, M., Ngo, G. H., Jamison, K., Kuceyeski, A., and Sabuncu, M. R · 2020
Later among the works it cites.
Searching through functional space reveals distributed visual, auditory, and semantic coding in the human brain
Kumar, S., Ellis, C., O’Connell, T. P., Chun, M. M., and Turk-Browne, N. B · 2020
Later among the works it cites.
Big data suggest strong constraints of linguistic similarity on adult language learning
Schepens, J., van Hout, R., and Jaeger, T. F · 2020
Later among the works it cites.
Coupled training of sequence-to-sequence models for accented speech recognition
Unni, V., Joshi, N., and Jyothi, P · 2020
Later among the works it cites.
Learning noise invariant features through transfer learning for robust end-to-end speech recognition
Zhang, S., Do, C.-T., Doddipatla, R., and Renals, S · 2020
Later among the works it cites.