Fetching the paper…
Reading the bibliography…
A key function of auditory cognition is the association of characteristic sounds with their corresponding semantics over time.
“Echoic memory explored and applied,”
Terry Clark, · 1987
Earlier work this paper cites.
“Auditory sensory (” echoic”) memory dysfunction in schizophrenia.,”
Rael D Strous, Nelson Cowan, Walter Ritter, and Daniel C Javitt, · 1995
Earlier work this paper cites.
Human Memory
Gabriel A Radvansky, · 2005
Earlier work this paper cites.
“Visual concept recognition and localization via iterative introspection,”
Amir Rosenfeld and Shimon Ullman, · 2016
Earlier work this paper cites.
“Deep residual learning for image recognition,”
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, · 2016
Earlier work this paper cites.
“Audio set: An ontology and human-labeled dataset for audio events,”
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, et al., · 2017
Earlier work this paper cites.
“Autoscaler: Scale-attention networks for visual correspondence,”
Shenlong Wang, Linjie Luo, Ning Zhang, and Jia Li, · 2017
Earlier work this paper cites.
“Look closer to see better: Recurrent attention convolutional neural network for fine-grained image recognition,”
Jianlong Fu, Heliang Zheng, and Tao Mei, · 2017
Earlier work this paper cites.
“Learning to zoom: a saliency-based sampling layer for neural networks,”
Adria Recasens, Petr Kellnhofer, Simon Stent, Wojciech Matusik, and Antonio Torralba, · 2018
Earlier work this paper cites.
“mixup: Beyond empirical risk minimization,”
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz, · 2018
Earlier work this paper cites.
“A comparison of five multiple instance learning pooling functions for sound event detection with weak labeling,”
Yun Wang, Juncheng Li, and Florian Metze, · 2019
Earlier work this paper cites.
“Efficientnet: Rethinking model scaling for convolutional neural networks,”
Mingxing Tan and Quoc Le, · 2019
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, et al., · 2020
Cited alongside, same era.
“Object-centric learning with slot attention,”
Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, et al., · 2020
Cited alongside, same era.
“VGGSound: A large-scale audio-visual dataset,”
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman, · 2020
Cited alongside, same era.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, et al., · 2020
Cited alongside, same era.
“Sound event detection of weakly labelled data with cnn-transformer and automatic threshold optimization,”
Qiuqiang Kong, Yong Xu, Wenwu Wang, and Mark D Plumbley, · 2020
Cited alongside, same era.
“Conformer: Convolution-augmented transformer for speech recognition,”
“Perceiver: General perception with iterative attention,”
Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, et al., · 2021
Later among the works it cites.
“Training data-efficient image transformers & distillation through attention,”
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, et al., · 2021
Later among the works it cites.
“Swin transformer: Hierarchical vision transformer using shifted windows,”
Ze Liu, Yutong Lin, Yue Cao, Han Hu, et al., · 2021
Later among the works it cites.
“Polyvit: Co-training vision transformers on images, videos and audio,”
Valerii Likhosherstov, Anurag Arnab, Krzysztof Choromanski, Mario Lucic, et al., · 2021
Later among the works it cites.
“Rescaling egocentric vision: collection, pipeline and challenges for EPIC-KITCHENS-100,”
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, et al., · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, et al., · 2020
Cited alongside, same era.
“Convolution augmented transformer for semi-supervised sound event detection,”
Koichi Miyazaki, Tatsuya Komatsu, Tomoki Hayashi, Shinji Watanabe, et al., · 2020
Cited alongside, same era.
“Acoustic scene classification using deep residual networks with late fusion of separated high and low frequency paths,”
Mark D McDonnell and Wei Gao, · 2020
Cited alongside, same era.
“Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,”
Yuan Gong, Yu-An Chung, and James Glass, · 2021
Cited alongside, same era.
“Slow-fast auditory streams for audio recognition,”
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen, · 2021
Cited alongside, same era.
“Efficient training of audio transformers with patchout,”
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer, · 2021
Cited alongside, same era.
“Attention bottlenecks for multimodal fusion,”
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, et al., · 2021
Cited alongside, same era.
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, et al., · 2022
Closest in time.
“Learning the spectrogram temporal resolution for audio classification,”
Haohe Liu, Xubo Liu, Qiuqiang Kong, Wenwu Wang, and Mark D Plumbley, · 2022
Closest in time.
“Mae-ast: Masked autoencoding audio spectrogram transformer,”
Alan Baade, Puyuan Peng, and David Harwath, · 2022
Closest in time.
“Masked spectrogram prediction for self-supervised audio pre-training,”
Dading Chong, Helin Wang, Peilin Zhou, and Qingcheng Zeng, · 2022
Closest in time.
“Masked autoencoders that listen,”
Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, et al., · 2022
Closest in time.
“Mvitv2: Improved multiscale vision transformers for classification and detection,”
Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, et al., · 2022
Closest in time.
“Balanced multimodal learning via on-the-fly gradient modulation,”
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu, · 2022
Closest in time.