Fetching the paper…
Reading the bibliography…
Imagine hearing a dog bark and turning toward the sound only to see a parked car, while the real, silent dog sits elsewhere.
Hrft measurements of a kemar dummy-head microphone, 1994
Bill Gardner, Keith Martin, et al · 1994
Earlier work this paper cites.
The cocktail party problem
Simon Haykin and Zhe Chen · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A phenomenological model of the synapse between the inner hair cell and auditory nerve: long-term adaptation with power-law dynamics
Muhammad SA Zilany, Ian C Bruce, Paul C Nelson, and Laurel H Carney · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Sound texture perception via statistics of the auditory periphery: evidence from sound synthesis
Josh H McDermott and Eero P Simoncelli · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context, 2015
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2015
Earlier work this paper cites.
A method for stochastic optimization
Diederik Kinga, Jimmy Ba Adam, et al · 2015
Earlier work this paper cites.
Soundnet: Learning sound representations from unlabeled video
Yusuf Aytar, Carl Vondrick, and Antonio Torralba · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Scaper: A library for soundscape synthesis and augmentation
Justin Salamon, Duncan MacConnell, Mark Cartwright, Peter Li, and Juan Pablo Bello · 2017
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelovic and Andrew Zisserman · 2017
Earlier work this paper cites.
Headphone screening to facilitate web-based auditory experiments
Kevin JP Woods, Max H Siegel, James Traer, and Josh H McDermott · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Earlier work this paper cites.
The locata challenge data corpus for acoustic source localization and tracking
Heinrich W Löllmann, Christine Evers, Alexander Schmidt, Heinrich Mellmann, Hendrik Barfuss, Patrick A Naylor, and Walter Kellermann · 2018
Earlier work this paper cites.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Earlier work this paper cites.
Cortical mechanisms of spatial hearing
Kiki van der Heijden, Josef P Rauschecker, Beatrice de Gelder, and Elia Formisano · 2019
Earlier work this paper cites.
Multi-task self-supervised learning for human activity detection
Aaqib Saeed, Tanir Ozcelebi, and Johan Lukkien · 2019
Earlier work this paper cites.
Deformable convnets v2: More deformable, better results
Xizhou Zhu, Han Hu, Stephen Lin, and Jifeng Dai · 2019
Earlier work this paper cites.
Semantic understanding of scenes through the ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Tete Xiao, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2019
Earlier work this paper cites.
Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity
Ethan Manilow, Gordon Wichern, Prem Seetharaman, and Jonathan Le Roux · 2019
Earlier work this paper cites.
Learning sound event classifiers from web audio with noisy labels
Eduardo Fonseca, Manoj Plakal, Daniel PW Ellis, Frederic Font, Xavier Favory, and Xavier Serra · 2019
Earlier work this paper cites.
Sound event detection in domestic environments with weakly labeled data and soundscape synthesis
Nicolas Turpault, Romain Serizel, Ankit Parag Shah, and Justin Salamon · 2019
Earlier work this paper cites.
Learning sound event classifiers from web audio with noisy labels
Eduardo Fonseca, Manoj Plakal, Daniel PW Ellis, Frederic Font, Xavier Favory, and Xavier Serra · 2019
Earlier work this paper cites.
Illusory sound texture reveals multi-second statistical completion in auditory scene analysis
Richard McWalter and Josh H McDermott · 2019
Earlier work this paper cites.
Multiple sound sources localization from coarse to fine
Rui Qian, Di Hu, Heinrich Dinkel, Mengyue Wu, Ning Xu, and Weiyao Lin · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman · 2020
Earlier work this paper cites.
Sep-stereo: Visually guided stereophonic audio generation by associating source separation
Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang, and Ziwei Liu · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Cited alongside, same era.
Perceptual inference, learning, and attention in a multisensory world
Uta Noppeney · 2021
Cited alongside, same era.
Self-supervised soft obstacle detection for safe navigation of visually impaired people
George Dimas, Eirini Cholopoulou, and Dimitris K Iakovidis · 2021
Cited alongside, same era.
Space-time memory network for sounding object localization in videos
Sizhe Li, Yapeng Tian, and Chenliang Xu · 2021
Cited alongside, same era.
Pvt v2: Improved baselines with pyramid vision transformer
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2022
Later among the works it cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Audio-visual contrastive learning with temporal self-supervision
Simon Jenni, Alexander Black, and John Collomosse · 2023
Later among the works it cites.
Robots understanding contextual information in human-centered environments using weakly supervised mask data distillation
Daniel Dworakowski, Angus Fung, and Goldie Nejat · 2023
Later among the works it cites.
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning
Sangho Lee, Jiwan Chung, Youngjae Yu, Gunhee Kim, Thomas Breuel, Gal Chechik, and Yale Song · 2021
Cited alongside, same era.
A cappella: Audio-visual singing voice separation
Juan F Montesinos, Venkatesh S Kadandale, and Gloria Haro · 2021
Cited alongside, same era.
Localizing visual sounds the hard way
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman · 2021
Cited alongside, same era.
Audio-visual localization by synthetic acoustic image generation
Valentina Sanguineti, Pietro Morerio, Alessio Del Bue, and Vittorio Murino · 2021
Cited alongside, same era.
Deep neural network models reveal interplay of peripheral coding and stimulus statistics in pitch perception
Mark R Saddler, Ray Gonzalez, and Josh H McDermott · 2021
Cited alongside, same era.
When pigs fly: Contextual reasoning in synthetic and natural scenes
Philipp Bomatter, Mengmi Zhang, Dimitar Karev, Spandan Madan, Claire Tseng, and Gabriel Kreiman · 2021
Cited alongside, same era.
Enhancing geometric factors in model learning and inference for object detection and instance segmentation
Zhaohui Zheng, Ping Wang, Dongwei Ren, Wei Liu, Rongguang Ye, Qinghua Hu, and Wangmeng Zuo · 2021
Cited alongside, same era.
Mavil: Masked audio-video learners
Po-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali, Yanghao Li, Shang-Wen Li, Gargi Ghosh, Jitendra Malik, Christoph Feichtenhofer, et al · 2023
Later among the works it cites.
Unsupervised sound localization via iterative contrastive learning
Yan-Bo Lin, Hung-Yu Tseng, Hsin-Ying Lee, Yen-Yu Lin, and Ming-Hsuan Yang · 2023
Later among the works it cites.
Catr: Combinatorial-dependence audio-queried transformer for audio-visual video segmentation
Kexin Li, Zongxin Yang, Lei Chen, Yi Yang, and Jun Xiao · 2023
Later among the works it cites.
Robust self-supervised multi-instance learning with structure awareness
Yejiang Wang, Yuhai Zhao, Zhengkui Wang, and Meixia Wang · 2023
Later among the works it cites.
Languagebind: Extending video-language pretraining to n-modality by language-based semantic alignment, 2023
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, Wang HongFa, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Cai Wan Zhang, Zhifeng Li, Wei Liu, and Li Yuan · 2023
Later among the works it cites.
Learning audio-visual source localization via false negative aware contrastive learning
Weixuan Sun, Jiayi Zhang, Jianyuan Wang, Zheyuan Liu, Yiran Zhong, Tianpeng Feng, Yandong Guo, Yanhao Zhang, and Nick Barnes · 2023
Later among the works it cites.
Sound source localization is all about cross-modal alignment
Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh, Hanspeter Pfister, and Joon Son Chung · 2023
Later among the works it cites.
Imagebind: One embedding space to bind them all
Rohit Girdhar, Alaaeldin El-Nouby, Zhuang Liu, Mannat Singh, Kalyan Vasudev Alwala, Armand Joulin, and Ishan Misra · 2023
Later among the works it cites.
Reproducible scaling laws for contrastive language-image learning
Mehdi Cherti, Romain Beaumont, Ross Wightman, Mitchell Wortsman, Gabriel Ilharco, Cade Gordon, Christoph Schuhmann, Ludwig Schmidt, and Jenia Jitsev · 2023
Later among the works it cites.
Separating the" chirp" from the" chat": Self-supervised visual grounding of sound and language
Mark Hamilton, Andrew Zisserman, John R Hershey, and William T Freeman · 2024
Later among the works it cites.
Avsegformer: Audio-visual segmentation with transformer
Shengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang, and Tong Lu · 2024
Later among the works it cites.
Unveiling and mitigating bias in audio visual segmentation
Peiwen Sun, Honggang Zhang, and Di Hu · 2024
Later among the works it cites.
Audio-visual segmentation with semantics
Jinxing Zhou, Xuyang Shen, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, et al · 2024
Later among the works it cites.
Ref-avs: Refer and segment objects in audio-visual scenes
Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li, Honggang Zhang, and Di Hu · 2024
Later among the works it cites.
Peavs: Perceptual evaluation of audio-visual synchrony grounded in viewers’ opinion scores, 2024
Lucas Goncalves, Prashant Mathur, Chandrashekhar Lavania, Metehan Cekic, Marcello Federico, and Kyu J. Han · 2024
Later among the works it cites.
Unraveling instance associations: A closer look for audio-visual segmentation
Yuanhong Chen, Yuyuan Liu, Hu Wang, Fengbei Liu, Chong Wang, Helen Frazer, and Gustavo Carneiro · 2024
Later among the works it cites.
Models optimized for real-world tasks reveal the task-dependent necessity of precise temporal coding in hearing
Mark R Saddler and Josh H McDermott · 2024
Later among the works it cites.
Da Mu, Zhicheng Zhang, Haobo Yue, Zehao Wang, Jin Tang, and Jianqin Yin · 2024
Later among the works it cites.
Depth anything: Unleashing the power of large-scale unlabeled data
Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Hengshuang Zhao · 2024
Later among the works it cites.
Human-in-the-loop robot learning for smart manufacturing: A human-centric perspective
Hongpeng Chen, Shufei Li, Junming Fan, Anqing Duan, Chenguang Yang, David Navarro-Alarcon, and Pai Zheng · 2025
Closest in time.
Uni-retrieval: A multi-style retrieval framework for stem’s education, 2025
Yanhao Jia, Xinyi Wu, Hao Li, Qinglin Zhang, Yuxiao Hu, Shuai Zhao, and Wenqi Fan · 2025
Closest in time.
Freestyleret: Retrieving images from style-diversified queries
Hao Li, Yanhao Jia, Jin Peng, Zesen Cheng, Kehan Li, Jialu Sui, Chang Liu, and Li Yuan · 2025
Closest in time.
Toward interactive sound source localization: Better align sight and sound!
Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh, Hanspeter Pfister, and Joon Son Chung · 2025
Closest in time.
Stereo sound event localization and detection with onscreen/offscreen classification
Kazuki Shimada, Archontis Politis, Iran R Roman, Parthasaarathy Sudarsanam, David Diaz-Guerra, Ruchi Pandey, Kengo Uchida, Yuichiro Koyama, Naoya Takahashi, Takashi Shibuya, et al · 2025
Closest in time.
A two-step learning framework for enhancing sound event localization and detection
Hogeon Yu · 2025
Closest in time.