Fetching the paper…
Reading the bibliography…
Machines that can represent and describe environmental soundscapes have practical potential, e.g., for audio tagging and captioning systems.
Aligning sentences in parallel corpora
Peter F. Brown, Jennifer C. Lai, and Robert L. Mercer. 1991 · 1991
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
A comparison of pivot methods for phrase-based statistical machine translation
Masao Utiyama and Hitoshi Isahara. 2007 · 2007
Earlier work this paper cites.
Pivot language approach for phrase-based statistical machine translation
Hua Wu and Haifeng Wang. 2007 · 2007
Earlier work this paper cites.
Machine hearing: An emerging field [exploratory dsp]
Richard F. Lyon. 2010 · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely. 2011 · 2011
Earlier work this paper cites.
Freesound technical demo
Frederic Font, Gerard Roma, and Xavier Serra. 2013 · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollar, and Larry Zitnick. 2014 · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello. 2014 · 2014
Earlier work this paper cites.
Video captions benefit everyone
Morton Ann Gernsbacher. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
ESC: Dataset for Environmental Sound Classification
Karol J. Piczak. 2015 · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. 2015 · 2015
Earlier work this paper cites.
Multimodal pivots for image caption translation
Julian Hitschler, Shigehiko Schamoni, and Stefan Riezler. 2016 · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class n-pair loss objective
Kihyuk Sohn. 2016 · 2016
Earlier work this paper cites.
A shared task on multimodal machine translation and crosslingual image description
Lucia Specia, Stella Frank, Khalil Sima’an, and Desmond Elliott. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Classifying environmental sounds using image recognition networks
Venkatesh Boddapati, Andrej Petef, Jim Rasmusson, and Lars Lundberg. 2017 · 2017
Cited alongside, same era.
Very deep convolutional neural networks for raw waveforms
Wei Dai, Chia Dai, Shuhui Qu, Juncheng Li, and Samarjit Das. 2017 · 2017
Cited alongside, same era.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017 · 2017
Cited alongside, same era.
Audio event and scene recognition: A unified approach using strongly and weakly labeled data
Anurag Kumar and Bhiksha Raj. 2017 · 2017
Cited alongside, same era.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. 2020 · 2020
Later among the works it cites.
Clotho: an audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen. 2020 · 2020
Later among the works it cites.
Esresnet: Environmental sound classification based on visual domain models
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel. 2021b · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
VATT: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hideki Nakayama and Noriki Nishida. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Scaling SGD batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg. 2017 · 2017
Cited alongside, same era.
Audio set classification with attention model: A probabilistic perspective
Qiuqiang Kong, Yong Xu, Wenwu Wang, and Mark D. Plumbley. 2018 · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani. 2018 · 2018
Cited alongside, same era.
Knowledge transfer from weakly labeled audio using convolutional neural network for sound events and scenes
Anurag Kumar, Maksim Khadkevich, and Christian Fügen. 2018 · 2018
Cited alongside, same era.
Adaptive pooling operators for weakly labeled sound event detection
Brian McFee, Justin Salamon, and Juan Pablo Bello. 2018 · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Closest in time.
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass. 2021 · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. 2021 · 2021
Closest in time.
Attention bottlenecks for multimodal fusion
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. 2021 · 2021
Closest in time.
Queryd: A video dataset with high-quality text and audio narrations
Andreea-Maria Oncescu, João F. Henriques, Yang Liu, Andrew Zisserman, and Samuel Albanie. 2021a · 2021
Closest in time.
Audio Retrieval with Natural Language Queries
Andreea-Maria Oncescu, A. Sophia Koepke, João F. Henriques, Zeynep Akata, and Samuel Albanie. 2021b · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Closest in time.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Closest in time.
Multimodal self-supervised learning of general audio representations
Luyu Wang, Pauline Luc, Adrià Recasens, Jean-Baptiste Alayrac, and Aäron van den Oord. 2021 · 2021
Closest in time.
Wav2clip: Learning robust audio representations from CLIP
Ho-Hsiang Wu, Prem Seetharaman, Kundan Kumar, and Juan Pablo Bello. 2021 · 2021
Closest in time.