Fetching the paper…
Reading the bibliography…
As humans, we navigate a multimodal world, building a holistic understanding from all our senses.
The origins of intelligence in children
Jean Piaget and Margaret Trans Cook · 1952
Earlier work this paper cites.
Scripts, plans, and knowledge
Roger C. Schank and Robert P. Abelson · 1975
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
Situated knowledges: The science question in feminism and the privilege of partial perspective
Donna Haraway · 1988
Earlier work this paper cites.
Neural darwinism: selection and reentrant signaling in higher brain function
Gerald M Edelman · 1993
Earlier work this paper cites.
Crime in black and white: The violent, scary world of local news
Franklin D Gilliam Jr, Shanto Iyengar, Adam Simon, and Oliver Wright · 1996
Earlier work this paper cites.
Children’s language learning: An interactionist perspective
Robin S Chapman · 2000
Earlier work this paper cites.
Overrepresentation and underrepresentation of african americans and latinos as lawbreakers on television news
Travis L Dixon and Daniel Linz · 2000
Earlier work this paper cites.
The development of embodied cognition: Six lessons from babies
Linda Smith and Michael Gasser · 2005
Earlier work this paper cites.
Crime news and racialized beliefs: Understanding the relationship between local news viewing and perceptions of african americans and crime
Travis L Dixon · 2008
Earlier work this paper cites.
Exploring the gender divide on youtube: An analysis of the creation and reception of vlogs
Heather Molyneaux, Susan O’Donnell, Kerri Gibson, Janice Singer, et al · 2008
Earlier work this paper cites.
Clustered synopsis of surveillance video
Yael Pritch, Sarit Ratovitch, Avishai Hendel, and Shmuel Peleg · 2009
Earlier work this paper cites.
An alternative view of privacy on facebook
Christian Fuchs · 2011
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme · 2013
Earlier work this paper cites.
White news: Why local news programs don’t cover people of color
Don Heider · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Networked privacy: How teenagers negotiate context in social media
Alice E Marwick and danah boyd · 2014
Earlier work this paper cites.
A dataset and taxonomy for urban sound research
Justin Salamon, Christopher Jacoby, and Juan Pablo Bello · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
“my data just goes everywhere:” user mental models of the internet and implications for privacy and security
Ruogu Kang, Laura Dabbish, Nathaniel Fruchter, and Sara Kiesler · 2015
Earlier work this paper cites.
Drawing a chip environmental profile: environmental indicators for the semiconductor industry
Aurélie Villard, Alan Lelah, and Daniel Brissaud · 2015
Earlier work this paper cites.
Big other: surveillance capitalism and the prospects of an information civilization
Shoshana Zuboff · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
An uncertain future: Forecasting from static images using variational autoencoders
Jacob Walker, Carl Doersch, Abhinav Gupta, and Martial Hebert · 2016
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter · 2017
Earlier work this paper cites.
Making the V in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Computer vision for assistive technologies
Marco Leo, G Medioni, M Trivedi, Takeo Kanade, and Giovanni Maria Farinella · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
Tegan Maharaj, Nicolas Ballas, Anna Rohrbach, Aaron C Courville, and Christopher Joseph Pal · 2017
Earlier work this paper cites.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Earlier work this paper cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Chris Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Earlier work this paper cites.
Platform capitalism
Nick Srnicek · 2017
Earlier work this paper cites.
Gender and dialect bias in youtube’s automatic captions
Rachael Tatman · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Neural domain adaptation for biomedical question answering
Georg Wiese, Dirk Weissenborn, and Mariana Neves · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Anxiety, panic and self-optimization: Inequalities and the youtube algorithm
Sophie Bishop · 2018
Earlier work this paper cites.
Vggface2: A dataset for recognising faces across pose and age
Qiong Cao, Li Shen, Weidi Xie, Omkar M Parkhi, and Andrew Zisserman · 2018
Earlier work this paper cites.
A short note about kinetics-600
Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman · 2018
Earlier work this paper cites.
The limits of transparency: Data brokers and commodification
Matthew Crain · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
From lifestyle vlogs to everyday interactions
David F Fouhey, Wei-cheng Kuo, Alexei A Efros, and Jitendra Malik · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumeé III, and Kate Crawford · 2018
Cited alongside, same era.
Gender recognition or gender reductionism? the social implications of embedded gender recognition systems
Foad Hamidi, Morgan Klaus Scheuerman, and Stacy M Branham · 2018
Cited alongside, same era.
Avlnet: Learning audio-visual language representations from instructional videos
Andrew Rouditchenko, Angie Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogerio Feris, et al · 2020
Later among the works it cites.
Watching YouTube
Michael Strangelove · 2020
Later among the works it cites.
Mmft-bert: Multimodal fusion transformer with bert encodings for visual question answering
Aisha Urooj, Amir Mazaheri, Mubarak Shah, et al · 2020
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Cited alongside, same era.
Neural motifs: Scene graph parsing with global context
Rowan Zellers, Mark Yatskar, Sam Thomson, and Yejin Choi · 2018
Cited alongside, same era.
Fusion of detected objects in text for visual question answering
Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter · 2019
Cited alongside, same era.
Finding microaggressions in the wild: A case for locating elusive phenomena in social media posts
Luke Breitfeller, Emily Ahn, David Jurgens, and Yulia Tsvetkov · 2019
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Good” isn’t good enough
Ben Green · 2019
Cited alongside, same era.
A case study on combining ASR and visual features for generating instructional video captions
Jack Hessel, Bo Pang, Zhenhai Zhu, and Radu Soricut · 2019
Cited alongside, same era.
Later among the works it cites.
Contrastive learning of medical visual representations from paired images and text
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz · 2020
Later among the works it cites.
VATT: transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Linagzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Later among the works it cites.
Rotary embeddings: A relative revolution, 2021
Stella Biderman, Sid Black, Charles Foster, Leo Gao, Eric Hallahan, Horace He, Ben Wang, and Phil Wang · 2021
Later among the works it cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe · 2021
Later among the works it cites.
Data efficient masked language modeling for vision and language
Yonatan Bitton, Gabriel Stanovsky, Michael Elhadad, and Roy Schwartz · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal · 2021
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, , Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2021
Later among the works it cites.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Later among the works it cites.
Harms of gender exclusivity and challenges in non-binary representation in language technologies
Sunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian, Jeff M Phillips, and Kai-Wei Chang · 2021
Later among the works it cites.
Documenting the english colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner · 2021
Later among the works it cites.
Anticipative video transformer
Rohit Girdhar and Kristen Grauman · 2021
Later among the works it cites.
Ast: Audio spectrogram transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Later among the works it cites.
Toward user-driven sound recognizer personalization with people who are d/deaf or hard of hearing
Steven M Goodman, Ping Liu, Dhruv Jain, Emma J McDonnell, Jon E Froehlich, and Leah Findlater · 2021
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2021
Later among the works it cites.
Five sources of bias in natural language processing
Dirk Hovy and Shrimai Prabhumoye · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Truth from the machine: artificial intelligence and the materialization of identity
Os Keyes, Zoë Hitzig, and Mwenza Blell · 2021
Later among the works it cites.
Self-supervised pre-training and contrastive representation learning for multiple-choice video qa
Seonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang, and Nojun Kwak · 2021
Later among the works it cites.
Cross-modal learning for audio-visual video parsing
Jatin Lamba, Jayaprakash Akula, Rishabh Dabral, Preethi Jyothi, Ganesh Ramakrishnan, et al · 2021
Later among the works it cites.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Later among the works it cites.
Vx2text: End-to-end learning of video-based text generation from multimodal inputs
Xudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang, Devi Parikh, and Lorenzo Torresani · 2021
Later among the works it cites.
Opt: Omni-perception pre-trainer for cross-modal understanding and generation
Jing Liu, Xinxin Zhu, Fei Liu, Longteng Guo, Zijia Zhao, Mingzhen Sun, Weining Wang, Jinqiao Wang, and Hanqing Lu · 2021
Later among the works it cites.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Later among the works it cites.
Understanding the behaviour of contrastive loss
Feng Wang and Huaping Liu · 2021
Later among the works it cites.
Multimodal self-supervised learning of general audio representations
Luyu Wang, Pauline Luc, Adria Recasens, Jean-Baptiste Alayrac, and Aaron van den Oord · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao · 2021
Later among the works it cites.
Disembodied machine learning: On the illusion of objectivity in nlp
Zeerak Waseem, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein · 2021
Later among the works it cites.
Star: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B. Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, and Florian Metze Luke Zettlemoyer Christoph Feichtenhofer · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
Scaling vision transformers, 2021
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
Multiview transformers for video recognition
Shen Yan, Xuehan Xiong, Anurag Arnab, Zhichao Lu, Mi Zhang, Chen Sun, and Cordelia Schmid · 2022
Closest in time.