Fetching the paper…
Reading the bibliography…
As humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 1906
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Logic and conversation
Herbert P Grice · 1975
Earlier work this paper cites.
Scripts, plans, and knowledge
Roger C. Schank and Robert P. Abelson · 1975
Earlier work this paper cites.
Situated knowledges: The science question in feminism and the privilege of partial perspective
Donna Haraway · 1988
Earlier work this paper cites.
Crime in black and white: The violent, scary world of local news
Franklin D Gilliam Jr, Shanto Iyengar, Adam Simon, and Oliver Wright · 1996
Earlier work this paper cites.
Overrepresentation and underrepresentation of african americans and latinos as lawbreakers on television news
Travis L Dixon and Daniel Linz · 2000
Earlier work this paper cites.
Mallet: A machine learning for language toolkit
Andrew Kachites McCallum · 2002
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
Dynamic time warping
Meinard Müller · 2007
Earlier work this paper cites.
Crime news and racialized beliefs: Understanding the relationship between local news viewing and perceptions of african americans and crime
Travis L Dixon · 2008
Earlier work this paper cites.
Exploring the gender divide on youtube: An analysis of the creation and reception of vlogs
Heather Molyneaux, Susan O’Donnell, Kerri Gibson, Janice Singer, et al · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
HMDB: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Verbs: Aspect and causal structure
William Croft · 2012
Earlier work this paper cites.
Activity forecasting
Kris M Kitani, Brian D Ziebart, James Andrew Bagnell, and Martial Hebert · 2012
Earlier work this paper cites.
Learning object class detectors from weakly annotated video
Alessandro Prest, Christian Leistner, Javier Civera, Cordelia Schmid, and Vittorio Ferrari · 2012
Earlier work this paper cites.
The dangers of surveillance
Neil M Richards · 2012
Earlier work this paper cites.
A dataset of 101 human action classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme · 2013
Earlier work this paper cites.
White news: Why local news programs don’t cover people of color
Don Heider · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Networked privacy: How teenagers negotiate context in social media
Alice E Marwick and danah boyd · 2014
Earlier work this paper cites.
Instructional videos for unsupervised harvesting and learning of action examples
Shoou-I Yu, Lu Jiang, and Alexander Hauptmann · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles · 2015
Earlier work this paper cites.
“my data just goes everywhere:” user mental models of the internet and implications for privacy and security
Ruogu Kang, Laura Dabbish, Nathaniel Fruchter, and Sara Kiesler · 2015
Earlier work this paper cites.
What’s cookin’? interpreting cooking videos using text, speech and vision
Jonathan Malmaud, Jonathan Huang, Vivek Rathod, Nick Johnston, Andrew Rabinovich, and Kevin Murphy · 2015
Earlier work this paper cites.
Unsupervised semantic parsing of video collections
Ozan Sener, Amir R Zamir, Silvio Savarese, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Big other: surveillance capitalism and the prospects of an information civilization
Shoshana Zuboff · 2015
Earlier work this paper cites.
Youtube-8m: A large-scale video classification benchmark
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan · 2016
Earlier work this paper cites.
Sort Story: Sorting Jumbled Images and Captions into Stories
Harsh Agrawal, Arjun Chandrasekaran, Dhruv Batra, Devi Parikh, and Mohit Bansal · 2016
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
Interaction networks for learning about objects, relations and physics
Peter Battaglia, Razvan Pascanu, Matthew Lai, Danilo Jimenez Rezende, and Koray kavukcuoglu · 2016
Earlier work this paper cites.
Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C Lawrence Zitnick, et al · 2016
Earlier work this paper cites.
Unsupervised learning for physical interaction through video prediction
Chelsea Finn, Ian Goodfellow, and Sergey Levine · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual storytelling
Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Ishan Misra, C Lawrence Zitnick, and Martial Hebert · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Learning language-visual embedding for movie understanding with natural-language
Atousa Torabi, Niket Tandon, and Leon Sigal · 2016
Cited alongside, same era.
An uncertain future: Forecasting from static images using variational autoencoders
Jacob Walker, Carl Doersch, Abhinav Gupta, and Martial Hebert · 2016
Cited alongside, same era.
Joint discovery of object states and manipulation actions
Jean-Baptiste Alayrac, Ivan Laptev, Josef Sivic, and Simon Lacoste-Julien · 2017
Cited alongside, same era.
Story Comprehension for Predicting What Happens Next
Snigdha Chaturvedi, Haoruo Peng, and Dan Roth · 2017
Cited alongside, same era.
Unsupervised visual-linguistic reference resolution in instructional videos
Multimodal abstractive summarization for how2 videos
Shruti Palaskar, Jindrich Libovickỳ, Spandana Gella, and Florian Metze · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A Smith · 2019
Later among the works it cites.
Dense procedure captioning in narrated instructional videos
Botian Shi, Lei Ji, Yaobo Liang, Nan Duan, Peng Chen, Zhendong Niu, and Ming Zhou · 2019
Later among the works it cites.
Next sentence prediction helps implicit discourse relation classification within and across domains
Wei Shi and Vera Demberg · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in nlp
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
De-An Huang, Joseph J. Lim, Li Fei-Fei, and Juan Carlos Niebles · 2017
Cited alongside, same era.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
Tegan Maharaj, Nicolas Ballas, Anna Rohrbach, Aaron C Courville, and Christopher Joseph Pal · 2017
Cited alongside, same era.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Chris Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Cited alongside, same era.
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Later among the works it cites.
ActivityNet-QA: a dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Later among the works it cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman · 2020
Later among the works it cites.
Noise estimation using density estimation for self-supervised multimodal learning
Elad Amrani, Rami Ben-Ari, Daniel Rotman, and Alex Bronstein · 2020
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2020
Later among the works it cites.
DramaQA: character-centered video story understanding with hierarchical qa
Seongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo, Youwon Jang, Seungchan Lee, Minsu Lee, and Byoung-Tak Zhang · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Large-scale adversarial training for vision-and-language representation learning
Zhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu, Yu Cheng, and Jingjing Liu · 2020
Later among the works it cites.
Beyond instructional videos: Probing for more diverse visual-textual grounding on youtube
Jack Hessel, Zhenhai Zhu, Bo Pang, and Radu Soricut · 2020
Later among the works it cites.
In defense of grid features for visual question answering
Huaizu Jiang, Ishan Misra, Marcus Rohrbach, Erik Learned-Miller, and Xinlei Chen · 2020
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Later among the works it cites.
Self-supervised pre-training and contrastive representation learning for multiple-choice video qa
Seonhoon Kim, Seohyeong Jeong, Eun-Byul Kim, Inho Kang, and Nojun Kwak · 2020
Later among the works it cites.
What is more likely to happen next? video-and-language future event prediction
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Later among the works it cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, Daxin Jiang, and Ming Zhou · 2020
Later among the works it cites.
End-to-end learning of visual representations from uncurated instructional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman · 2020
Later among the works it cites.
Auditing radicalization pathways on youtube
Manoel Horta Ribeiro, Raphael Ottoni, Robert West, Virgílio AF Almeida, and Wagner Meira Jr · 2020
Later among the works it cites.
Watching YouTube
Michael Strangelove · 2020
Later among the works it cites.
Just ask: Learning to answer questions from millions of narrated videos
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Contrastive learning of medical visual representations from paired images and text
Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz · 2020
Later among the works it cites.
ActBERT: Learning global-local video-text representations
Linchao Zhu and Yi Yang · 2020
Later among the works it cites.
VATT: transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Linagzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Closest in time.
Rspnet: Relative speed perception for unsupervised video representation learning
Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, and Chuang Gan · 2021
Closest in time.
Learning temporal dynamics from cycles in narrated video
Dave Epstein, Jiajun Wu, Cordelia Schmid, and Chen Sun · 2021
Closest in time.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Closest in time.
Less is more: Clipbert for video-and-language learning via sparse sampling
Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L Berg, Mohit Bansal, and Jingjing Liu · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Closest in time.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Closest in time.
Decembert: Learning from noisy instructional videos via dense captions and entropy minimization
Zineng Tang, Jie Lei, and Mohit Bansal · 2021
Closest in time.
Disembodied machine learning: On the illusion of objectivity in nlp
Zeerak Waseem, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein · 2021
Closest in time.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Closest in time.