Fetching the paper…
Reading the bibliography…
Visual understanding goes well beyond object recognition.
Algorithms for the assignment and transportation problems
James Munkres · 1957
Earlier work this paper cites.
A shortest augmenting path algorithm for dense and sparse linear assignment problems
Roy Jonker and Anton Volgenant · 1987
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Hamed Pirsiavash, Carl Vondrick, and Antonio Torralba · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Commonsense reasoning and commonsense knowledge in artificial intelligence
Ernest Davis and Gary Marcus · 2015
Earlier work this paper cites.
Exploring nearest neighbor approaches for image captioning
Jacob Devlin, Saurabh Gupta, Ross B. Girshick, Margaret Mitchell, and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
The Most Common Unisex Names In America: Is Yours One Of Them?, June 2015
Andrew Flowers · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
Licheng Yu, Eunbyung Park, Alexander C. Berg, and Tamara L. Berg · 2015
Earlier work this paper cites.
Temporal perception and prediction in ego-centric video
Yipin Zhou and Tamara L Berg · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Richard Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Social lstm: Human trajectory prediction in crowded spaces
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese · 2016
Earlier work this paper cites.
Unsupervised learning from narrated instruction videos
Jean-Baptiste Alayrac, Piotr Bojanowski, Nishant Agrawal, Josef Sivic, Ivan Laptev, and Simon Lacoste-Julien · 2016
Earlier work this paper cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Generating visual explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2016
Earlier work this paper cites.
Pororoqa: Cartoon video series dataset for story understanding
K Kim, C Nan, MO Heo, SH Choi, and BT Zhang · 2016
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Earlier work this paper cites.
Leveraging visual question answering for image-caption ranking
Xiao Lin and Devi Parikh · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy · 2016
Earlier work this paper cites.
A corpus and evaluation framework for deeper understanding of commonsense stories
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James Allen · 2016
Earlier work this paper cites.
“what happens if…” learning to predict the effect of forces in images
Roozbeh Mottaghi, Mohammad Rastegari, Abhinav Gupta, and Ali Farhadi · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Cited alongside, same era.
Gender-distinguishing features in film dialogue
Alexandra Schofield and Leo Mehr · 2016
Cited alongside, same era.
Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks
Krishna Kumar Singh, Kayvon Fatahalian, and Alexei A Efros · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Cited alongside, same era.
Anticipating visual representations from unlabeled video
Nematus: a toolkit for neural machine translation
Rico Sennrich, Orhan Firat, Kyunghyun Cho, Alexandra Birch, Barry Haddow, Julian Hitschler, Marcin Junczys-Dowmunt, Samuel Läubli, Antonio Valerio Miceli Barone, Jozef Mokry, and Maria Nadejde · 2017
Later among the works it cites.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony R. Dick · 2017
Later among the works it cites.
Learning paraphrastic sentence embeddings from back-translated bitext
John Wieting, Jonathan Mallinson, and Kevin Gimpel · 2017
Later among the works it cites.
A joint speakerlistener-reinforcer model for referring expressions
Licheng Yu, Hao Tan, Mohit Bansal, and Tamara L Berg · 2017
Later among the works it cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba · 2016
Cited alongside, same era.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Qi Wu, Peng Wang, Chunhua Shen, Anthony R. Dick, and Anton van den Hengel · 2016
Cited alongside, same era.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Cited alongside, same era.
Visual7W: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
MUTAN: Multimodal Tucker Fusion for Visual Question Answering
Hedi Ben-younes, Remi Cadene, Matthieu Cord, and Nicolas Thome · 2017
Cited alongside, same era.
Explanation and justification in machine learning: A survey
Or Biran and Courtenay Cotton · 2017
Cited alongside, same era.
Enhanced lstm for natural language inference
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Si Wei, Hui Jiang, and Diana Inkpen · 2017
Cited alongside, same era.
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Closest in time.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Closest in time.
Do explanations make vqa models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh · 2018
Closest in time.
Learning to act properly: Predicting and explaining affordances from images
Ching-Yao Chuang, Jiaman Li, Antonio Torralba, and Sanja Fidler · 2018
Closest in time.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Closest in time.
Who let the dogs out? modeling dog behavior from visual data
Kiana Ehsani, Hessam Bagherinezhad, Joseph Redmon, Roozbeh Mottaghi, and Ali Farhadi · 2018
Closest in time.
From lifestyle vlogs to everyday interactions
David F. Fouhey, Weicheng Kuo, Alexei A. Efros, and Jitendra Malik · 2018
Closest in time.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumeé III, and Kate Crawford · 2018
Closest in time.
Detectron
Ross Girshick, Ilija Radosavovic, Georgia Gkioxari, Piotr Dollár, and Kaiming He · 2018
Closest in time.
Social gan: Socially acceptable trajectories with generative adversarial networks
Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi · 2018
Closest in time.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel R. Bowman, and Noah A. Smith · 2018
Closest in time.
Grounding visual explanations
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Closest in time.
Explainable neural computation via stack neural module networks
Ronghang Hu, Jacob Andreas, Trevor Darrell, and Kate Saenko · 2018
Closest in time.
Multimodal explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2018
Closest in time.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata · 2018
Closest in time.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg · 2018
Closest in time.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Closest in time.
Hypothesis Only Baselines in Natural Language Inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme · 2018
Closest in time.
Overcoming language priors in visual question answering with adversarial regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee · 2018
Closest in time.
Moviegraphs: Towards understanding human-centric situations from videos
Paul Vicol, Makarand Tapaswi, Lluis Castrejon, and Sanja Fidler · 2018
Closest in time.
Answering visual what-if questions: From actions to predicted scene descriptions
Misha Wagner, Hector Basevi, Rakshith Shetty, Wenbin Li, Mateusz Malinowski, Mario Fritz, and Ales Leonardis · 2018
Closest in time.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Closest in time.
Interpretable intuitive physics model
Tian Ye, Xiaolong Wang, James Davidson, and Abhinav Gupta · 2018
Closest in time.
Stair actions: A video dataset of everyday home actions
Yuya Yoshikawa, Jiaqing Lin, and Akikazu Takeuchi · 2018
Closest in time.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Closest in time.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2018
Closest in time.