Fetching the paper…
Reading the bibliography…
The ability to sequence unordered events is an essential skill to comprehend and reason about real world task procedures, which often requires thorough understanding of temporal common sense and multimodal information, as these procedures are often communicated through a combination of texts and images.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 1908
Earlier work this paper cites.
Children’s representation and imitation of events: How goal organization influences 3-year-old children’s memory for action sequences
Jeff Loucks, Christina Mutschler, and Andrew N Meltzoff. 2017 · 1933
Earlier work this paper cites.
The tomkins-horn picture arrangement test
Silvan S Tomkins. 1952 · 1952
Earlier work this paper cites.
Mechanical, behavioural and intentional understanding of picture stories in autistic children
Simon Baron-Cohen, Alan M Leslie, and Uta Frith. 1986 · 1986
Earlier work this paper cites.
Probabilistic text structuring: Experiments with sentence ordering
Mirella Lapata. 2003 · 2003
Earlier work this paper cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020 · 2004
Earlier work this paper cites.
Multimodal pretraining unmasked: Unifying the vision and language BERTs
Emanuele Bugliarello, Ryan Cotterell, Naoaki Okazaki, and Desmond Elliott. 2020 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Sort story: Sorting jumbled images and captions into stories
Harsh Agrawal, Arjun Chandrasekaran, Dhruv Batra, Devi Parikh, and Mohit Bansal. 2016 · 2016
Earlier work this paper cites.
Xinchi Chen, Xipeng Qiu, and Xuanjing Huang. 2016 · 2016
Earlier work this paper cites.
End-to-end neural sentence ordering using pointer network
Jingjing Gong, Xinchi Chen, Xipeng Qiu, and Xuanjing Huang. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Visual storytelling
Ting-Hao K. Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Jacob Devlin, Aishwarya Agrawal, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al. 2016 · 2016
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Order matters: Sequence to sequence for sets
Oriol Vinyals, Samy Bengio, and Manjunath Kudlur. 2016 · 2016
Earlier work this paper cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Cited alongside, same era.
Unsupervised representation learning by sorting sequences
Hsin-Ying Lee, Jia-Bin Huang, Maneesh Singh, and Ming-Hsuan Yang. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deep attentive sentence ordering network
Baiyun Cui, Yingming Li, Ming Chen, and Zhongfei Zhang. 2018 · 2018
Cited alongside, same era.
Autonomous task sequencing in a robot swarm
Lorenzo Garattoni and Mauro Birattari. 2018 · 2018
Cited alongside, same era.
Slm: Learning a discourse language representation with sentence unshuffling
Haejun Lee, Drew A Hudson, Kangwook Lee, and Christopher D Manning. 2020 · 2020
Later among the works it cites.
Hero: Hierarchical encoder for video+ language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Multi-level multimodal transformer network for multimodal recipe comprehension
Ao Liu, Shuai Yuan, Chenbin Zhang, Congjian Luo, Yaqing Liao, Kun Bai, and Zenglin Xu. 2020 · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
A dataset for tracking entities in open domain procedural text
Niket Tandon, Keisuke Sakaguchi, Bhavana Dalvi, Dheeraj Rajagopal, Peter Clark, Michal Guerquin, Kyle Richardson, and Eduard Hovy. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mahnaz Koupaee and William Yang Wang. 2018 · 2018
Cited alongside, same era.
Sentence ordering and coherence modeling using recurrent neural networks
Lajanugen Logeswaran, Honglak Lee, and Dragomir Radev. 2018 · 2018
Cited alongside, same era.
Recipeqa: A challenge dataset for multimodal comprehension of cooking recipes
Semih Yagcioglu, Aykut Erdem, Erkut Erdem, and Nazli Ikizler-Cinbis. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Cited alongside, same era.
Topic-guided coherence modeling for sentence ordering by preserving global and local information
Byungkook Oh, Seungmin Seo, Cheolheon Shin, Eunju Jo, and Kyong-Ho Lee. 2019 · 2019
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019 · 2019
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei. 2021 · 2021
Closest in time.
Ordering sentences and paragraphs with pre-trained encoder-decoder transformers and pointer ensembles
Rémi Calizzano, Malte Ostendorff, and Georg Rehm. 2021 · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2021 · 2021
Closest in time.
Learning temporal dynamics from cycles in narrated video
Dave Epstein, Jiajun Wu, Cordelia Schmid, and Chen Sun. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Closest in time.
proScript: Partially ordered scripts generation
Keisuke Sakaguchi, Chandra Bhagavatula, Ronan Le Bras, Niket Tandon, Peter Clark, and Yejin Choi. 2021 · 2021
Closest in time.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. 2021 · 2021
Closest in time.
Cookie: Contrastive cross-modal knowledge sharing pre-training for vision-language representation
Keyu Wen, Jin Xia, Yuanyuan Huang, Linyang Li, Jiayan Xu, and Jie Shao. 2021 · 2021
Closest in time.
Visual goal-step inference using wikihow
Yue Yang, Artemis Panagopoulou, Qing Lyu, Li Zhang, Mark Yatskar, and Chris Callison-Burch. 2021 · 2021
Closest in time.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021 · 2021
Closest in time.