Fetching the paper…
Reading the bibliography…
In contrast with previous approaches where information flows only towards deeper layers of a stack, we consider a multi-pass transformer (MPT) architecture in which earlier layers are allowed to process information in light of the output of later layers.
David R So, Chen Liang, and Quoc V Le. 2019 · 1901
Earlier work this paper cites.
Dynamic layer aggregation for neural machine translation with routing-by-agreement
Zi-Yi Dou, Zhaopeng Tu, Xing Wang, Longyue Wang, Shuming Shi, and Tong Zhang. 2019 · 1902
Earlier work this paper cites.
Exploring randomly wired neural networks for image recognition
Saining Xie, Alexander Kirillov, Ross Girshick, and Kaiming He. 2019 · 1904
Earlier work this paper cites.
Exploiting sentential context for neural machine translation
Xing Wang, Zhaopeng Tu, Longyue Wang, and Shuming Shi. 2019 · 1906
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Multiscale collaborative deep models for neural machine translation
Xiangpeng Wei, Heng Yu, Yue Hu, Yue Zhang, Rongxiang Weng, and Weihua Luo. 2020 · 2004
Earlier work this paper cites.
Character matters: Video story understanding with character-aware relations
Shijie Geng, Ji Zhang, Zuohui Fu, Peng Gao, Hang Zhang, and Gerard de Melo. 2020b · 2005
Earlier work this paper cites.
Spatio-temporal scene graphs for video dialog
Shijie Geng, Peng Gao, Chiori Hori, Jonathan Le Roux, and Anoop Cherian. 2020a · 2007
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le. 2016 · 2016
Cited alongside, same era.
Dual path networks
Yunpeng Chen, Jianan Li, Huaxin Xiao, Xiaojie Jin, Shuicheng Yan, and Jiashi Feng. 2017 · 2017
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017 · 2017
Later among the works it cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Relation networks for object detection
Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei. 2018 · 2018
Later among the works it cites.
Non-local neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xception: Deep learning with depthwise separable convolutions
François Chollet. 2017 · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. 2017 · 2017
Cited alongside, same era.
Dynamic fusion with intra- and inter-modality attention flow for visual question answering
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven C. H. Hoi, Xiaogang Wang, and Hongsheng Li. 2019a
Cited in the paper.
Multi-modality latent interaction network for visual question answering
Peng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang, and Hongsheng Li. 2019b
Cited in the paper.
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018 · 2018
Later among the works it cites.