Fetching the paper…
Reading the bibliography…
Large-scale pre-trained models (PTMs) such as BERT and GPT have recently achieved great success and become a milestone in the field of artificial intelligence (AI).
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
Pre-training with whole word masking for chinese bert
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. 2019 · 1906
Earlier work this paper cites.
Inducing syntactic trees from bert representations
Rudolf Rosa and David Mareček. 2019 · 1906
Earlier work this paper cites.
Ernie 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang. 2019d · 1907
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Nezha: Neural contextualized representation for chinese language understanding
Junqiu Wei, Xiaozhe Ren, Xiaoguang Li, Wenyong Huang, Yi Liao, Yasheng Wang, Jiashu Lin, Xin Jiang, Xiao Chen, and Qun Liu. 2019 · 1909
Earlier work this paper cites.
Yung-Sung Chuang, Chi-Liang Liu, Hung-Yi Lee, and Lin-shan Lee. 2019 · 1910
Earlier work this paper cites.
Understanding multi-head attention in abstractive summarization
Joris Baan, Maartje ter Hoeve, Marlies van der Wees, Anne Schuth, and Maarten de Rijke. 2019 · 1911
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019 · 1911
Earlier work this paper cites.
Cloze procedure: A new tool for measuring readability
Wilson L Taylor. 1953 · 1953
Earlier work this paper cites.
Some tests of the decay theory of immediate memory
John Brown. 1958 · 1958
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Working memory
Alan Baddeley. 1992 · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, and Paolo Frasconi. 1994 · 1994
Earlier work this paper cites.
Below the surface: Analogical similarity and retrieval competition in reminding
Charles M Wharton, Keith J Holyoak, Paul E Downing, Trent E Lange, Thomas D Wickens, and Eric R Melz. 1994 · 1994
Earlier work this paper cites.
Learning to learn: Introduction and overview
Sebastian Thrun and Lorien Pratt. 1998 · 1998
Earlier work this paper cites.
Improving predictive inference under covariate shift by weighting the log-likelihood function
Hidetoshi Shimodaira. 2000 · 2000
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Wei-Tsung Kao, Tsung-Han Wu, Po-Han Chi, Chun-Cheng Hsieh, and Hung-Yi Lee. 2020 · 2001
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti. 2020 · 2001
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2002
Earlier work this paper cites.
K-adapter: Infusing knowledge into pre-trained models with adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Cuihong Cao, Daxin Jiang, Ming Zhou, et al. 2020b · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin. 2003 · 2003
Earlier work this paper cites.
Xgpt: Cross-modal generative pre-training for image captioning
Qiaolin Xia, Haoyang Huang, Nan Duan, Dongdong Zhang, Lei Ji, Zhifang Sui, Edward Cui, Taroon Bharti, Xin Liu, and Ming Zhou. 2020 · 2003
Earlier work this paper cites.
Time constraints and resource sharing in adults’ working memory spans
Pierre Barrouillet, Sophie Bernardin, and Valérie Camos. 2004 · 2004
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Regularized multi–task learning
Theodoros Evgeniou and Massimiliano Pontil. 2004 · 2004
Earlier work this paper cites.
Learning to learn with the informative vector machine
Neil D Lawrence and John C Platt. 2004 · 2004
Earlier work this paper cites.
Learning and evaluating classifiers under sample selection bias
Bianca Zadrozny. 2004 · 2004
Earlier work this paper cites.
A high-performance semi-supervised learning method for text chunking
Rie Johnson and Tong Zhang. 2005 · 2005
Earlier work this paper cites.
Domain adaptation for statistical classifiers
Hal Daume III and Daniel Marcu. 2006 · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh. 2006 · 2006
Earlier work this paper cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Haoyang Huang, Lin Su, Di Qi, Nan Duan, Edward Cui, Taroon Bharti, Lei Zhang, Lijuan Wang, Jianfeng Gao, Bei Liu, et al. 2020b · 2006
Earlier work this paper cites.
Self-supervised learning: Generative or contrastive
Xiao Liu, Fanjin Zhang, Zhenyu Hou, Zhaoyu Wang, Li Mian, Jing Zhang, and Jie Tang. 2020b · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Li, Madian Khabsa, Han Fang, and Hao Ma. 2020c · 2006
Earlier work this paper cites.
Infoxlm: An information-theoretic framework for cross-lingual language model pre-training
Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, Heyan Huang, and Ming Zhou. 2020b · 2007
Earlier work this paper cites.
Co-clustering based classification for out-of-domain documents
Wenyuan Dai, Gui-Rong Xue, Qiang Yang, and Yong Yu. 2007 · 2007
Earlier work this paper cites.
Multi-task feature learning
An Evgeniou and Massimiliano Pontil. 2007 · 2007
Earlier work this paper cites.
Mapping and revising markov logic networks for transfer learning
Lilyana Mihalkova, Tuyen Huynh, and Raymond J Mooney. 2007 · 2007
Earlier work this paper cites.
Self-taught learning: transfer learning from unlabeled data
Rajat Raina, Alexis Battle, Honglak Lee, Benjamin Packer, and Andrew Y Ng. 2007 · 2007
Earlier work this paper cites.
Knowledge-aware language model pretraining
Corby Rosset, Chenyan Xiong, Minh Phan, Xia Song, Paul Bennett, and Saurabh Tiwary. 2020 · 2007
Earlier work this paper cites.
Facts as experts: Adaptable and interpretable neural memory over symbolic knowledge
Pat Verga, Haitian Sun, Livio Baldini Soares, and William W Cohen. 2020 · 2007
Earlier work this paper cites.
Multi-task gaussian process prediction
Chris Williams, Edwin V Bonilla, and Kian M Chai. 2007 · 2007
Earlier work this paper cites.
Variance-reduced language pretraining via a mask proposal network
Liang Chen, Tianyuan Zhang, Di He, Guolin Ke, Liwei Wang, and Tie-Yan Liu. 2020b · 2008
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
Self-taught clustering
Wenyuan Dai, Qiang Yang, Gui-Rong Xue, and Yong Yu. 2008 · 2008
Earlier work this paper cites.
Knowledge transfer via multiple model local structure mapping
Jing Gao, Wei Fan, Jing Jiang, and Jiawei Han. 2008 · 2008
Earlier work this paper cites.
Transferred dimensionality reduction
Zheng Wang, Yangqiu Song, and Changshui Zhang. 2008 · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009 · 2009
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen. 2009 · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2009 · 2009
Earlier work this paper cites.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze. 2020 · 2009
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020 · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent. 2010 · 2010
Earlier work this paper cites.
Word representations: a simple and general method for semi-supervised learning
Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Language models are open knowledge graphs
Chenguang Wang, Xiao Liu, and Dawn Song. 2020a · 2010
Earlier work this paper cites.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He. 2020 · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011 · 2011
Earlier work this paper cites.
Towards a universal continuous knowledge base
Gang Chen, Maosong Sun, and Yang Liu. 2020a · 2012
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012 · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller. 2012 · 2012
Earlier work this paper cites.
Generating adversarial examples in chinese texts using sentence-pieces
Linyang Li, Yunfan Shao, Demin Song, Xipeng Qiu, and Xuanjing Huang. 2020c · 2012
Earlier work this paper cites.
Xuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang. 2020 · 2012
Earlier work this paper cites.
Chenglei Si, Zhengyan Zhang, Fanchao Qi, Zhiyuan Liu, Yasheng Wang, Qun Liu, and Maosong Sun. 2020 · 2012
Earlier work this paper cites.
Cpm: A large-scale generative chinese pre-trained language model
Zhengyan Zhang, Xu Han, Hao Zhou, Pei Ke, Yuxian Gu, Deming Ye, Yujia Qin, Yusheng Su, Haozhe Ji, Jian Guan, et al. 2020c · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli. 2013 · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Overfeat: Integrated recognition, localization and detection using convolutional networks
Pierre Sermanet, David Eigen, Xiang Zhang, Michaël Mathieu, Rob Fergus, and Yann LeCun. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Object detection via a multi-region and semantic segmentation-aware cnn model
Spyros Gidaris and Nikos Komodakis. 2015 · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan. 2015 · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy. 2015 · 2015
Earlier work this paper cites.
Deeply-supervised nets
Chen-Yu Lee, Saining Xie, Patrick Gallagher, Zhengyou Zhang, and Zhuowen Tu. 2015 · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Conditional random fields as recurrent neural networks
Shuai Zheng, Sadeep Jayasumana, Bernardino Romera-Paredes, Vibhav Vineet, Zhizhong Su, Dalong Du, Chang Huang, and Philip HS Torr. 2015 · 2015
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model
Wenhan Xiong, Jingfei Du, William Yang Wang, and Veselin Stoyanov. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Q8bert: Quantized 8bit bert
Ofir Zafrir, Guy Boudoukh, Peter Izsak, and Moshe Wasserblat. 2019 · 2019
Later among the works it cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Philip Pham, Anirudh Ravula, and Sumit Sanghai. 2020 · 2020
Later among the works it cites.
PLATO: Pre-trained dialogue generation model with discrete latent variable
Siqi Bao, Huang He, Fan Wang, Hua Wu, and Haifeng Wang. 2020 · 2020
Later among the works it cites.
Palm: Pre-training an autoencoding&autoregressive language model for context-conditioned generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016 · 2016
Cited alongside, same era.
Layer normalization
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele. 2016 · 2016
Cited alongside, same era.
Probing for semantic evidence of composition by means of simple classification tasks
Allyson Ettinger, Ahmed Elgohary, and Philip Resnik. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling. 2016 · 2016
Cited alongside, same era.
Bin Bi, Chenliang Li, Chen Wu, Ming Yan, and Wei Wang. 2020 · 2020
Later among the works it cites.
Inducing relational knowledge from BERT
Zied Bouraoui, José Camacho-Collados, and Steven Schockaert. 2020 · 2020
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Later among the works it cites.
Differentiable reasoning over a virtual knowledge base
Bhuwan Dhingra, Manzil Zaheer, Vidhisha Balachandran, Graham Neubig, Ruslan Salakhutdinov, and William W Cohen. 2020 · 2020
Later among the works it cites.
Cogltx: Applying bert to long texts
Ming Ding, Chang Zhou, Hongxia Yang, and Jie Tang. 2020 · 2020
Later among the works it cites.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Later among the works it cites.
Entities as experts: Sparse memory access with entity supervision
Thibault Févry, Livio Baldini Soares, Nicholas FitzGerald, Eunsol Choi, and Tom Kwiatkowski. 2020 · 2020
Later among the works it cites.
Compressing BERT: studying the effects of weight pruning on transfer learning
Mitchell A. Gordon, Kevin Duh, and Nicholas Andrews. 2020 · 2020
Later among the works it cites.
Train no evil: Selective masking for task-guided pre-training
Yuxian Gu, Zhengyan Zhang, Xiaozhi Wang, Zhiyuan Liu, and Maosong Sun. 2020 · 2020
Later among the works it cites.
A knowledge-enhanced pretraining model for commonsense story generation
Jian Guan, Fei Huang, Zhihao Zhao, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020 · 2020
Later among the works it cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020 · 2020
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Later among the works it cites.
Sentilare: Linguistic knowledge enhanced language representation for sentiment analysis
Pei Ke, Haozhe Ji, Siyang Liu, Xiaoyan Zhu, and Minlie Huang. 2020 · 2020
Later among the works it cites.
Are pre-trained language models aware of phrases? simple but strong baselines for grammar induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang-goo Lee. 2020 · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020 · 2020
Later among the works it cites.
A mutual information maximization perspective of language representation learning
Lingpeng Kong, Cyprien de Masson d’Autume, Lei Yu, Wang Ling, Zihang Dai, and Dani Yogatama. 2020 · 2020
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee. 2020 · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Later among the works it cites.
Contextual and non-contextual word embeddings: an in-depth linguistic investigation
Alessio Miaschi and Felice Dell’Orletta. 2020 · 2020
Later among the works it cites.
Adversarial nli: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
Later among the works it cites.
Merlin: A gpu accelerated recommendation framework
Even Oldridge, J. Perez, Ben Frederickson, Nicolas Koumchatzky, M. Lee, Z.-H. Wang, Lei Wu, F. Yu, Rick Zamora, O. Yılmaz, Alec M. Gunny, Vinh Phu Nguyen, and S. Lee. 2020 · 2020
Later among the works it cites.
Rethinking softmax cross-entropy loss for adversarial robustness
Tianyu Pang, Kun Xu, Yinpeng Dong, Chao Du, Ning Chen, and Jun Zhu. 2020 · 2020
Later among the works it cites.
E-BERT: efficient-yet-effective entity embeddings for BERT
Nina Pörner, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
When BERT plays the lottery, all tickets are winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
Pre-trained models for natural language processing: A survey
Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020 · 2020
Later among the works it cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020 · 2020
Later among the works it cites.
And the bit goes down: Revisiting the quantization of neural networks
Pierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham, and Hervé Jégou. 2020 · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Colake: Contextualized language and knowledge embedding
Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang, and Zheng Zhang. 2020 · 2020
Later among the works it cites.
Parsing as pretraining
David Vilares, Michalina Strzyz, Anders Søgaard, and Carlos Gómez-Rodríguez. 2020 · 2020
Later among the works it cites.
Can neural networks acquire a structural bias from raw linguistic data?
Alex Warstadt and Samuel R. Bowman. 2020 · 2020
Later among the works it cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Zhiyong Wu, Yun Chen, Ben Kao, and Qun Liu. 2020 · 2020
Later among the works it cites.
Alternating language modeling for cross-lingual pre-training
Jian Yang, Shuming Ma, D. Zhang, Shuangzhi Wu, Zhou jun Li, and M. Zhou. 2020 · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2020 · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al. 2020 · 2020
Later among the works it cites.
Word-level textual adversarial attacking as combinatorial optimization
Yuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu, Meng Zhang, Qun Liu, and Maosong Sun. 2020 · 2020
Later among the works it cites.
Accelerating training of transformer-based language models with progressive layer dropping
Minjia Zhang and Yuxiong He. 2020 · 2020
Later among the works it cites.
Rethinking pre-training and self-training
Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk, and Quoc Le. 2020 · 2020
Later among the works it cites.
Insightface project
2021 · 2021
Closest in time.
MindSpore Deep Learning Framework
2021 · 2021
Closest in time.
OneFlow Deep Learning Framework
2021 · 2021
Closest in time.
Plato-2: Towards building an open-domain chatbot via curriculum learning
Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhen Guo, Zhibin Liu, and Xinchao Xu. 2021 · 2021
Closest in time.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al. 2021 · 2021
Closest in time.
All nlp tasks are generation tasks: A general pretraining framework
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2021 · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2021 · 2021
Closest in time.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Closest in time.
Is supervised syntactic parsing beneficial for language understanding tasks? an empirical investigation
Goran Glavaš and Ivan Vulić. 2021 · 2021
Closest in time.
Ptr: Prompt tuning with rules for text classification
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Closest in time.
Fastmoe: A fast mixture-of-expert training system
Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, and Jie Tang. 2021 · 2021
Closest in time.
Knowledgeable prompt-tuning: Incorporating knowledge into prompt verbalizer for text classification
Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Juanzi Li, and Maosong Sun. 2021 · 2021
Closest in time.
Wenlan: Bridging vision and language by large-scale multi-modal pre-training
Yuqi Huo, Manli Zhang, Guangzhen Liu, Haoyu Lu, Yizhao Gao, Guoxing Yang, Jingyuan Wen, Heng Zhang, Baogui Xu, Weihao Zheng, et al. 2021 · 2021
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021 · 2021
Closest in time.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Closest in time.
Token-aware virtual adversarial training in natural language understanding
Linyang Li and Xipeng Qiu. 2021 · 2021
Closest in time.
Terapipe: Token-level pipeline parallelism for training large-scale language models
Zhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo, Hao Zhang, Dawn Song, and Ion Stoica. 2021 · 2021
Closest in time.
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2021 · 2021
Closest in time.
Efficient large-scale language model training on gpu clusters
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021 · 2021
Closest in time.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Erica: Improving entity and relation understanding for pre-trained language models via contrastive learning
Yujia Qin, Yankai Lin, Ryuichi Takanobu, Zhiyuan Liu, Peng Li, Heng Ji, Minlie Huang, Maosong Sun, and Jie Zhou. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand hini Agarwal, Girish Sastry, Amand a Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
Closest in time.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Closest in time.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Closest in time.
Zero-offload: Democratizing billion-scale model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 · 2021
Closest in time.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021 · 2021
Closest in time.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Closest in time.
Reasoning over virtual knowledge bases with open predicate relations
Haitian Sun, Pat Verga, Bhuwan Dhingra, Ruslan Salakhutdinov, and William W Cohen. 2021 · 2021
Closest in time.
On learning universal representations across languages
Xiangpeng Wei, Yue Hu, Rongxiang Weng, Luxi Xing, Heng Yu, and Weihua Luo. 2021 · 2021
Closest in time.
How neural networks extrapolate: From feedforward to graph neural networks
Keyulu Xu, Jingling Li, Mozhi Zhang, Simon S Du, Ken-ichi Kawarabayashi, and Stefanie Jegelka. 2021 · 2021
Closest in time.
Adversarial language games for advanced natural language intelligence
Yuan Yao, Haoxi Zhong, Zhengyan Zhang, Xu Han, Xiaozhi Wang, Chaojun Xiao, Guoyang Zeng, Zhiyuan Liu, and Maosong Sun. 2021 · 2021
Closest in time.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al. 2021 · 2021
Closest in time.
Controllable generation from pre-trained language models via inverse prompting
Xu Zou, Da Yin, Qingyang Zhong, Hongxia Yang, Zhilin Yang, and Jie Tang. 2021 · 2021
Closest in time.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu. 2015 · 2025
Closest in time.
What’s in an embedding? analyzing word embeddings through multilingual evaluation
Arne Köhn. 2015 · 2073
Closest in time.