Fetching the paper…
Reading the bibliography…
Pre-trained models have achieved state-of-the-art results in various Natural Language Processing (NLP) tasks.
Roberta: A robustly optimized bert pretraining approach, 2019
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Sang and Fien De Meulder · 2003
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2006
Earlier work this paper cites.
Distant supervision for relation extraction without labeled data
Mike Mintz, Steven Bills, Rion Snow, and Dan Jurafsky · 2009
Earlier work this paper cites.
Ontonotes release 4.0
Ralph Weischedel, Sameer Pradhan, Lance Ramshaw, Martha Palmer, Nianwen Xue, Mitchell Marcus, Ann Taylor, Craig Greenberg, Eduard Hovy, Robert Belvin, et al · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S. Gordon · 2011
Earlier work this paper cites.
The Winograd schema challenge
Hector J Levesque, Ernest Davis, and Leora Morgenstern · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning · 2015
Earlier work this paper cites.
Lcsts: A large scale chinese short text summarization dataset
Baotian Hu, Qingcai Chen, and Fangze Zhu · 2015
Earlier work this paper cites.
Named entity recognition for chinese social media with jointly trained embeddings
Nanyun Peng and Mark Dredze · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Dataset and neural recurrent sequence labeling model for open-domain factoid question answering
Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, and Wei Xu · 2016
Earlier work this paper cites.
http://web.archive.org/save/http://commoncrawl.org/2016/10/news-dataset-available
Sebastian nagel. 2016. cc-news · 2016
Earlier work this paper cites.
Jingjing Xu, Ji Wen, Xu Sun, and Qi Su · 2017
Earlier work this paper cites.
Chinese medical question answer matching using end-to-end character-level multi-scale cnns
Sheng Zhang, Xin Zhang, Hui Wang, Jiajun Cheng, Pei Li, and Zhaoyun Ding · 2017
Earlier work this paper cites.
Dureader: a chinese machine reading comprehension dataset from real-world applications
Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, et al · 2017
Earlier work this paper cites.
End-to-end neural ad-hoc ranking with kernel pooling
Chenyan Xiong, Zhuyun Dai, Jamie Callan, Zhiyuan Liu, and Russell Power · 2017
Earlier work this paper cites.
Deep neural solver for math word problems
Yan Wang, Xiaojiang Liu, and Shuming Shi · 2017
Earlier work this paper cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2018
Earlier work this paper cites.
Character-based bilstm-crf incorporating pos and dictionaries for chinese opinion target extraction
Yanzeng Li, Tingwen Liu, Diying Li, Quangang Li, Jinqiao Shi, and Yanqiu Wang · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
Lcqmc: A large-scale chinese question matching corpus
Xin Liu, Qingcai Chen, Chong Deng, Huajun Zeng, Jing Chen, Dongfang Li, and Buzhou Tang · 2018
Earlier work this paper cites.
The bq corpus: A large-scale domain-specific chinese corpus for sentence semantic equivalence identification
Jing Chen, Qingcai Chen, Xin Liu, Haijun Yang, Daohe Lu, and Buzhou Tang · 2018
Earlier work this paper cites.
Matching article pairs with graphical decomposition and convolutions
Bang Liu, Di Niu, Haojie Wei, Jinghong Lin, Yancheng He, Kunfeng Lai, and Yu Xu · 2018
Cited alongside, same era.
Multi-scale attentive interaction networks for chinese medical question answer selection
Sheng Zhang, Xin Zhang, Hui Wang, Lixiang Guo, and Shanshan Liu · 2018
Cited alongside, same era.
A span-extraction dataset for chinese machine reading comprehension
Yiming Cui, Ting Liu, Li Xiao, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu · 2018
Cited alongside, same era.
Drcd: a chinese machine reading comprehension dataset
Chih Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai · 2018
Cited alongside, same era.
Cail2018: A large-scale legal dataset for judgment prediction
Cpm: A large-scale generative chinese pre-trained language model
Zhengyan Zhang, Xu Han, Hao Zhou, Pei Ke, Yuxian Gu, Deming Ye, Yujia Qin, Yusheng Su, Haozhe Ji, Jian Guan, et al · 2020
Later among the works it cites.
Colake: Contextualized language and knowledge embedding
Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang, and Zheng Zhang · 2020
Later among the works it cites.
Pre-training text-to-text transformers for concept-centric common sense
Wangchunshu Zhou, Dong-Ho Lee, Ravi Kiran Selvam, Seyeon Lee, Bill Yuchen Lin, and Xiang Ren · 2020
Later among the works it cites.
K-adapter: Infusing knowledge into pre-trained models with adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Cuihong Cao, Daxin Jiang, Ming Zhou, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xianpei Han, Zhen Hu, Heng Wang, et al · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth · 2018
Cited alongside, same era.
ReCoRD: Bridging the gap between human and machine commonsense reading comprehension
Sheng Zhang, Xiaodong Liu, Jingjing Liu, Jianfeng Gao, Kevin Duh, and Benjamin Van Durme · 2018
Cited alongside, same era.
A simple method for commonsense reasoning
Trieu H Trinh and Quoc V Le · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Cited alongside, same era.
Ernie: Enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu · 2019
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Later among the works it cites.
Ernie 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Segabert: Pre-training of segment-aware BERT for language understanding
He Bai, Peng Shi, Jimmy Lin, Luchen Tan, Kun Xiong, Wen Gao, and Ming Li · 2020
Later among the works it cites.
Ernie-doc: The retrospective long-document modeling transformer
Siyu Ding, Junyuan Shang, Shuohuan Wang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Later among the works it cites.
Ernie-gen: An enhanced multi-flow pre-training and fine-tuning framework for natural language generation, 2020
Dongling Xiao, Han Zhang, Yukun Li, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Clue: A chinese language understanding evaluation benchmark
Liang Xu, Hai Hu, Xuanwei Zhang, Lu Li, Chenjie Cao, Yudong Li, Yechen Xu, Kai Sun, Dian Yu, Cong Yu, et al · 2020
Later among the works it cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
A sentence cloze dataset for chinese machine reading comprehension
Yiming Cui, Ting Liu, Ziqing Yang, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu · 2020
Later among the works it cites.
Hongxuan Tang, Jing Liu, Hongyu Li, Yu Hong, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Investigating prior knowledge for challenging chinese machine reading comprehension
Kai Sun, Dian Yu, Dong Yu, and Claire Cardie · 2020
Later among the works it cites.
Canwen Xu, Jiaxin Pei, Hongtao Wu, Yiyu Liu, and Chenliang Li · 2020
Later among the works it cites.
Findings of the 2020 conference on machine translation (wmt20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, et al · 2020
Later among the works it cites.
Kdconv: a chinese multi-domain dialogue dataset towards multi-turn knowledge-driven conversation
Hao Zhou, Chujie Zheng, Kaili Huang, Minlie Huang, and Xiaoyan Zhu · 2020
Later among the works it cites.
Dongling Xiao, Yu-Kun Li, Han Zhang, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
Skep: Sentiment knowledge enhanced pre-training for sentiment analysis
Hao Tian, Can Gao, Xinyan Xiao, Hao Liu, Bolei He, Hua Wu, Haifeng Wang, and Feng Wu · 2020
Later among the works it cites.
Revisiting pre-trained models for chinese natural language processing
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Shijin Wang, and Guoping Hu · 2020
Later among the works it cites.
Chinese medical question answer matching based on interactive sentence representation learning
Xiongtao Cui and Jungang Han · 2020
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention, 2020
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Robo-writers: the rise and risks of language-generating ai
Matthew Hutson · 2021
Closest in time.
Cpm-2: Large-scale cost-effective pre-trained language models
Zhengyan Zhang, Yuxian Gu, Xu Han, Shengqi Chen, Chaojun Xiao, Zhenbo Sun, Yuan Yao, Fanchao Qi, Jian Guan, Pei Ke, et al · 2021
Closest in time.
M6: A chinese multimodal pretrainer
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, et al · 2021
Closest in time.
Wei Zeng, Xiaozhe Ren, Teng Su, Hui Wang, Yi Liao, Zhiwei Wang, Xin Jiang, ZhenZhang Yang, Kaisheng Wang, Xiaoda Zhang, et al · 2021
Closest in time.
Kepler: A unified model for knowledge embedding and pre-trained language representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang · 2021
Closest in time.
Efficientnetv2: Smaller models and faster training
Mingxing Tan and Quoc V. Le · 2021
Closest in time.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
Blow the dog whistle: A Chinese dataset for cant understanding with common sense and world knowledge
Canwen Xu, Wangchunshu Zhou, Tao Ge, Ke Xu, Julian McAuley, and Furu Wei · 2021
Closest in time.
Zen 2.0: Continue training and adaption for n-gram enhanced text encoders
Yan Song, Tong Zhang, Yonggang Wang, and Kai-Fu Lee · 2021
Closest in time.