Roberta: A robustly optimized bert pretraining approach, 2019
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Dataset and neural recurrent sequence labeling model for open-domain factoid question answering
Original
Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, and Wei Xu · 2016
Earlier work this paper cites.
Consensus attention-based neural networks for chinese reading comprehension
Original
Yiming Cui, Ting Liu, Zhipeng Chen, Shijin Wang, and Guoping Hu · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A discourse-level named entity recognition and relation extraction dataset for chinese literature text
Original
Jingjing Xu, Ji Wen, Xu Sun, and Qi Su · 2017
Earlier work this paper cites.
Chinese medical question answer matching using end-to-end character-level multi-scale cnns
Sheng Zhang, Xin Zhang, Hui Wang, Jiajun Cheng, Pei Li, and Zhaoyun Ding · 2017
Earlier work this paper cites.
Dataset for the first evaluation on chinese machine reading comprehension
Original
Yiming Cui, Ting Liu, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu · 2017
Earlier work this paper cites.
Dureader: a chinese machine reading comprehension dataset from real-world applications
Original
Wei He, Kai Liu, Jing Liu, Yajuan Lyu, Shiqi Zhao, Xinyan Xiao, Yuan Liu, Yizhong Wang, Hua Wu, Qiaoqiao She, et al · 2017
Earlier work this paper cites.
Deep contextualized word representations
Original
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Original
Alexander Sergeev and Mike Del Balso · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Original
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al · 2018
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel dnn training
Original
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons · 2018
Earlier work this paper cites.
Character-based bilstm-crf incorporating pos and dictionaries for chinese opinion target extraction
Yanzeng Li, Tingwen Liu, Diying Li, Quangang Li, Jinqiao Shi, and Yanqiu Wang · 2018
Earlier work this paper cites.
Lcqmc: A large-scale chinese question matching corpus
Xin Liu, Qingcai Chen, Chong Deng, Huajun Zeng, Jing Chen, Dongfang Li, and Buzhou Tang · 2018
Earlier work this paper cites.
The bq corpus: A large-scale domain-specific chinese corpus for sentence semantic equivalence identification
Jing Chen, Qingcai Chen, Xin Liu, Haijun Yang, Daohe Lu, and Buzhou Tang · 2018
Earlier work this paper cites.
Matching article pairs with graphical decomposition and convolutions
Original
Bang Liu, Di Niu, Haojie Wei, Jinghong Lin, Yancheng He, Kunfeng Lai, and Yu Xu · 2018
Earlier work this paper cites.
Multi-scale attentive interaction networks for chinese medical question answer selection
Sheng Zhang, Xin Zhang, Hui Wang, Lixiang Guo, and Shanshan Liu · 2018
Earlier work this paper cites.
Drcd: a chinese machine reading comprehension dataset
Original
Chih Chieh Shao, Trois Liu, Yuting Lai, Yiying Tseng, and Sam Tsai · 2018
Earlier work this paper cites.
A span-extraction dataset for chinese machine reading comprehension
Original
Yiming Cui, Ting Liu, Li Xiao, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu · 2018
Earlier work this paper cites.
Cail2018: A large-scale legal dataset for judgment prediction
Original
Chaojun Xiao, Haoxi Zhong, Zhipeng Guo, Cunchao Tu, Zhiyuan Liu, Maosong Sun, Yansong Feng, Xianpei Han, Zhen Hu, Heng Wang, et al · 2018
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Original
Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
A span-extraction dataset for chinese machine reading comprehension
Original
Yiming Cui, Ting Liu, Wanxiang Che, Li Xiao, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu · 2018
Earlier work this paper cites.
Paddlepaddle: An open-source deep learning platform from industrial practice
Yanjun Ma, Dianhai Yu, Tian Wu, and Haifeng Wang · 2019
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Original
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu · 2019
Earlier work this paper cites.
Improved knowledge distillation via teacher assistant: Bridging the gap between student and teacher
Original
Seyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, and Hassan Ghasemzadeh · 2019
Earlier work this paper cites.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Earlier work this paper cites.
Knowledge distillation via route constrained optimization
Xiao Jin, Baoyun Peng, Yichao Wu, Yu Liu, Jiaheng Liu, Ding Liang, Junjie Yan, and Xiaolin Hu · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Original
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Earlier work this paper cites.