Fetching the paper…
Reading the bibliography…
Recently, the emergence of pre-trained models (PTMs) has brought natural language processing (NLP) to a new era.
Distilling task-specific knowledge from BERT into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin · 1903
Earlier work this paper cites.
ERNIE: enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu · 1904
Earlier work this paper cites.
ClinicalBERT: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath · 1904
Earlier work this paper cites.
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 1904
Earlier work this paper cites.
Contrastive bidirectional transformer for temporal representation learning
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 1906
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 1908
Earlier work this paper cites.
Matthew Tang, Priyanka Gandhi, Md Ahsanul Kabir, Christopher Zou, Jordyn Blakey, and Xiao Luo · 1910
Earlier work this paper cites.
KEPLER: A unified model for knowledge embedding and pre-trained language representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhiyuan Liu, Juanzi Li, and Jian Tang · 1911
Earlier work this paper cites.
PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J Liu · 1912
Earlier work this paper cites.
“cloze procedure”: A new tool for measuring readability
Wilson L. Taylor · 1953
Earlier work this paper cites.
Distributed representations
GE Hinton, JL McClelland, and DE Rumelhart · 1986
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer · 2001
Earlier work this paper cites.
MiniLM: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2002
Earlier work this paper cites.
BERT-of-Theseus: Compressing BERT by progressive module replacing
Canwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei, and Ming Zhou · 2002
Earlier work this paper cites.
K-adapter: Infusing knowledge into pre-trained models with adapters
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou · 2002
Earlier work this paper cites.
Improving BERT fine-tuning via self-ensemble and self-distillation
Yige Xu, Xipeng Qiu, Ligao Zhou, and Xuanjing Huang · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
A survey on contextual embeddings
Qi Liu, Matt J Kusner, and Phil Blunsom · 2003
Earlier work this paper cites.
Adv-BERT: BERT is not robust on misspellings! generating nature adversarial samples on BERT
Lichao Sun, Kazuma Hashimoto, Wenpeng Yin, Akari Asai, Jia Li, Philip Yu, and Caiming Xiong · 2003
Earlier work this paper cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2004
Earlier work this paper cites.
BERT-ATTACK: Adversarial attack against BERT using BERT
Linyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue, and Xipeng Qiu · 2004
Earlier work this paper cites.
Adversarial training for large neural language models
Xiulei Liu, Hao Cheng, Peng cheng He, Weizhu Chen, Yu Wang, Hoifung Poon, and Jianfeng Gao · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
Model compression
Cristian Bucilua, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Dumitru Erhan, Yoshua Bengio, Aaron C. Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel P. Kuksa · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu · 2013
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc VV Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov · 2014
Earlier work this paper cites.
Knowledge graph and text jointly embedding
Zhen Wang, Jianwen Zhang, Jianlin Feng, and Zheng Chen · 2014
Earlier work this paper cites.
Improving vector space word representations using multilingual correlation
Manaal Faruqui and Chris Dyer · 2014
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning · 2015
Earlier work this paper cites.
Long short-term memory over recursive structures
Xiaodan Zhu, Parinaz Sobihani, and Hongyu Guo · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le · 2015
Earlier work this paper cites.
How well do distributional models capture different types of semantic knowledge?
Dana Rubinstein, Effi Levi, Roy Schwartz, and Ari Rappoport · 2015
Earlier work this paper cites.
Distributional vectors encode referential attributes
Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Padó · 2015
Earlier work this paper cites.
Aligning knowledge and text embeddings by entity descriptions
Huaping Zhong, Jianwen Zhang, Zhen Wang, Hai Wan, and Zheng Chen · 2015
Earlier work this paper cites.
Bilingual word representations with monolingual quality in mind
Minh-Thang Luong, Hieu Pham, and Christopher D Manning · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Recurrent neural network for text classification with multi-task learning
Pengfei Liu, Xipeng Qiu, and Xuanjing Huang · 2016
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Context2Vec: Learning generic context embedding with bidirectional LSTM
Oren Melamud, Jacob Goldberger, and Ido Dagan · 2016
Earlier work this paper cites.
Representation learning of knowledge graphs with entity descriptions
Ruobing Xie, Zhiyuan Liu, Jia Jia, Huanbo Luan, and Maosong Sun · 2016
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher · 2017
Earlier work this paper cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2017
Earlier work this paper cites.
Unsupervised pretraining for sequence to sequence learning
Prajit Ramachandran, Peter J Liu, and Quoc Le · 2017
Earlier work this paper cites.
Knowledge graph representation with jointly structural and textual encoding
Jiacheng Xu, Xipeng Qiu, Kan Chen, and Xuanjing Huang · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass · 2017
Earlier work this paper cites.
Allennlp: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke S. Zettlemoyer · 2017
Earlier work this paper cites.
Semi-supervised sequence tagging with bidirectional language models
Matthew E. Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2017
Earlier work this paper cites.
A survey of model compression and acceleration for deep neural networks
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2017
Earlier work this paper cites.
Exploiting semantics in neural machine translation with graph convolutional networks
Diego Marcheggiani, Joost Bastings, and Ivan Titov · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Contextual string embeddings for sequence labeling
Alan Akbik, Duncan Blythe, and Roland Vollgraf · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
A multi-task approach to learning multilingual representations
Karan Singla, Doğan Can, and Shrikanth Narayanan · 2018
Cited alongside, same era.
Sentence encoders on STILTs: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Tanda: Transfer and adapt pre-trained transformer models for answer sentence selection
Siddhant Garg, Thuy Vu, and Alessandro Moschitti · 2019
Later among the works it cites.
An embarrassingly simple approach for transfer learning from pretrained language models
Alexandra Chronopoulou, Christos Baziotis, and Alexandros Potamianos · 2019
Later among the works it cites.
Specializing word embeddings (for parsing) by information bottleneck
Xiang Lisa Li and Jason Eisner · 2019
Later among the works it cites.
CTRL: A conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R Varshney, Caiming Xiong, and Richard Socher · 2019
Later among the works it cites.
A multiscale visualization of attention in the transformer model
Jesse Vig · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Cited alongside, same era.
Efficient contextualized representation: Language model pruning for sequence labeling
Liyuan Liu, Xiang Ren, Jingbo Shang, Xiaotao Gu, Jian Peng, and Jiawei Han · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Star-transformer
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang · 2019
Cited alongside, same era.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, and Michael Auli · 2019
Cited alongside, same era.
MASS: masked sequence to sequence pre-training for language generation
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2019
Cited alongside, same era.
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann · 2019
Later among the works it cites.
CoQA: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D. Manning · 2019
Later among the works it cites.
Technical report on conversational question answering
Ying Ju, Fubang Zhao, Shijie Chen, Bowen Zheng, Xuefeng Yang, and Yunfeng Liu · 2019
Later among the works it cites.
An investigation of transfer learning-based sentiment analysis in japanese
Enkhbold Bataa and Joshua Wu · 2019
Later among the works it cites.
BERT post-training for review reading comprehension and aspect-based sentiment analysis
Hu Xu, Bing Liu, Lei Shu, and Philip S. Yu · 2019
Later among the works it cites.
Alexander Rietzler, Sebastian Stabinger, Paul Opitz, and Stefan Engl · 2019
Later among the works it cites.
Biomedical named entity recognition with multilingual BERT
Kai Hakala and Sampo Pyysalo · 2019
Later among the works it cites.
Pre-trained language model representations for language generation
Sergey Edunov, Alexei Baevski, and Michael Auli · 2019
Later among the works it cites.
On the use of BERT for neural machine translation
Stephane Clinchant, Kweon Woo Jung, and Vassilina Nikoulina · 2019
Later among the works it cites.
Recycling a pre-trained BERT encoder for neural machine translation
Kenji Imamura and Eiichiro Sumita · 2019
Later among the works it cites.
Text summarization with pretrained encoders
Yang Liu and Mirella Lapata · 2019
Later among the works it cites.
Is BERT really robust? natural language attack on text classification and entailment
Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits · 2019
Later among the works it cites.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Later among the works it cites.
Megatron-LM: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A Smith · 2019
Later among the works it cites.
UniLMv2: Pseudo-masked language models for unified language model pre-training
Hangbo Bao, Li Dong, Furu Wei, Wenhui Wang, Nan Yang, Xiaodong Liu, Yu Wang, Songhao Piao, Jianfeng Gao, Ming Zhou, et al · 2020
Closest in time.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning · 2020
Closest in time.
Pretrained encyclopedia: Weakly supervised knowledge-pretrained language model
Wenhan Xiong, Jingfei Du, William Yang Wang, and Veselin Stoyanov · 2020
Closest in time.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Closest in time.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith · 2020
Closest in time.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh · 2020
Closest in time.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2020
Closest in time.
Colake: Contextualized language and knowledge embedding
Tianxiang Sun, Yunfan Shao, Xipeng Qiu, Qipeng Guo, Yaru Hu, Xuanjing Huang, and Zheng Zhang · 2020
Closest in time.
RobBERT: a Dutch RoBERTa-based language model
Pieter Delobelle, Thomas Winters, and Bettina Berendt · 2020
Closest in time.
VL-BERT: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Closest in time.
Compressing BERT: Studying the effects of weight pruning on transfer learning
Mitchell A Gordon, Kevin Duh, and Nicholas Andrews · 2020
Closest in time.
Q-BERT: Hessian based ultra low precision quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Closest in time.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin · 2020
Closest in time.
The right tool for the job: Matching model and instance complexities
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith · 2020
Closest in time.
Bert loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei · 2020
Closest in time.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger · 2020
Closest in time.
Are pre-trained language models aware of phrases? simple but strong baselines for grammar induction
Taeuk Kim, Jihun Choi, Daniel Edmiston, and Sang goo Lee · 2020
Closest in time.
A knowledge-enhanced pretraining model for commonsense story generation
Jian Guan, Fei Huang, Zhihao Zhao, Xiaoyan Zhu, and Minlie Huang · 2020
Closest in time.
Cross-lingual ability of multilingual BERT: An empirical study
Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth · 2020
Closest in time.
AraBERT: Transformer-based model for Arabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj · 2020
Closest in time.
UniViLM: A unified video and language pre-training model for multimodal understanding and generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Xilin Chen, and Ming Zhou · 2020
Closest in time.
Compressing large-scale transformer-based models: A case study on BERT
Prakhar Ganesh, Yao Chen, Xin Lou, Mohammad Ali Khan, Yin Yang, Deming Chen, Marianne Winslett, Hassan Sajjad, and Preslav Nakov · 2020
Closest in time.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Closest in time.
TwinBERT: Distilling knowledge to twin-structured BERT models for efficient retrieval
Wenhao Lu, Jian Jiao, and Ruofei Zhang · 2020
Closest in time.
Depth-adaptive transformer
Maha Elbayad, Jiatao Gu, Edouard Grave, and Michael Auli · 2020
Closest in time.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith · 2020
Closest in time.
How can we know what language models know
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig · 2020
Closest in time.
Textbrewer: An open-source knowledge distillation toolkit for natural language processing
Ziqing Yang, Yiming Cui, Zhipeng Chen, Wanxiang Che, Ting Liu, Shijin Wang, and Guoping Hu · 2020
Closest in time.
Retrospective reader for machine reading comprehension
Zhuosheng Zhang, Junjie Yang, and Hai Zhao · 2020
Closest in time.
Select, answer and explain: Interpretable multi-hop reading comprehension over multiple documents
Ming Tu, Kevin Huang, Guangtao Wang, Jing Huang, Xiaodong He, and Bowen Zhou · 2020
Closest in time.
Adversarial training for aspect-based sentiment analysis with BERT
Akbar Karimi, Leonardo Rossi, Andrea Prati, and Katharina Full · 2020
Closest in time.
Youwei Song, Jiahai Wang, Zhiwei Liang, Zhiyue Liu, and Tao Jiang · 2020
Closest in time.
Extractive summarization as text matching
Ming Zhong, Pengfei Liu, Yiran Chen, Danqing Wang, Xipeng Qiu, and Xuan-Jing Huang · 2020
Closest in time.
Data augmentation using pre-trained transformer models
Varun Kumar, Ashutosh Choudhary, and Eunah Cho · 2020
Closest in time.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al · 2020
Closest in time.
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2021
Closest in time.
It’s not just size that matters: Small language models are also few-shot learners
Timo Schick and Hinrich Schütze · 2021
Closest in time.
WARP: word-level adversarial reprogramming
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May · 2021
Closest in time.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Closest in time.
A global past-future early exit method for accelerating inference of pre-trained language models
Kaiyuan Liao, Yi Zhang, Xuancheng Ren, Qi Su, Xu Sun, and Bin He · 2021
Closest in time.
Early exiting with ensemble internal classifiers
Tianxiang Sun, Yunhua Zhou, Xiangyang Liu, Xinyu Zhang, Hao Jiang, Zhao Cao, Xuanjing Huang, and Xipeng Qiu · 2021
Closest in time.
Accelerating bert inference for sequence labeling via early-exit
Xiaonan Li, Yunfan Shao, Tianxiang Sun, Hang Yan, Xipeng Qiu, and Xuanjing Huang · 2021
Closest in time.
Faster depth-adaptive transformers
Yijin Liu, Fandong Meng, Jie Zhou, Yufeng Chen, and Jinan Xu · 2021
Closest in time.
Elbert: Fast albert with confidence-window based early exit
Keli Xie, Siyuan Lu, Meiqi Wang, and Zhongfeng Wang · 2021
Closest in time.
How many data points is a prompt worth?
Teven Le Scao and Alexander M. Rush · 2021
Closest in time.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze · 2021
Closest in time.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner · 2021
Closest in time.
Factual probing is [MASK]: learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen · 2021
Closest in time.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Closest in time.