Fetching the paper…
Reading the bibliography…
Masked language models (MLMs) such as BERT and RoBERTa have revolutionized the field of Natural Language Understanding in the past few years.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, and Yann LeCun. 2005 · 2005
Earlier work this paper cites.
The second international chinese word segmentation bakeoff
Thomas Emerson. 2005 · 2005
Earlier work this paper cites.
The third international chinese language processing bakeoff: Word segmentation and named entity recognition
Gina-Anne Levow. 2006 · 2006
Earlier work this paper cites.
Ontonotes release 4.0. ldc2011t03
Ralph Weischedel, Sameer Pradhan, Lance Ramshaw, Martha Palmer, Nianwen Xue, Mitchell Marcus, Ann Taylor, Craig Greenberg, Eduard Hovy, and Robert Belvin. 2011 · 2011
Earlier work this paper cites.
Clear: Contrastive learning for sentence representation
Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Unsupervised learning of visual representations using videos
Xiaolong Wang and Abhinav Gupta. 2015 · 2015
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
F-score driven max margin neural network for named entity recognition in chinese social media
Hangfeng He and Xu Sun. 2017 · 2017
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Chinese NER using lattice LSTM
Yue Zhang and Jie Yang. 2018 · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? comparing the geometry of bert, elmo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Earlier work this paper cites.
Glyce: Glyph-vectors for chinese character representations
Yuxian Meng, Wei Wu, Fei Wang, Xiaoya Li, Ping Nie, Fan Yin, Muyu Li, Qinghong Han, Xiaofei Sun, and Jiwei Li. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
BERT post-training for review reading comprehension and aspect-based sentiment analysis
Hu Xu, Bing Liu, Lei Shu, and Philip Yu. 2019 · 2019
Cited alongside, same era.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Cited alongside, same era.
Self-alignment pretraining for biomedical entity representations
Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Collier. 2021a · 2021
Closest in time.
Fast, effective, and self-supervised: Transforming masked language models into universal lexical and sentence encoders
Fangyu Liu, Ivan Vulić, Anna Korhonen, and Nigel Collier. 2021b · 2021
Closest in time.
SimCLS: A simple framework for contrastive learning of abstractive summarization
Yixin Liu and Pengfei Liu. 2021 · 2021
Closest in time.
Multilingual bert post-pretraining alignment
Lin Pan, Chung-Wei Hang, Haode Qi, Abhishek Shah, Saloni Potdar, and Mo Yu. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ELECTRA: pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Cited alongside, same era.
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
FLAT: chinese NER using flat-lattice transformer
Xiaonan Li, Hang Yan, Xipeng Qiu, and Xuanjing Huang. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Semantic re-tuning with contrastive tension
Fredrik Carlsson, Amaru Cuba Gyllensten, Evangelia Gogoulou, Erik Ylipää Hellqvist, and Magnus Sahlgren. 2021 · 2021
Cited alongside, same era.
Container: Few-shot named entity recognition via contrastive learning
Sarkar Snigdha Sarathi Das, Arzoo Katiyar, Rebecca J Passonneau, and Rui Zhang. 2021 · 2021
Cited alongside, same era.
Closest in time.
Non-autoregressive text generation with pre-trained language models
Yixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, and Nigel Collier. 2021a · 2021
Closest in time.
Dialogue response selection with hierarchical curriculum learning
Yixuan Su, Deng Cai, Qingyu Zhou, Zibo Lin, Simon Baker, Yunbo Cao, Shuming Shi, Nigel Collier, and Yan Wang. 2021b · 2021
Closest in time.
Few-shot table-to-text generation with prototype memory
Yixuan Su, Zaiqiao Meng, Simon Baker, and Nigel Collier. 2021c · 2021
Closest in time.
Keep the primary, rewrite the secondary: A two-stage approach for paraphrase generation
Yixuan Su, David Vandyke, Simon Baker, Yan Wang, and Nigel Collier. 2021e · 2021
Closest in time.
Plan-then-generate: Controlled data-to-text generation via planning
Yixuan Su, David Vandyke, Sihui Wang, Yimai Fang, and Nigel Collier. 2021f · 2021
Closest in time.
LexFit: Lexical fine-tuning of pretrained language models
Ivan Vulić, Edoardo Maria Ponti, Anna Korhonen, and Goran Glavaš. 2021 · 2021
Closest in time.
Phrase-bert: Improved phrase embeddings from bert with an application to corpus exploration
Shufan Wang, Laure Thompson, and Mohit Iyyer. 2021 · 2021
Closest in time.
Videoclip: Contrastive pre-training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021 · 2021
Closest in time.
ConSERT: A contrastive framework for self-supervised sentence representation transfer
Yuanmeng Yan, Rumei Li, Sirui Wang, Fuzheng Zhang, Wei Wu, and Weiran Xu. 2021 · 2021
Closest in time.
Taco: Token-aware cascade contrastive learning for video-text alignment
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. 2021 · 2021
Closest in time.
Dialoglm: Pre-trained model for long dialogue understanding and summarization
Ming Zhong, Yang Liu, Yichong Xu, Chenguang Zhu, and Michael Zeng. 2021 · 2021
Closest in time.
A contrastive framework for neural text generation
Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022 · 2022
Closest in time.