Fetching the paper…
Reading the bibliography…
Prompt-based probing has been widely used in evaluating the abilities of pretrained language models (PLMs).
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Models, reasoning and inference
Judea Pearl et al. 2000 · 2000
Earlier work this paper cites.
Correlating automated and human assessments of machine translation quality
Deborah Coughlin. 2003 · 2003
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
A survey of evaluation metrics used for NLG systems
Ananya B. Sai, Akash Kumar Mohankumar, and Mitesh M. Khapra. 2020 · 2008
Earlier work this paper cites.
Few-Shot Text Generation with Pattern-Exploiting Training
Timo Schick and Hinrich Schütze. 2020 · 2012
Earlier work this paper cites.
Illinois-LH: A denotational and distributional approach to semantics
Alice Lai and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
SemEval-2014 task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment
Marco Marelli, Luisa Bentivogli, Marco Baroni, Raffaella Bernardi, Stefano Menini, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro. 2016 · 2016
Earlier work this paper cites.
Annotating relation inference in context via question answering
Omer Levy and Ido Dagan. 2016 · 2016
Earlier work this paper cites.
Causal inference in statistics: A primer
Judea Pearl, Madelyn Glymour, and Nicholas P Jewell. 2016 · 2016
Earlier work this paper cites.
Avoiding discrimination through causal reasoning
Niki Kilbertus, Mateo Rojas-Carulla, Giambattista Parascandolo, Moritz Hardt, Dominik Janzing, and Bernhard Schölkopf. 2017 · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J. Kusner, Joshua R. Loftus, Chris Russell, and Ricardo Silva. 2017 · 2017
Earlier work this paper cites.
Adversarial learning for neural dialogue generation
Jiwei Li, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, and Dan Jurafsky. 2017 · 2017
Earlier work this paper cites.
The effect of different writing tasks on linguistic style: A case study of the ROC story cloze task
Roy Schwartz, Maarten Sap, Ioannis Konstas, Leila Zilles, Yejin Choi, and Noah A. Smith. 2017 · 2017
Earlier work this paper cites.
Neural natural language inference models enhanced with external knowledge
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Diana Inkpen, and Si Wei. 2018 · 2018
Earlier work this paper cites.
Visual referring expression recognition: What do systems actually learn?
Volkan Cirik, Louis-Philippe Morency, and Taylor Berg-Kirkpatrick. 2018 · 2018
Earlier work this paper cites.
Collecting diverse natural language inference problems for sentence representation evaluation
Adam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Event2Mind: Commonsense inference on events, intents, and reactions
Hannah Rashkin, Maarten Sap, Emily Allaway, Noah A. Smith, and Yejin Choi. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Commonsense knowledge mining from pretrained models
Joe Davison, Joshua Feldman, and Alexander Rush. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Later among the works it cites.
oLMpics-on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Later among the works it cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart M. Shieber. 2020 · 2020
Later among the works it cites.
Evaluating commonsense in pre-trained language models
Xuhui Zhou, Yue Zhang, Leyang Cui, and Dandan Huang. 2020 · 2020
Later among the works it cites.
Shortcutted commonsense: Data spuriousness in deep learning of commonsense reasoning
Ruben Branco, António Branco, João António Rodrigues, and João Ricardo Silva. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Re-evaluating ADEM: A deeper look at scoring dialogue responses
Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, and Mukundhan Srinivasan. 2019 · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models
Allyson Ettinger. 2020 · 2020
Cited alongside, same era.
Knowledgeable or educated guess? revisiting language models as knowledge bases
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021 · 2021
Later among the works it cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Amir Feder, Katherine A Keith, Emaad Manzoor, Reid Pryzant, Dhanya Sridhar, Zach Wood-Doughty, Jacob Eisenstein, Justin Grimmer, Roi Reichart, Margaret E Roberts, et al. 2021 · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Later among the works it cites.
BERTese: Learning to speak to BERT
Adi Haviv, Jonathan Berant, and Amir Globerson. 2021 · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
Visually grounded reasoning across languages and cultures
Fangyu Liu, Emanuele Bugliarello, Edoardo Maria Ponti, Siva Reddy, Nigel Collier, and Desmond Elliott. 2021a · 2021
Later among the works it cites.
Learning how to ask: Querying LMs with mixtures of soft prompts
Guanghui Qin and Jason Eisner. 2021 · 2021
Later among the works it cites.
ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Yu Sun, Shuohuan Wang, Shikun Feng, Siyu Ding, Chao Pang, Junyuan Shang, Jiaxiang Liu, Xuyi Chen, Yanbin Zhao, Yuxiang Lu, Weixin Liu, Zhihua Wu, Weibao Gong, Jianzhong Liang, Zhizhou Shang, Peng Sun, Wei Liu, Xuan Ouyang, Dianhai Yu, Hao Tian, Hua Wu, and Haifeng Wang. 2021 · 2021
Later among the works it cites.
Can language models be biomedical knowledge bases?
Mujeen Sung, Jinhyuk Lee, Sean Yi, Minji Jeon, Sungdong Kim, and Jaewoo Kang. 2021 · 2021
Later among the works it cites.
Factual probing is [MASK]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021 · 2021
Later among the works it cites.