Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit remarkable fluency and competence across various natural language tasks.
How Can We Know What Language Models Know?
Jiang, Z.; Xu, F. F.; Araki, J.; and Neubig, G. 2020 · 1911
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, A.; Nangia, N.; and Bowman, S. R. 2017 · 2017
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2018 · 2018
Earlier work this paper cites.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Riedel, S.; et al. 2019 · 2019
Earlier work this paper cites.
PAWS: Paraphrase Adversaries from Word Scrambling
Zhang, Y.; Baldridge, J.; and He, L. 2019 · 2019
Earlier work this paper cites.
Prompt Consistency for Zero-Shot Task Generalization
Zhou, C.; He, J.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2022 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; et al. 2020 · 2020
Earlier work this paper cites.
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
He, P.; Liu, X.; Gao, J.; and Chen, W. 2020 · 2020
Earlier work this paper cites.
BLEURT: Learning Robust Metrics for Text Generation
Sellam, T.; Das, D.; and Parikh, A. P. 2020 · 2020
Earlier work this paper cites.
Measuring and Improving Consistency in Pretrained Language Models
Elazar, Y.; Kassner, N.; Ravfogel, S.; et al. 2021 · 2021
Earlier work this paper cites.
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
He, P.; Gao, J.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models
Jang, M.; Kwon, D. S.; and Lukasiewicz, T. 2021 · 2021
Earlier work this paper cites.
BeliefBank: Adding Memory to a Pre-Trained Language Model for a Systematic Notion of Belief
Kassner, N.; Tafjord, O.; Schütze, H.; and Clark, P. 2021 · 2021
Cited alongside, same era.
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; Webson, A.; Gu, S. S.; Dai, Z.; Suzgun, M.; Chen, X.; Chowdhery, A.; Castro-Ros, A.; Pellat, M.; Robinson, K.; Valter, D.; Narang, S.; Mishra, G.; Yu, A.; Zhao, V.; Huang, Y.; Dai, A.; Yu, H.; Petrov, S.; Chi, E. H.; Dean, J.; Devlin, J.; Roberts, A.; Zhou, D.; Le, Q. V.; and Wei, J. 2022 · 2022
Cited alongside, same era.
Factual Consistency of Multilingual Pretrained Language Models
Fierro, C.; and Søgaard, A. 2022 · 2022
Cited alongside, same era.
Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts
Guo, Y.; Yang, Y.; and Abbasi, A. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Taxonomy of Risks Posed by Language Models
Weidinger, L.; Uesato, J.; Rauh, M.; et al. 2022 · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; et al. 2022 · 2022
Later among the works it cites.
Let’s Sample Step by Step: Adaptive-Consistency for Efficient Reasoning with LLMs
Aggarwal, P.; Madaan, A.; Yang, Y.; and Mausam. 2023 · 2023
Closest in time.
Survey on Sociodemographic Bias in Natural Language Processing
Gupta, V.; Venkit, P. N.; Wilson, S.; and Passonneau, R. J. 2023 · 2023
Closest in time.
Keleg, A.; and Magdy, W. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference
Mitchell, E.; Noh, J.; Li, S.; Armstrong, W.; Agarwal, A.; Liu, P.; Finn, C.; and Manning, C. 2022 · 2022
Cited alongside, same era.
P-Adapters: Robustly Extracting Factual Information from Language Models with Diverse Prompts
Newman, B.; Choubey, P. K.; and Rajani, N. 2022 · 2022
Cited alongside, same era.
Measuring Reliability of Large Language Models through Semantic Consistency
Raj, H.; Rosati, D.; and Majumdar, S. 2022 · 2022
Cited alongside, same era.
Evaluating the Factual Consistency of Large Language Models Through Summarization
Tam, D.; Mascarenhas, A.; Zhang, S.; Kwan, S.; Bansal, M.; and Raffel, C. 2022 · 2022
Cited alongside, same era.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Wang, X.; Wei, J.; Schuurmans, D.; Le, Q.; Chi, E.; and Zhou, D. 2022 · 2022
Cited alongside, same era.
Do Prompt-Based Models Really Understand the Meaning of Their Prompts?
Webson, A.; and Pavlick, E. 2022 · 2022
Cited alongside, same era.
Kuhn, L.; Gal, Y.; and Farquhar, S. 2023 · 2023
Closest in time.
Illustrating Reinforcement Learning from Human Feedback (RLHF)
Lambart, N.; et al. 2023 · 2023
Closest in time.
How do text-davinci-002 and text-davinci-003 differ?
OpenAI. 2023 · 2023
Closest in time.
Prompting GPT-3 To Be Reliable
Si, C.; Gan, Z.; Yang, Z.; Wang, S.; Wang, J.; Boyd-Graber, J.; and Wang, L. 2023 · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; et al. 2023 · 2023
Closest in time.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; and Zhou, D. 2023 · 2023
Closest in time.