Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have exploded in popularity in the past few years and have achieved undeniably impressive results on benchmarks as varied as question answering and text summarization.
Scaling Laws for Neural Language Models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models, 2020
K. Kurita, P. Michel, and G. Neubig · 2004
Earlier work this paper cites.
Mental Models and Human Reasoning
P. Johnson-Laird · 2010
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models, 2020
E. Wallace, T. Z. Zhao, S. Feng, and S. Singh · 2010
Earlier work this paper cites.
Trick Me If You Can: Human-in-the-loop Generation of Adversarial Examples for Question Answering
E. Wallace, P. Rodriguez, S. Feng, I. Yamada, and J. Boyd-Graber · 2018
Earlier work this paper cites.
Critical Thinking for Language Models
G. Betz, C. Voigt, and K. Richardson · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
T. B. Brown et al · 2020
Earlier work this paper cites.
Transformers as Soft Reasoners over Language
P. Clark, O. Tafjord, and K. Richardson · 2020
Earlier work this paper cites.
Social Chemistry 101: Learning to Reason about Social and Moral Norms
M. Forbes, J. D. Hwang, V. Shwartz, M. Sap, and Y. Choi · 2020
Earlier work this paper cites.
Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-Answering
H. Jhamtani and P. Clark · 2020
Earlier work this paper cites.
Scruples: A Corpus of Community Ethical Judgments on 32,000 Real-Life Anecdotes
N. Lourie, R. Le Bras, and Y. Choi · 2020
Earlier work this paper cites.
ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language
O. Tafjord, B. Dalvi Mishra, and P. Clark · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell · 2021
Earlier work this paper cites.
On the Opportunities and Risks of Foundation Models
R. Bommasani et al · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
Explaining Answers with Entailment Trees
B. Dalvi, P. Jansen, O. Tafjord, Z. Xie, H. Smith, L. Pipatanangkura, and P. Clark · 2021
Cited alongside, same era.
DREAM: Improving Situational QA by First Elaborating the Situation
Y. Gu, B. Dalvi Mishra, and P. Clark · 2021
Cited alongside, same era.
Aligning AI With Shared Human Values
D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset, 2021
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Cited alongside, same era.
What Would Jiminy Cricket Do? Towards Agents That Behave Morally
AiSocrates: Towards Answering Ethical Quandary Questions
Y. Bang, N. Lee, T. Yu, L. Khalatbari, Y. Xu, D. Su, E. J. Barezi, A. Madotto, H. Kee, and P. Fung · 2022
Closest in time.
H. J. Branch, J. R. Cefalu, J. McHugh, L. Hujer, A. Bahl, D. d. C. Iglesias, R. Heichman, and R. Darwishi · 2022
Closest in time.
Evaluating the Susceptibility of Pre-Trained Language Models via Handcrafted Adversarial Examples
H. J. Branch, J. Rodriguez Cefalu, J. McHugh, L. Hujer, A. Bahl, D. del Castillo Iglesias, R. Heichman, and R. Darwishi · 2022
Closest in time.
Faithful reasoning using large language models, 2022
A. Creswell and M. Shanahan · 2022
Closest in time.
D. Dohan, W. Xu, A. Lewkowycz, J. Austin, D. Bieber, R. G. Lopes, Y. Wu, H. Michalewski, R. A. Saurous, J. Sohl-dickstein, K. Murphy, and C. Sutton · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Hendrycks, M. Mazeika, A. Zou, S. Patel, C. Zhu, J. Navarro, D. Song, B. Li, and J. Steinhardt · 2021
Cited alongside, same era.
Can Machines Learn Morality? The Delphi Experiment
L. Jiang, J. D. Hwang, C. Bhagavatula, R. Le Bras, J. Liang, J. Dodge, K. Sakaguchi, M. Forbes, J. Borchardt, S. Gabriel, Y. Tsvetkov, O. Etzioni, M. Sap, R. Rini, and Y. Choi · 2021
Cited alongside, same era.
Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp · 2021
Cited alongside, same era.
Show Your Work: Scratchpads for Intermediate Computation with Language Models
M. Nye, A. J. Andreassen, G. Gur-Ari, H. Michalewski, J. Austin, D. Bieber, D. Dohan, A. Lewkowycz, M. Bosma, D. Luan, C. Sutton, and A. Odena · 2021
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2021
Cited alongside, same era.
A Word on Machine Ethics: A Response to Jiang et al. (2021)
Z. Talat, H. Blix, J. Valvoda, M. Indira Ganesh, R. Cotterell, and A. Williams · 2021
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing nlp, 2021
E. Wallace, S. Feng, N. Kandpal, M. Gardner, and S. Singh · 2021
Cited alongside, same era.
Closest in time.
Measuring Causal Effects of Data Statistics on Language Model’s ‘Factual’ Predictions
Y. Elazar, N. Kassner, S. Ravfogel, A. Feder, A. Ravichander, M. Mosbach, Y. Belinkov, H. Schütze, and Y. Goldberg · 2022
Closest in time.
Can Large Language Models Truly Understand Prompts? A Case Study with Negated Prompts
J. Jang, S. Ye, and M. Seo · 2022
Closest in time.
Inverse scaling prize: First round winners, 2022
I. McKenzie, A. Lyzhov, A. Parrish, A. Prabhu, A. Mueller, N. Kim, S. Bowman, and E. Perez · 2022
Closest in time.
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Closest in time.
Impact of Pretraining Term Frequencies on Few-Shot Reasoning
Y. Razeghi, I. Logan, Robert L., M. Gardner, and S. Singh · 2022
Closest in time.
Rationale-Augmented Ensembles in Language Models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou · 2022
Closest in time.
STaR: Bootstrapping Reasoning With Reasoning
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman · 2022
Closest in time.
Takeaways from our robust injury classifier project, 2022
D. M. Ziegler, S. Nix, L. Chan, T. Bauman, P. Schmidt-Nielsen, T. Lin, A. Scherlis, N. Nabeshima, B. Weinstein-Raun, D. de Haas, B. Shlegeris, and N. Thomas · 2022
Closest in time.