Fetching the paper…
Reading the bibliography…
Much of the recent discourse within the ML community has been centered around Large Language Models (LLMs), their functionality and potential -- yet not only do we not have a working definition of LLMs, but much of this discourse relies on claims and assumptions that are worth re-examining.
On the thresholds of knowledge
Lenat, D. and Feigenbaum, E. A · 1981
Earlier work this paper cites.
Economic transformations: general purpose technologies and long-term economic growth
Lipsey, R. G., Carlaw, K. I., and Bekar, C. T · 2005
Earlier work this paper cites.
A Survey on Transfer Learning
Pan, S. J. and Yang, Q · 2009
Earlier work this paper cites.
Syntactic annotations for the google books ngram corpus
Lin, Y., Michel, J.-B., Lieberman, E. A., Orwant, J., Brockman, W., and Petrov, S · 2012
Earlier work this paper cites.
One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling
Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T · 2013
Earlier work this paper cites.
Cohesion in English
Halliday, M. and Hasan, R · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
General purpose technologies in theory, application and controversy: a review
Bekar, C., Carlaw, K., and Lipsey, R · 2018
Earlier work this paper cites.
Three dimensions of reproducibility in natural language processing
Cohen, K. B., Xia, J., Zweigenbaum, P., Callahan, T., Hargraves, O., Goss, F., Ide, N., Névéol, A., Grouin, C., and Hunter, L · 2018
Earlier work this paper cites.
Questionable Answers in Question Answering Research: Reproducibility and Variability of Published Results
Crane, M · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
SciBERT: A pretrained language model for scientific text
Beltagy, I., Lo, K., and Cohan, A · 2019
Earlier work this paper cites.
Unreproducible research is reproducible
Bouthillier, X., Laurent, C., and Vincent, P · 2019
Earlier work this paper cites.
Artificial intelligence technologies and aggregate growth prospects
Bresnahan, T · 2019
Earlier work this paper cites.
On the Measure of Intelligence
Chollet, F · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Show Your Work: Improved Reporting of Experimental Results
Dodge, J., Gururangan, S., Card, D., Schwartz, R., and Smith, N. A · 2019
Earlier work this paper cites.
Why deep-learning AIs are so easy to fool
Heaven, D. et al · 2019
Earlier work this paper cites.
Using pre-training can improve model robustness and uncertainty
Hendrycks, D., Lee, K., and Mazeika, M · 2019
Earlier work this paper cites.
Learning The Difference That Makes A Difference With Counterfactually-Augmented Data
Kaushik, D., Hovy, E., and Lipton, Z · 2019
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in Natural Language Inference
McCoy, R. T., Pavlick, E., and Linzen, T · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
The bitter lesson (blog post)
Sutton, R · 2019
Earlier work this paper cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Earlier work this paper cites.
Ahmed, N. and Wahed, M · 2020
Earlier work this paper cites.
A Survey on Transfer Learning in Natural Language Processing
Alyafeai, Z., AlShaibani, M. S., and Ahmad, I · 2020
Earlier work this paper cites.
Bianchini, S., Müller, M., and Pelletier, P · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Earlier work this paper cites.
Utility is in the eye of the user: A critique of nlp leaderboards
Ethayarajh, K. and Jurafsky, D · 2020
Earlier work this paper cites.
Evaluating NLP Models via Contrast Sets
Gardner, M., Artzi, Y., Basmova, V., Berant, J., Bogin, B., Chen, S., Dasigi, P., Dua, D., Elazar, Y., Gottumukkala, A., Gupta, N., Hajishirzi, H., Ilharco, G., Khashabi, D., Lin, K., Liu, J., Liu, N. F., Mulcaire, P., Ning, Q., Singh, S., Smith, N. A., Subramanian, S., Tsarfaty, R., Wallace, E., Zhang, A., and Zhou, B · 2020
Earlier work this paper cites.
Overview of the Transformer-based Models for NLP Tasks
Gillioz, A., Casas, J., Mugellini, E., and Abou Khaled, O · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for natural language understanding
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance
McCoy, R. T., Min, J., and Linzen, T · 2020
Earlier work this paper cites.
Imagining the thinking machine: Technological myths and the rise of artificial intelligence
Natale, S. and Ballatore, A · 2020
Earlier work this paper cites.
Emergent Properties
O’Connor, T. and Wong, H. Y · 2020
Earlier work this paper cites.
Meta-KD: A meta knowledge distillation framework for language model compression across domains
Pan, H., Wang, C., Qiu, M., Zhang, Y., Li, Y., and Huang, J · 2020
Earlier work this paper cites.
Neural Unsupervised Domain Adaptation in NLP—A Survey
Ramponi, A. and Plank, B · 2020
Earlier work this paper cites.
Beyond Accuracy: Behavioral Testing of NLP Models with CheckList
Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S · 2020
Earlier work this paper cites.
Getting closer to AI complete question answering: A set of prerequisite real tasks
Rogers, A., Kovaleva, O., Downey, M., and Rumshisky, A · 2020
Earlier work this paper cites.
A Comprehensive Survey on Transfer Learning
Zhuang, F., Qi, Z., Duan, K., Xi, D., Zhu, Y., Zhu, H., Xiong, H., and He, Q · 2020
Earlier work this paper cites.
The Grey Hoodie Project: Big tobacco, big tech, and the threat on academic integrity
Abdalla, M. and Abdalla, M · 2021
Earlier work this paper cites.
Persistent anti-Muslim bias in Large Language Models
Abid, A., Farooqi, M., and Zou, J · 2021
Earlier work this paper cites.
A systematic review of reproducibility research in Natural Language Processing
Belz, A., Agarwal, S., Shimorina, A., and Reiter, E · 2021
Earlier work this paper cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
Generalization in NLI: Ways (Not) To Go Beyond Simple Heuristics
Bhargava, P., Drozd, A., and Rogers, A · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Shortcutted commonsense: Data spuriousness in deep learning of commonsense reasoning
Branco, R., Branco, A., Rodrigues, J., and Silva, J · 2021
Earlier work this paper cites.
Artificial intelligence as a general-purpose technology: an historical perspective
Crafts, N · 2021
Earlier work this paper cites.
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus, 2021
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D · 2021
Cited alongside, same era.
Competency problems: On finding and removing artifacts in language data
Gardner, M., Merrill, W., Dodge, J., Peters, M. E., Ross, A., Singh, S., and Smith, N. A · 2021
Cited alongside, same era.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C · 2021
Cited alongside, same era.
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Cited alongside, same era.
Question and answer test-train overlap in open-domain question answering datasets
Lewis, P., Stenetorp, P., and Riedel, S · 2021
Cited alongside, same era.
Sparks of Artificial General Intelligence: Early experiments with GPT-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4, 2023
Chang, K. K., Cramer, M., Soni, S., and Bamman, D · 2023
Closest in time.
Language model behavior: A comprehensive survey
Chang, T. A. and Bergen, B. K · 2023
Closest in time.
How is ChatGPT’s behavior changing over time?
Chen, L., Zaharia, M., and Zou, J · 2023
Closest in time.
ML.ENERGY leaderboard, 2023
Chung, J.-W., Liu, J., Wu, Z., Xia, Y., and Chowdhury, M · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2021
Cited alongside, same era.
Language Model Evaluation Beyond Perplexity
Meister, C. and Cotterell, R · 2021
Cited alongside, same era.
Show Your Work: Scratchpads for Intermediate Computation with Language Models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Cited alongside, same era.
True Few-Shot Learning with Language Models
Perez, E., Kiela, D., and Cho, K · 2021
Cited alongside, same era.
ICT’s Wide Web: a System-Level Analysis of ICT’s Industrial Diffusion with Algorithmic Links
Prytkova, E · 2021
Cited alongside, same era.
AI and the Everything in the Whole Wide World Benchmark
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A · 2021
Cited alongside, same era.
"Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI
Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., and Aroyo, L. M · 2021
Cited alongside, same era.
Honey, I shrunk the language: Language model behavior at reduced scale
Deshpande, V., Pechi, D., Thatte, S., Lialin, V., and Rumshisky, A · 2023
Closest in time.
GPTs are GPTs: An early look at the labor market impact potential of large language models
Eloundou, T., Manning, S., Mishkin, P., and Rock, D · 2023
Closest in time.
Emnlp 2023 call for main conference papers theme track: Large language models and the future of nlp, 2023
EMNLP · 2023
Closest in time.
Sensitivity and Robustness of Large Language Models to Prompt Template in Japanese Text Classification Tasks, June 2023
Gan, C. and Mori, T · 2023
Closest in time.
Is the world ready for ChatGPT therapists?
Graber-Stiehl, I · 2023
Closest in time.
Textbooks are all you need, 2023
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Giorno, A. D., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T., and Li, Y · 2023
Closest in time.
Towards inferential reproducibility of machine learning research, 2023
Hagmann, M., Meier, P., and Riezler, S · 2023
Closest in time.
Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine
Harrer, S · 2023
Closest in time.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2023
Closest in time.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Closest in time.
ChatGPT for good? on opportunities and challenges of large language models for education
Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., et al · 2023
Closest in time.
GPT-4 Passes the Bar Exam, March 2023
Katz, D. M., Bommarito, M. J., Gao, S., and Arredondo, P · 2023
Closest in time.
OpenAI’s CEO Says the Age of Giant AI Models Is Already Over
Knight, W · 2023
Closest in time.
How Enterprises Can Become Ready To Work With LLMs
Kuttan, J · 2023
Closest in time.
Ethics and deep learning, 2023
LaCroix, T. and Prince, S. J. D · 2023
Closest in time.
ChatGPT Beyond English: Towards a Comprehensive Evaluation of Large Language Models in Multilingual Learning, 2023
Lai, V. D., Ngo, N. T., Veyseh, A. P. B., Man, H., Dernoncourt, F., Bui, T., and Nguyen, T. H · 2023
Closest in time.
BLOOM: A 176b-parameter open-access multilingual language model
Le Scao, T., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2023
Closest in time.
Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine
Lee, P., Bubeck, S., and Petro, J · 2023
Closest in time.
Are emergent abilities in large language models just in-context learning?, 2023
Lu, S., Bigoulaeva, I., Sachdeva, R., Madabushi, H. T., and Gurevych, I · 2023
Closest in time.
Reproducibility in NLP: What have we learned from the checklist?
Magnusson, I., Smith, N. A., and Dodge, J · 2023
Closest in time.
Data Portraits: Recording Foundation Model Training Data, March 2023
Marone, M. and Van Durme, B · 2023
Closest in time.
Re-Evaluating GPT-4’s Bar Exam Performance, May 2023
Martínez, E · 2023
Closest in time.
Embers of autoregression: Understanding large language models through the problem they are trained to solve, 2023
McCoy, R. T., Yao, S., Friedman, D., Hardy, M., and Griffiths, T. L · 2023
Closest in time.
News coverage of artificial intelligence reflects business and government hype — not critical voices
McKelvey, F., Dandurand, G., and Roberge, J · 2023
Closest in time.
Why GPT Should Stand For ‘General Purpose Technology’ For All
McKendrick, J · 2023
Closest in time.
GPT-4 Technical Report, 2023
OpenAI · 2023
Closest in time.
The ROOTS Search Tool: Data Transparency for LLMs
Piktus, A., Akiki, C., Villegas, P., Laurençon, H., Dupont, G., Luccioni, A. S., Jernite, Y., and Rogers, A · 2023
Closest in time.
Closed AI Models Make Bad Baselines
Rogers, A · 2023
Closest in time.
Program chairs’ report on peer review at ACL 2023
Rogers, A., Karpinska, M., Boyd-Graber, J., and Okazaki, N · 2023
Closest in time.
How does ChatGPT really work?, 2023
Roose, K · 2023
Closest in time.
Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs, 2023
Sap, M., LeBras, R., Fried, D., and Choi, Y · 2023
Closest in time.
Are Emergent Abilities of Large Language Models a Mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2023
Closest in time.
Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models, May 2023
Shapira, N., Levy, M., Alavi, S. H., Zhou, X., Choi, Y., Goldberg, Y., Sap, M., and Shwartz, V · 2023
Closest in time.
The gradient of generative AIrelease: Methods and considerations
Solaiman, I · 2023
Closest in time.
Training Large Language Models Efficiently with Sparsity and Dataflow
Srinivasan, V., Gandhi, D., Thakker, U., and Prabhakar, R · 2023
Closest in time.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, 2023
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., Kluska, A., Lewkowycz, A., Agarwal, A., Power, A., Ray, A., Warstadt, A., Kocurek, A. W., Safaya, A., Tazarv, A., Xiang, A., Parrish, A., Nie, A., Hussain, A., Askell, A., Dsouza, A., Slone, A., Rahane, A., Iyer, A. S., Andreassen, A., Madotto, A., Santilli, A., Stuhlmüller, A., Dai, A., La, A., Lampinen, A., Zou, A., Jiang, A., Chen, A., Vuong, A., Gupta, A., Gottardi, A., Norelli, A., Venkatesh, A., Gholamidavoodi, A., Tabassum, A., Menezes, A., Kirubarajan, A., Mullokandov, A., Sabharwal, A., Herrick, A., Efrat, A., Erdem, A., Karakaş, A., Roberts, B. R., Loe, B. S., Zoph, B., Bojanowski, B., Özyurt, B., Hedayatnia, B., Neyshabur, B., Inden, B., Stein, B., Ekmekci, B., Lin, B. Y., Howald, B., Orinion, B., Diao, C., Dour, C., Stinson, C., Argueta, C., Ramírez, C. F., Singh, C., Rathkopf, C., Meng, C., Baral, C., Wu, C., Callison-Burch, C., Waites, C., Voigt, C., Manning, C. D., Potts, C., Ramirez, C., Rivera, C. E., Siro, C., Raffel, C., Ashcraft, C., Garbacea, C., Sileo, D., Garrette, D., Hendrycks, D., Kilman, D., Roth, D., Freeman, D., Khashabi, D., Levy, D., González, D. M., Perszyk, D., Hernandez, D., Chen, D., Ippolito, D., Gilboa, D., Dohan, D., Drakard, D., Jurgens, D., Datta, D., Ganguli, D., Emelin, D., Kleyko, D., Yuret, D., Chen, D., Tam, D., Hupkes, D., Misra, D., Buzan, D., Mollo, D. C., Yang, D., Lee, D.-H., Schrader, D., Shutova, E., Cubuk, E. D., Segal, E., Hagerman, E., Barnes, E., Donoway, E., Pavlick, E., Rodola, E., Lam, E., Chu, E., Tang, E., Erdem, E., Chang, E., Chi, E. A., Dyer, E., Jerzak, E., Kim, E., Manyasi, E. E., Zheltonozhskii, E., Xia, F., Siar, F., Martínez-Plumed, F., Happé, F., Chollet, F., Rong, F., Mishra, G., Winata, G. I., de Melo, G., Kruszewski, G., Parascandolo, G., Mariani, G., Wang, G., Jaimovitch-López, G., Betz, G., Gur-Ari, G., Galijasevic, H., Kim, H., Rashkin, H., Hajishirzi, H., Mehta, H., Bogar, H., Shevlin, H., Schütze, H., Yakura, H., Zhang, H., Wong, H. M., Ng, I., Noble, I., Jumelet, J., Geissinger, J., Kernion, J., Hilton, J., Lee, J., Fisac, J. F., Simon, J. B., Koppel, J., Zheng, J., Zou, J., Kocoń, J., Thompson, J., Wingfield, J., Kaplan, J., Radom, J., Sohl-Dickstein, J., Phang, J., Wei, J., Yosinski, J., Novikova, J., Bosscher, J., Marsh, J., Kim, J., Taal, J., Engel, J., Alabi, J., Xu, J., Song, J., Tang, J., Waweru, J., Burden, J., Miller, J., Balis, J. U., Batchelder, J., Berant, J., Frohberg, J., Rozen, J., Hernandez-Orallo, J., Boudeman, J., Guerr, J., Jones, J., Tenenbaum, J. B., Rule, J. S., Chua, J., Kanclerz, K., Livescu, K., Krauth, K., Gopalakrishnan, K., Ignatyeva, K., Markert, K., Dhole, K. D., Gimpel, K., Omondi, K., Mathewson, K., Chiafullo, K., Shkaruta, K., Shridhar, K., McDonell, K., Richardson, K., Reynolds, L., Gao, L., Zhang, L., Dugan, L., Qin, L., Contreras-Ochando, L., Morency, L.-P., Moschella, L., Lam, L., Noble, L., Schmidt, L., He, L., Colón, L. O., Metz, L., Şenel, L. K., Bosma, M., Sap, M., ter Hoeve, M., Farooqi, M., Faruqui, M., Mazeika, M., Baturan, M., Marelli, M., Maru, M., Quintana, M. J. R., Tolkiehn, M., Giulianelli, M., Lewis, M., Potthast, M., Leavitt, M. L., Hagen, M., Schubert, M., Baitemirova, M. O., Arnaud, M., McElrath, M., Yee, M. A., Cohen, M., Gu, M., Ivanitskiy, M., Starritt, M., Strube, M., Swędrowski, M., Bevilacqua, M., Yasunaga, M., Kale, M., Cain, M., Xu, M., Suzgun, M., Walker, M., Tiwari, M., Bansal, M., Aminnaseri, M., Geva, M., Gheini, M., T, M. V., Peng, N., Chi, N. A., Lee, N., Krakover, N. G.-A., Cameron, N., Roberts, N., Doiron, N., Martinez, N., Nangia, N., Deckers, N., Muennighoff, N., Keskar, N. S., Iyer, N. S., Constant, N., Fiedel, N., Wen, N., Zhang, O., Agha, O., Elbaghdadi, O., Levy, O., Evans, O., Casares, P. A. M., Doshi, P., Fung, P., Liang, P. P., Vicol, P., Alipoormolabashi, P., Liao, P., Liang, P., Chang, P., Eckersley, P., Htut, P. M., Hwang, P., Miłkowski, P., Patil, P., Pezeshkpour, P., Oli, P., Mei, Q., Lyu, Q., Chen, Q., Banjade, R., Rudolph, R. E., Gabriel, R., Habacker, R., Risco, R., Millière, R., Garg, R., Barnes, R., Saurous, R. A., Arakawa, R., Raymaekers, R., Frank, R., Sikand, R., Novak, R., Sitelew, R., LeBras, R., Liu, R., Jacobs, R., Zhang, R., Salakhutdinov, R., Chi, R., Lee, R., Stovall, R., Teehan, R., Yang, R., Singh, S., Mohammad, S. M., Anand, S., Dillavou, S., Shleifer, S., Wiseman, S., Gruetter, S., Bowman, S. R., Schoenholz, S. S., Han, S., Kwatra, S., Rous, S. A., Ghazarian, S., Ghosh, S., Casey, S., Bischoff, S., Gehrmann, S., Schuster, S., Sadeghi, S., Hamdan, S., Zhou, S., Srivastava, S., Shi, S., Singh, S., Asaadi, S., Gu, S. S., Pachchigar, S., Toshniwal, S., Upadhyay, S., Shyamolima, Debnath, Shakeri, S., Thormeyer, S., Melzi, S., Reddy, S., Makini, S. P., Lee, S.-H., Torene, S., Hatwar, S., Dehaene, S., Divic, S., Ermon, S., Biderman, S., Lin, S., Prasad, S., Piantadosi, S. T., Shieber, S. M., Misherghi, S., Kiritchenko, S., Mishra, S., Linzen, T., Schuster, T., Li, T., Yu, T., Ali, T., Hashimoto, T., Wu, T.-L., Desbordes, T., Rothschild, T., Phan, T., Wang, T., Nkinyili, T., Schick, T., Kornev, T., Tunduny, T., Gerstenberg, T., Chang, T., Neeraj, T., Khot, T., Shultz, T., Shaham, U., Misra, V., Demberg, V., Nyamai, V., Raunak, V., Ramasesh, V., Prabhu, V. U., Padmakumar, V., Srikumar, V., Fedus, W., Saunders, W., Zhang, W., Vossen, W., Ren, X., Tong, X., Zhao, X., Wu, X., Shen, X., Yaghoobzadeh, Y., Lakretz, Y., Song, Y., Bahri, Y., Choi, Y., Yang, Y., Hao, Y., Chen, Y., Belinkov, Y., Hou, Y., Hou, Y., Bai, Y., Seid, Z., Zhao, Z., Wang, Z., Wang, Z. J., Wang, Z., and Wu, Z · 2023
Closest in time.
Is ChatGPT good at search? investigating large language models as re-ranking agent
Sun, W., Yan, L., Ma, X., Ren, P., Yin, D., and Ren, Z · 2023
Closest in time.
GPT-RE: In-context learning for relation extraction using large language models
Wan, Z., Cheng, F., Mao, Z., Liu, Q., Song, H., Li, J., and Kurohashi, S · 2023
Closest in time.
GPT-NER: Named Entity Recognition via Large Language Models, 2023
Wang, S., Sun, X., Li, X., Ouyang, R., Wu, F., Zhang, T., Li, J., and Wang, G · 2023
Closest in time.
Common arguments regarding emergent abilities, May 2023
Wei, J · 2023
Closest in time.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks, 2023
Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y · 2023
Closest in time.
OpenAI CEO tells Senate that he fears AI’s potential to manipulate views, 2023
Zakrzewski, C. e. a · 2023
Closest in time.
Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models, 2023
Zhang, S. J., Florin, S., Lee, A. N., Niknafs, E., Marginean, A., Wang, A., Tyser, K., Chin, Z., Hicke, Y., Singh, N., Udell, M., Kim, Y., Buonassisi, T., Solar-Lezama, A., and Drori, I · 2023
Closest in time.
Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis, 2023
Zhu, W., Liu, H., Dong, Q., Xu, J., Huang, S., Kong, L., Chen, J., and Li, L · 2023
Closest in time.
Multi-VALUE: A Framework for Cross-Dialectal English NLP, 2023
Ziems, C., Held, W., Yang, J., Dhamala, J., Gupta, R., and Yang, D · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Closest in time.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al · 2024
Closest in time.
Speech and Language Processing (3rd Ed. Draft)
Jurafsky, D. and Martin, J. H · 2024
Closest in time.
Dolma: An open corpus of three trillion tokens for language model pretraining research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., et al · 2024
Closest in time.