Fetching the paper…
Reading the bibliography…
The widespread public deployment of large language models (LLMs) in recent months has prompted a wave of new attention and engagement from advocates, policymakers, and scholars from many fields.
American parenting of language-learning children: Persisting differences in family-child interactions observed in natural home environments
Hart, B. and Risley, T. R · 1992
Earlier work this paper cites.
The Winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Concrete problems in AI safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Deep learning
Bengio, Y., Goodfellow, I., and Courville, A · 2017
Earlier work this paper cites.
Mapping the early language environment using all-day recordings and automated analysis
Gilkerson, J., Richards, J. A., Warren, S. F., Montgomery, J. K., Greenwood, C. R., Kimbrough Oller, D., Hansen, J. H., and Paul, T. D · 2017
Earlier work this paper cites.
Pathologies of neural models make interpretations difficult
Feng, S., Wallace, E., Grissom II, A., Iyyer, M., Rodriguez, P., and Boyd-Graber, J · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Lipton, Z. C · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Dinan, E., Humeau, S., Chintagunta, B., and Weston, J · 2019
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S · 2019
Earlier work this paper cites.
Attention is not Explanation
Jain, S. and Wallace, B. C · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
McCoy, T., Pavlick, E., and Linzen, T · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Bender, E. M. and Koller, A · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Underspecification presents challenges for credibility in modern machine learning
D’Amour, A., Heller, K., Moldovan, D., Adlam, B., Alipanahi, B., Beutel, A., Chen, C., Deaton, J., Eisenstein, J., Hoffman, M. D., et al · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift
Krueger, D., Maharaj, T., and Leike, J · 2020
Earlier work this paper cites.
To dissect an octopus: Making sense of the form/meaning debate
Michael, J · 2020
Earlier work this paper cites.
Is it possible for language models to achieve language understanding
Potts, C · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Designing AI with Rights, Consciousness, Self-Respect, and Freedom
Schwitzgebel, E. and Garza, M · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F · 2020
Earlier work this paper cites.
Can language models encode perceptual structure without grounding? a case study in color
Abdou, M., Kulmizev, A., Hershcovich, D., Frank, S., Pavlick, E., and Søgaard, A · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
An interpretability illusion for BERT
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., et al · 2021
Earlier work this paper cites.
A survey of race, racism, and anti-racism in NLP
Field, A., Blodgett, S. L., Waseem, Z., and Tsvetkov, Y · 2021
Earlier work this paper cites.
DeBERTa: Decoding-enhanced BERT with Disentangled Attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Earlier work this paper cites.
Alignment of language agents
Kenton, Z., Everitt, T., Weidinger, L., Gabriel, I., Mikulik, V., and Irving, G · 2021
Earlier work this paper cites.
Implicit representations of meaning in neural language models
Li, B. Z., Nye, M., and Andreas, J · 2021
Earlier work this paper cites.
Do language models know the way to Rome?
Liétard, B., Abdou, M., and Søgaard, A · 2021
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., et al · 2021
Cited alongside, same era.
Shaking the foundations: delusions in sequence models for interaction and control
Ortega, P. A., Kunesch, M., Delétang, G., Genewein, T., Grau-Moya, J., Veness, J., Buchli, J., Degrave, J., Piot, B., Perolat, J., et al · 2021
Cited alongside, same era.
Sorting through the noise: Testing robustness of information processing in pre-trained language models
Pandia, L. and Ettinger, A · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Prompt programming for large language models: Beyond the few-shot paradigm
Reynolds, L. and McDonell, K · 2021
Cited alongside, same era.
Skill induction and planning with latent language
Sharma, P., Torralba, A., and Andreas, J · 2022
Later among the works it cites.
Language models seem to be much better than humans at next-token prediction
Shlegeris, B., Roger, F., and Chan, L · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
2022 expert survey on progress in AI
Stein-Perlman, Z., Weinstein-Raun, B., and Grace, K · 2022
Later among the works it cites.
AI forecasting: One year in
Steinhardt, J · 2022
Later among the works it cites.
Parametrically retargetable decision-makers tend to seek power
Turner, A. and Tadepalli, P · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How could we know when a robot was a moral patient?
Shevlin, H · 2021
Cited alongside, same era.
On the risks of emergent behavior in foundation models
Steinhardt, J · 2021
Cited alongside, same era.
Optimal policies tend to seek power
Turner, A. M., Smith, L. R., Shah, R., Critch, A., and Tadepalli, P · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Language models as agent models
Andreas, J · 2022
Cited alongside, same era.
The values encoded in machine learning research
Birhane, A., Kalluri, P., Card, D., Agnew, W., Dotan, R., and Bao, M · 2022
Cited alongside, same era.
The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail
Bowman, S · 2022
Cited alongside, same era.
Solving math word problems with process-and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Later among the works it cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Wang, B., Min, S., Deng, X., Shen, J., Wu, Y., Zettlemoyer, L., and Sun, H · 2022
Later among the works it cites.
Do prompt-based models really understand the meaning of their prompts?
Webson, A. and Pavlick, E · 2022
Later among the works it cites.
Taxonomy of risks posed by language models
Weidinger, L., Uesato, J., Rauh, M., Griffin, C., Huang, P.-S., Mellor, J., Glaese, A., Cheng, M., Balle, B., Kasirzadeh, A., Biles, C., Brown, S., Kenton, Z., Hawkins, W., Stepleton, T., Birhane, A., Hendricks, L. A., Rimell, L., Isaac, W., Haas, J., Legassick, S., Irving, G., and Gabriel, I · 2022
Later among the works it cites.
As ChatGPT’s popularity explodes, U.S. lawmakers take an interest
Bartz, D · 2023
Closest in time.
Pause giant AI experiments
Bengio, Y., Russell, S., Musk, E., Wozniak, S., et al · 2023
Closest in time.
ChatGPT and the future of medical writing
Biswas, S · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with GPT-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2023
Closest in time.
Microsoft announces new multibillion-dollar investment in ChatGPT-maker OpenAI
Capoot, A · 2023
Closest in time.
Could a large language model be conscious?
Chalmers, D. J · 2023
Closest in time.
Harms from increasingly agentic algorithmic systems
Chan, A., Salganik, R., Markelius, A., Pang, C., Rajkumar, N., Krasheninnikov, D., Langosco, L., He, Z., Duan, Y., Carroll, M., et al · 2023
Closest in time.
ChatGPT goes to law school
Choi, J. H., Hickman, K. E., Monahan, A., and Schwarcz, D · 2023
Closest in time.
PaLM-E: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Closest in time.
The capacity for moral self-correction in large language models
Ganguli, D., Askell, A., Schiefer, N., Liao, T., Lukošiūtė, K., Chen, A., Goldie, A., Mirhoseini, A., Olsson, C., Hernandez, D., et al · 2023
Closest in time.
ChatGPT and large language models: what’s the risk?
J, P. and C, D · 2023
Closest in time.
This changes everything
Klein, E · 2023
Closest in time.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R., Buckley, C. L., Phang, J., Bowman, S. R., and Perez, E · 2023
Closest in time.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Li, K., Hopkins, A. K., Bau, D., Viégas, F., Pfister, H., and Wattenberg, M · 2023
Closest in time.
I’m a congressman who codes. A.I. freaks me out
Lieu, T · 2023
Closest in time.
Chatting about ChatGPT: how may AI and GPT impact academia and libraries?
Lund, B. D. and Wang, T · 2023
Closest in time.
Reinventing search with a new AI-powered Microsoft Bing and Edge, your copilot for the web
Mehdi, Y · 2023
Closest in time.
The new AI-powered Bing is threatening users. that’s no laughing matter
Perrigo, B · 2023
Closest in time.
A conversation with Bing’s chatbot left me deeply unsettled
Roose, K · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Closest in time.
Prompting GPT-3 to be reliable
Si, C., Gan, Z., Yang, Z., Wang, S., Wang, J., Boyd-Graber, J. L., and Wang, L · 2023
Closest in time.
Grounding the vector space of an octopus: Word meaning from raw text
Søgaard, A · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Large language models fail on trivial alterations to theory-of-mind tasks
Ullman, T · 2023
Closest in time.
A prompt pattern catalog to enhance prompt engineering with ChatGPT
White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and Schmidt, D. C · 2023
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., and Chi, E. H · 2023
Closest in time.