Fetching the paper…
Reading the bibliography…
A rapidly growing number of applications rely on a small set of closed-source language models (LMs).
Chimpanzees: self-recognition
Gordon G Gallup Jr. 1970 · 1970
Earlier work this paper cites.
Verification
Rich Sutton. 2001 · 2001
Earlier work this paper cites.
Mike or me? self-recognition in a split-brain patient
David J Turk, Todd F Heatherton, William M Kelley, Margaret G Funnell, Michael S Gazzaniga, and C Neil Macrae. 2002 · 2002
Earlier work this paper cites.
Neural activity associated with self-reflection
Uwe Herwig, Tina Kaffenberger, Caroline Schell, Lutz Jäncke, and Annette B Brühl. 2012 · 2012
Earlier work this paper cites.
The significance of meaning: why do over 90% of behavioral neuroscience results fail to translate to humans, and what can we do to fix it?
Joseph P Garner. 2014 · 2014
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Unsupervised pretraining for sequence to sequence learning
Prajit Ramachandran, Peter J Liu, and Quoc V Le. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun. 2019 · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Earlier work this paper cites.
The continued need for animals to advance brain research
Judith R Homberg, Roger AH Adan, Natalia Alenina, Antonis Asiminas, Michael Bader, Tom Beckers, Denovan P Begg, Arjan Blokland, Marilise E Burger, Gertjan van Dijk, et al. 2021 · 2021
Earlier work this paper cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Language models as agent models
Jacob Andreas. 2022 · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 · 2022
Earlier work this paper cites.
The ethical need for watermarks in machine-generated language
Alexei Grinbaum and Laurynas Adomaitis. 2022 · 2022
Cited alongside, same era.
When should we prefer offline reinforcement learning over behavioral cloning?
Aviral Kumar, Joey Hong, Anikait Singh, and Sergey Levine. 2022 · 2022
Cited alongside, same era.
Mechanistic interpretability, variables, and the importance of interpretable bases
Chris Olah. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022 · 2022
Cited alongside, same era.
Claude 3 family
Anthropic. 2024 · 2024
Closest in time.
Evaluating language model agency through negotiations
Tim R. Davidson, Veniamin Veselovsky, Martin Josifoski, Maxime Peyrard, Antoine Bosselut, Michal Kosinski, and Robert West. 2024 · 2024
Closest in time.
Gemini: A family of highly capable multimodal models
Gemini. 2024 · 2024
Closest in time.
Watermark stealing in large language models
Nikola Jovanović, Robin Staab, and Martin Vechev. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2024 · 2024
Closest in time.
Llama 3
Meta. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mitigating label biases for in-context learning
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023 · 2023
Cited alongside, same era.
Turing mirror: Evaluating the ability of llms to recognize llm-generated text
Jason Hoelscher-Obermaier, Matthew J. Lutz, Quentin Feuillade-Montixi, and Sambita Modak. 2023 · 2023
Cited alongside, same era.
Large language models can self-improve
Jiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han. 2023 · 2023
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Cited alongside, same era.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023 · 2023
Cited alongside, same era.
Llms as narcissistic evaluators: When ego inflates evaluation scores
Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2023 · 2023
Cited alongside, same era.
Gpt-4 technical report
OpenAI, Josh Achiam, Steven Adler, and et al Sandhini Agarwal. 2023 · 2023
Cited alongside, same era.
The cost of training AI could soon become too much to bear
David Meyer. 2024 · 2024
Closest in time.
Language model inversion
John Xavier Morris, Wenting Zhao, Justin T Chiu, Vitaly Shmatikov, and Alexander M Rush. 2024 · 2024
Closest in time.
Introducing the GPT Store
OpenAI. 2024 · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng. 2024 · 2024
Closest in time.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. 2024 · 2024
Closest in time.
Beyond performance: Quantifying and mitigating label bias in llms
Yuval Reif and Roy Schwartz. 2024 · 2024
Closest in time.
Beyond memorization: Violating privacy via inference with large language models
Robin Staab, Mark Vero, Mislav Balunović, and Martin Vechev. 2024 · 2024
Closest in time.
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024 · 2024
Closest in time.
Provable robust watermarking for ai-generated text
Xuandong Zhao, Prabhanjan Ananth, Lei Li, and Yu-Xiang Wang. 2024 · 2024
Closest in time.
Large language models are not robust multiple choice selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2024 · 2024
Closest in time.