Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022 · 2022
Later among the works it cites.
Palm 2 technical report
Original
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023 · 2023
Later among the works it cites.
Scalable AI safety via doubly-efficient debate
Original
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras. 2023 · 2023
Later among the works it cites.
Be selfish, but wisely: Investigating the impact of agent personality in mixed-motive human-agent interactions
Original
Kushal Chawla, Ian Wu, Yu Rong, Gale M Lucas, and Jonathan Gratch. 2023 · 2023
Later among the works it cites.
Promptbreeder: Self-referential self-improvement via prompt evolution
Original
Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2023 · 2023
Later among the works it cites.
Pragmatics in language grounding: Phenomena, tasks, and modeling approaches
Daniel Fried, Nicholas Tomlin, Jennifer Hu, Roma Patel, and Aida Nematzadeh. 2023 · 2023
Later among the works it cites.
Strategic reasoning with language models
Original
Kanishk Gandhi, Dorsa Sadigh, and Noah D Goodman. 2023 · 2023
Later among the works it cites.
Mistral 7b
Original
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Later among the works it cites.
Reward design with language models
Original
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh. 2023 · 2023
Later among the works it cites.
Ai deception: A survey of examples, risks, and potential solutions
Original
Peter S Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2023 · 2023
Later among the works it cites.
Emergence and collapse of reciprocity in semiautomatic driving coordination experiments with humans
Hirokazu Shirado, Shunichi Kasahara, and Nicholas A Christakis. 2023 · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Original
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023 · 2023
Later among the works it cites.
Large language models as optimizers
Original
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen. 2023 · 2023
Later among the works it cites.
The illusion of artificial inclusion
William Agnew, A Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R McKee. 2024 · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Closest in time.