Fetching the paper…
Reading the bibliography…
Efficient exploration is essential for intelligent systems interacting with their environment, but existing language models often fall short in scenarios that require strategic information gathering.
Policy gradient search: Online planning and expert iteration without search trees, 2019
Anthony, T., Nishihara, R., Moritz, P., Salimans, T., and Schulman, J · 1904
Earlier work this paper cites.
Introduction to multi-armed bandits, 2024
Slivkins, A · 1904
Earlier work this paper cites.
Interactive fiction games: A colossal adventure, 2020a
Hausknecht, M., Ammanabrolu, P., Côté, M.-A., and Yuan, X · 1909
Earlier work this paper cites.
Interactive fiction games: A colossal adventure, 2020b
Hausknecht, M., Ammanabrolu, P., Côté, M.-A., and Yuan, X · 1909
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Reinforcement today
Skinner, B. F · 1958
Earlier work this paper cites.
Statistical mechanics of cellular automata
Wolfram, S · 1983
Earlier work this paper cites.
Curious model-building control systems
Schmidhuber, J · 1991
Earlier work this paper cites.
Instruction-following evaluation for large language models, 2023
Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., and Hou, L · 1991
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Using upper confidence bounds for online learning
Auer, P · 2000
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Automatic curriculum learning for deep rl: A short survey
Portelas, R., Colas, C., Weng, L., Hofmann, K., and Oudeyer, P.-Y · 2003
Earlier work this paper cites.
Universality in elementary cellular automata
Cook, M. et al · 2004
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
Schmidhuber, J · 2007
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Best Arm Identification in Multi-Armed Bandits
Audibert, J.-Y. and Bubeck, S · 2010
Earlier work this paper cites.
Biometry : the principles and practice of statistics in biological research / robert r. sokal and f. james rohlf, 04 2013
Sokal, R. and Rohlf, F · 2013
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Earlier work this paper cites.
Asking and evaluating natural language questions
Rothe, A., Lake, B., and Gureckis, T · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search, 2017
Anthony, T., Tian, Z., and Barber, D · 2017
Earlier work this paper cites.
Ucb exploration via q-ensembles
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Florensa, C., Held, D., Wulfmeier, M., Zhang, M., and Abbeel, P · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts, 2017
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Question asking as program generation, 2017
Rothe, A., Lake, B. M., and Gureckis, T. M · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Earlier work this paper cites.
Do people ask good questions?
Rothe, A., Lake, B. M., and Gureckis, T. M · 2018
Earlier work this paper cites.
Textworld: A learning environment for text-based games, 2019
Côté, M.-A., Ákos Kádár, Yuan, X., Kybartas, B., Barnes, T., Fine, E., Moore, J., Tao, R. Y., Hausknecht, M., Asri, L. E., Adada, M., Tay, W., and Trischler, A · 2019
Cited alongside, same era.
Curriculum-guided hindsight experience replay
Fang, M., Zhou, T., Du, Y., Han, L., and Zhang, Z · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Self-supervised exploration via disagreement
Pathak, D., Gandhi, D., and Gupta, A · 2019
Cited alongside, same era.
Asking goal-oriented questions and learning from answers
Rothe, A., Lake, B. M., and Gureckis, T. M · 2019
Cited alongside, same era.
Genqa: Generating millions of instructions from a handful of prompts, 2024
Chen, J., Qadri, R., Wen, Y., Jain, N., Kirchenbauer, J., Zhou, T., and Goldstein, T · 2024
Later among the works it cites.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Later among the works it cites.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Later among the works it cites.
Kto: Model alignment as prospect theoretic optimization, 2024
Ethayarajh, K., Xu, W., Muennighoff, N., Jurafsky, D., and Kiela, D · 2024
Later among the works it cites.
A llama sunk my battleship! asking rational questions with LLMs via bayesian inference
Grand, G., Pepe, V., Andreas, J., and Tenenbaum, J. B · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sharma, A., Gu, S., Levine, S., Kumar, V., and Hausman, K · 2019
Cited alongside, same era.
Wang, R., Lehman, J., Clune, J., and Stanley, K. O · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Emergent complexity and zero-shot transfer via unsupervised environment design
Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., and Levine, S · 2020
Cited alongside, same era.
Bridges-2: A platform for rapidly-evolving and data intensive research
Brown, S. T., Buitrago, P., Hanna, E., Sanielevici, S., Scibek, R., and Nystrom, N. A · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset, 2021
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
A systematic survey of text worlds as embodied natural language environments, 2021
Jansen, P. A · 2021
Cited alongside, same era.
Later among the works it cites.
Teaching large language models to reason with reinforcement learning, 2024
Havrilla, A., Du, Y., Raparthy, S. C., Nalmpantis, C., Dwivedi-Yu, J., Zhuravinskyi, M., Hambro, E., Sukhbaatar, S., and Raileanu, R · 2024
Later among the works it cites.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Later among the works it cites.
Jansen, P., Côté, M.-A., Khot, T., Bransom, E., Mishra, B. D., Majumder, B. P., Tafjord, O., and Clark, P · 2024
Later among the works it cites.
Can large language models explore in-context?, 2024
Krishnamurthy, A., Harris, K., Foster, D. J., Zhang, C., and Slivkins, A · 2024
Later among the works it cites.
Mt-eval: A multi-turn capabilities evaluation benchmark for large language models
Kwan, W.-C., Zeng, X., Jiang, Y., Wang, Y., Li, L., Shang, L., Jiang, X., Liu, Q., and Wong, K.-F · 2024
Later among the works it cites.
Supervised pretraining can learn in-context reinforcement learning
Lee, J., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E · 2024
Later among the works it cites.
Transformers as decision makers: Provable in-context reinforcement learning via supervised pretraining
Lin, L., Bai, Y., and Mei, S · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
MetaAI, Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., et al · 2024
Later among the works it cites.
BAGEL: Bootstrapping agents by guiding exploration with language
Murty, S., Manning, C. D., Shaw, P., Joshi, M., and Lee, K · 2024
Later among the works it cites.
Aviary: training language agents on challenging scientific tasks, 2024
Narayanan, S., Braza, J. D., Griffiths, R.-R., Ponnapati, M., Bou, A., Laurent, J., Kabeli, O., Wellawatte, G., Cox, S., Rodriques, S. G., and White, A. D · 2024
Later among the works it cites.
Turning up the heat: Min-p sampling for creative and coherent llm outputs
Nguyen, M., Baker, A., Neo, C., Roush, A., Kirsch, A., and Shwartz-Ziv, R · 2024
Later among the works it cites.
Evolve: Evaluating and optimizing llms for exploration, 2024
Nie, A., Su, Y., Chang, B., Lee, J. N., Chi, E. H., Le, Q. V., and Chen, M · 2024
Later among the works it cites.
Smaug: Fixing failure modes of preference optimisation with dpo-positive, 2024
Pal, A., Karkhanis, D., Dooley, S., Roberts, M., Naidu, S., and White, C · 2024
Later among the works it cites.
Iterative reasoning preference optimization, 2024
Pang, R. Y., Yuan, W., Cho, K., He, H., Sukhbaatar, S., and Weston, J · 2024
Later among the works it cites.
Unintentional unalignment: Likelihood displacement in direct preference optimization, 2024
Razin, N., Malladi, S., Bhaskar, A., Chen, D., Arora, S., and Hanin, B · 2024
Later among the works it cites.
Preference fine-tuning of llms should leverage suboptimal, on-policy data, 2024
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., and Kumar, A · 2024
Later among the works it cites.
Is dpo superior to ppo for llm alignment? a comprehensive study, 2024
Xu, S., Fu, W., Gao, J., Ye, W., Liu, W., Mei, Z., Wang, G., Yu, C., and Wu, Y · 2024
Later among the works it cites.
React meets actre: Autonomous annotation of agent trajectories for contrastive self-training
Yang, Z., Li, P., Yan, M., Zhang, J., Huang, F., and Liu, Y · 2024
Later among the works it cites.
Wildchat: 1m chatgpt interaction logs in the wild, 2024
Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., and Deng, Y · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., et al · 2025
Closest in time.
Learning to reason at the frontier of learnability, 2025
Foster, T. and Foerster, J · 2025
Closest in time.
Gemma 3 technical report, 2025
Gemma-Team, Kamath, A., Ferret, J., Pathak, S., Vieillard, N., Merhej, R., Perrin, S., Matejovicova, T., Ramé, A., Rivière, M., Rouillard, L., Mesnard, T., et al · 2025
Closest in time.
Shoot first, ask questions later? building rational agents that explore and act like people, 2025
Grand, G., Pepe, V., Andreas, J., and Tenenbaum, J. B · 2025
Closest in time.
Should you use your large language model to explore or exploit?, 2025
Harris, K. and Slivkins, A · 2025
Closest in time.
Aligning llms to ask good questions a case study in clinical reasoning, 2025
Li, S. S., Mun, J., Brahman, F., Ilgen, J. S., Tsvetkov, Y., and Sap, M · 2025
Closest in time.
Llms are in-context bandit reinforcement learners, 2025
Monea, G., Bosselut, A., Brantley, K., and Artzi, Y · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., et al · 2025
Closest in time.