Fetching the paper…
Reading the bibliography…
The emergence of large language models (LLMs) has substantially influenced natural language processing, demonstrating exceptional results across various tasks.
Introspective multistrategy learning: Constructing a learning strategy under reasoning failure
M. T. Cox · 1996
Earlier work this paper cites.
Cognitive load theory and complex learning: Recent developments and future directions
J. J. Van Merrienboer and J. Sweller · 2005
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Earlier work this paper cites.
Textworld: A learning environment for text-based games
M.-A. Côté, A. Kádár, X. Yuan, B. Kybartas, T. Barnes, E. Fine, J. Moore, M. Hausknecht, L. El Asri, M. Adada, et al · 2019
Earlier work this paper cites.
The limits of machine intelligence: Despite progress in machine intelligence, artificial general intelligence is still a major challenge
H. Shevlin, K. Vold, M. Crosby, and M. Halina · 2019
Earlier work this paper cites.
First textworld problems, the competition: Using text-based games to advance capabilities of ai agents
A. Trischler, M.-A. Côté, and P. Lima · 2019
Earlier work this paper cites.
Comprehensible context-driven text game playing
X. Yin and J. May · 2019
Earlier work this paper cites.
Learning dynamic belief graphs to generalize on text-based games
A. Adhikari, X. Yuan, M.-A. Côté, M. Zelinka, M.-A. Rondeau, R. Laroche, P. Poupart, J. Tang, A. Trischler, and W. Hamilton · 2020
Earlier work this paper cites.
Graph constrained reinforcement learning for natural language action spaces
P. Ammanabrolu and M. Hausknecht · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
D. Hendrycks, M. Mazeika, A. Zou, S. Patel, C. Zhu, J. Navarro, D. Song, B. Li, and J. Steinhardt · 2021
Earlier work this paper cites.
Neuro-symbolic reinforcement learning with first-order logic
D. Kimura, M. Ono, S. Chaudhury, R. Kohita, A. Wachi, D. J. Agravante, M. Tatsubori, A. Munawar, and A. Gray · 2021
Earlier work this paper cites.
Pretrained transformers as universal computation engines
K. Lu, A. Grover, P. Abbeel, and I. Mordatch · 2021
Earlier work this paper cites.
Improving sample efficiency in model-free reinforcement learning from images
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
E. Brooks, L. Walls, R. L. Lewis, and S. Singh · 2022
Cited alongside, same era.
Safe learning in robotics: From learning-based control to safe reinforcement learning
L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Cited alongside, same era.
In-context reinforcement learning with algorithm distillation
Perceiving the world: Question-guided reinforcement learning for text-based games
Y. Xu, M. Fang, L. Chen, Y. Du, J. T. Zhou, and C. Zhang · 2022
Later among the works it cites.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung · 2023
Closest in time.
Reward design with language models
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh · 2023
Closest in time.
Structured state space models for in-context reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Laskin, L. Wang, J. Oh, E. Parisotto, S. Spencer, R. Steigerwald, D. Strouse, S. Hansen, A. Filos, E. Brooks, et al · 2022
Cited alongside, same era.
Pre-trained language models for interactive decision-making
S. Li, X. Puig, C. Paxton, Y. Du, C. Wang, L. Fan, T. Chen, D.-A. Huang, E. Akyürek, A. Anandkumar, A. Jacob, M. Igor, T. Antonio, and Z. Yuke · 2022
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2022
Cited alongside, same era.
Learning object-oriented dynamics for planning from text
G. Liu, A. Adhikari, A.-m. Farahmand, and P. Poupart · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Cited alongside, same era.
A survey of text games for reinforcement learning informed by natural language
P. Osborne, H. Nõmm, and A. Freitas · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Cited alongside, same era.
C. Lu, Y. Schroecker, A. Gu, E. Parisotto, J. Foerster, S. Singh, and F. Behbahani · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al · 2023
Closest in time.
Introducing llama: A foundational, 65-billion-parameter large language model
A. Meta · 2023
Closest in time.
OpenAI · 2023
Closest in time.
B. Peng, M. Galley, P. He, H. Cheng, Y. Xie, Y. Hu, Q. Huang, L. Liden, Z. Yu, W. Chen, et al · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
N. Shinn, B. Labash, and A. Gopinath · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Human-timescale adaptation in an open-ended task space
A. A. Team, J. Bauer, K. Baumli, S. Baveja, F. Behbahani, A. Bhoopchand, N. Bradley-Schmieg, M. Chang, N. Clay, A. Collister, et al · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
S. Vemprala, R. Bonatti, A. Bucker, and A. Kapoor · 2023
Closest in time.
Foundation models for decision making: Problems, methods, and opportunities
S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans · 2023
Closest in time.