Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown promise as intelligent agents in interactive decision-making tasks.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J · 1999
Earlier work this paper cites.
Convergence of reinforcement learning algorithms and acceleration of learning
Potapov, A. and Ali, M · 2003
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, E · 2008
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Koller, D. and Friedman, N · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, M · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D · 2010
Earlier work this paper cites.
Optimal control theory and the linear bellman equation
Kappen, H. J · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, H. J., Gómez, V., and Opper, M · 2012
Earlier work this paper cites.
Reinforcement learning and markov decision processes
Van Otterlo, M. and Wiering, M · 2012
Earlier work this paper cites.
Openml: networked science in machine learning
Vanschoren, J., Van Rijn, J. N., Bischl, B., and Torgo, L · 2014
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R., Pakman, A., and Tishby, N · 2015
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Earlier work this paper cites.
Understanding auc-roc curve
Narkhede, S · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2018
Cited alongside, same era.
Automated machine learning: methods, systems, challenges
Hutter, F., Kotthoff, L., and Vanschoren, J · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
K2: A foundation language model for geoscience knowledge understanding and utilization
Deng, C., Zhang, T., He, Z., Chen, Q., Shi, Y., Zhou, L., Fu, L., Zhang, W., Wang, X., Zhou, C., Lin, Z., and He, J · 2023
Later among the works it cites.
Alphazero-like tree-search can guide large language model decoding and training
Feng, X., Wan, Z., Wen, M., Wen, Y., Zhang, W., and Wang, J · 2023
Later among the works it cites.
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z · 2023
Later among the works it cites.
Large language models for automated data science: Introducing caafe for context-aware automated feature engineering
Hollmann, N., Müller, S., and Hutter, F · 2023
Later among the works it cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, A., Pinto, L., and Abbeel, P · 2020
Cited alongside, same era.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M · 2020
Cited alongside, same era.
Accounting for variance in machine learning benchmarks
Bouthillier, X., Delaunay, P., Bronzi, M., Trofimov, A., Nichyporuk, B., Szeto, J., Mohammadi Sepahvand, N., Raff, E., Madan, K., Voleti, V., et al · 2021
Cited alongside, same era.
Coderl: Mastering code generation through pretrained models and deep reinforcement learning
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., and Hoi, S. C. H · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2022
Cited alongside, same era.
Multi-agent reinforcement learning is a sequence modeling problem
Wen, M., Kuba, J., Lin, R., Zhang, W., Wen, Y., Wang, J., and Yang, Y · 2022
Cited alongside, same era.
Later among the works it cites.
Geogalactica: A scientific large language model in geoscience
Lin, Z., Deng, C., Zhou, L., Zhang, T., Xu, Y., Xu, Y., He, Z., Shi, Y., Dai, B., Song, Y., Zeng, B., Chen, Q., Shi, T., Huang, T., Xu, Y., Wang, S., Fu, L., Zhang, W., He, J., Ma, C., Zhu, Y., Wang, X., and Zhou, C · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2023
Later among the works it cites.
Automatic prompt optimization with” gradient descent” and beam search
Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M · 2023
Later among the works it cites.
Communicative agents for software development
Qian, C., Cong, X., Yang, C., Chen, W., Su, Y., Xu, J., Liu, Z., and Sun, M · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J., Ellenberg, J. S., Wang, P., Fawzi, O., et al · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Chen, Z., Zhou, K., Zhao, W. X., Wan, J., Zhang, F., Zhang, D., and Wen, J.-R · 2024
Closest in time.
Reinforcing language agents via policy optimization with action decomposition, 2024
Wen, M., Wan, Z., Zhang, W., Wang, J., and Wen, Y · 2024
Closest in time.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.