Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) harness extensive data from the Internet, storing a broad spectrum of prior knowledge.
Exploration–exploitation tradeoff using variance estimates in multi-armed bandits
Audibert, J.-Y · 1902
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K · 1967
Earlier work this paper cites.
The harpy speech recognition system: performance with large vocabularies
Lowerre, B · 1976
Earlier work this paper cites.
Scout: A simple game-searching algorithm with proven optimal properties
Pearl, J · 1980
Earlier work this paper cites.
The history heuristic and alpha-beta search enhancements in practice
Schaeffer, J · 1989
Earlier work this paper cites.
Sample mean based index policies by o (log n) regret for the multi-armed bandit problem
Agrawal, R · 1995
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P · 2002
Earlier work this paper cites.
Transposition table driven work scheduling in distributed game-tree search
Kishimoto, A · 2002
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Coulom, R · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Kocsis, L · 2006
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M · 2010
Earlier work this paper cites.
Beam monte-carlo tree search
Baier, H · 2012
Cited alongside, same era.
Dynamic programming and optimal control: Volume I
Bertsekas, D · 2012
Cited alongside, same era.
Enhancements for monte-carlo tree search in ms pac-man
Pepels, T · 2012
Cited alongside, same era.
Portfolio greedy search and simulation for large-scale combat in starcraft
Churchill, D · 2013
Cited alongside, same era.
Monte-carlo tree search in total war: Rome ii’s campaign ai
Champandard, A. J · 2014
Cited alongside, same era.
Script-and cluster-based uct for starcraft
Justesen, N · 2014
Cited alongside, same era.
Heuristic move pruning in monte carlo tree search for the strategic card game lords of war
Towards reasoning in large language models: A survey
Huang, J · 2022
Later among the works it cites.
Achiam, J · 2023
Later among the works it cites.
Playing repeated games with large language models
Akata, E · 2023
Later among the works it cites.
The reversal curse: Llms trained on” a is b” fail to learn” b is a”
Berglund, L · 2023
Later among the works it cites.
Reasoning with language model is planning with world model
Hao, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sephton, N · 2014
Cited alongside, same era.
Enhancements for real-time monte-carlo tree search in general video game playing
Soemers, D. J · 2016
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D · 2017
Cited alongside, same era.
A hybrid search agent in pommerman
Zhou, H · 2018
Cited alongside, same era.
Non-asymptotic analysis of monte carlo tree search
Shah, D · 2020
Cited alongside, same era.
Chessgpt: Bridging policy learning and language modeling
Feng, X
Cited in the paper.
Huang, L · 2023
Later among the works it cites.
Liu, Z · 2023
Later among the works it cites.
Alympics: Language agents meet game theory
Mao, S · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S · 2023
Later among the works it cites.
Siren’s song in the ai ocean: A survey on hallucination in large language models
Zhang, Y · 2023
Later among the works it cites.
Grandmaster-level chess without search
Ruoss, A · 2024
Closest in time.