Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated significant capabilities in natural language processing and reasoning, yet their effectiveness in autonomous planning has been under debate.
Dynamic programming
Richard Bellman · 1966
Earlier work this paper cites.
A formal basis for the heuristic determination of minimum cost paths
Peter E Hart, Nils J Nilsson, and Bertram Raphael · 1968
Earlier work this paper cites.
Decisions with multiple objectives: Preferences and value tradeoffs
Ralph L Keeney · 1993
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Two Volume Set
Dimitri P. Bertsekas · 1995
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bandit based monte-carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Combining online and offline knowledge in uct
Sylvain Gelly and David Silver · 2007
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Attention is all you need
A Vaswani · 2017
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
The child as hacker: building more human-like models of learning
Joshua Stewart Rule · 2020
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Earlier work this paper cites.
System 3 Thinking: How to Choose Wisely when Facing Doubt, Dilemma, Or Disruption
P.J. Webb · 2021
Earlier work this paper cites.
Acre: Abstract causal reasoning beyond covariation
Chi Zhang, Baoxiong Jia, Mark Edmonds, Song-Chun Zhu, and Yixin Zhu · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
Discovering exfiltration paths using reinforcement learning with attack graphs
Tyler Cody, Abdul Rahman, Christopher Redino, Lanxiao Huang, Ryan Clark, Akshay Kakkar, Deepak Kushwaha, Paul Park, Peter Beling, and Edward Bowen · 2022
Earlier work this paper cites.
Magnetic field mapping of inaccessible regions using physics-informed neural networks
Umit H Coskun, Bilgehan Sel, and Brad Plaster · 2022
Earlier work this paper cites.
Compositional Semantic Parsing with Large Language Models
Andrew Drozdov, Nathanael Schärli, Ekin Akyürek, Nathan Scales, Xinying Song, Xinyun Chen, Olivier Bousquet, and Denny Zhou · 2022
Earlier work this paper cites.
Towards reasoning in large language models: A survey
Jie Huang and Kevin Chen-Chuan Chang · 2022
Earlier work this paper cites.
Exposing surveillance detection routes via reinforcement learning, attack graphs, and cyber terrain
Lanxiao Huang, Tyler Cody, Christopher Redino, Abdul Rahman, Akshay Kakkar, Deepak Kushwaha, Cheng Wang, Ryan Clark, Daniel Radke, Peter Beling, et al · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2022
Cited alongside, same era.
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them, October 2022
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V. Le, Ed H. Chi, Denny Zhou, and Jason Wei · 2022
Cited alongside, same era.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Cited alongside, same era.
Learning neural networks under input-output specifications
Zain ul Abdeen, He Yin, Vassilis Kekatos, and Ming Jin · 2022
Cited alongside, same era.
Self-Consistency Improves Chain of Thought Reasoning in Language Models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Certifiably robust neural ode with learning-based barrier function
Runing Yang, Ruoxi Jia, Xiangyu Zhang, and Ming Jin · 2023
Later among the works it cites.
Self-taught optimizer (stop): Recursively self-improving code generation
Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai · 2023
Later among the works it cites.
Does online gradient descent (and variants) still work with biased gradient and variance?
Ahmad Al-Tawaha and Ming Jin · 2024
Later among the works it cites.
Graph of thoughts: Solving elaborate problems with large language models
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al · 2024
Later among the works it cites.
Boosting of thoughts: Trial-and-error problem solving with large language models
Sijia Chen, Baochun Li, and Di Niu · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2022
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola · 2022
Cited alongside, same era.
Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi · 2022
Cited alongside, same era.
Decision-focused learning for inverse noncooperative games: Generalization bounds and convergence analysis
Ahmad Al-Tawaha, Harshal Kaushik, Bilgehan Sel, Ruoxi Jia, and Ming Jin · 2023
Cited alongside, same era.
A survey of chain of thought reasoning: Advances, frontiers and future
Zheng Chu, Jingchang Chen, Qianglong Chen, Weijiang Yu, Tao He, Haotian Wang, Weihua Peng, Ming Liu, Bing Qin, and Ting Liu · 2023
Cited alongside, same era.
Dynamic modeling and trajectory tracking of a quadcopter via linear and backstepping controller
Uygar Gunes, Artun Sel, Bilgehan Sel, and Cosku Kasnakoglu · 2023
Cited alongside, same era.
Balance reward and safety optimization for safe reinforcement learning: A perspective of gradient manipulation
Shangding Gu, Bilgehan Sel, Yuhao Ding, Lu Wang, Qingwei Lin, Ming Jin, and Alois Knoll · 2024
Later among the works it cites.
Can we trust the performance evaluation of uncertainty estimation methods in text summarization?
Jianfeng He, Runing Yang, Linlin Yu, Changbin Li, Ruoxi Jia, Feng Chen, Ming Jin, and Chang-Tien Lu · 2024
Later among the works it cites.
Democratizing energy management with llm-assisted optimization autoformalism
Ming Jin, Bilgehan Sel, Fnu Hardeep, and Wotao Yin · 2024
Later among the works it cites.
Llms can’t plan, but can help planning in llm-modulo frameworks
Subbarao Kambhampati, Karthik Valmeekam, Lin Guan, Kaya Stechly, Mudit Verma, Siddhant Bhambri, Lucas Saldyt, and Anil Murthy · 2024
Later among the works it cites.
When can llms actually correct their own mistakes? a critical survey of self-correction of llms
Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, and Rui Zhang · 2024
Later among the works it cites.
Optimization solution functions as deterministic policies for offline reinforcement learning
Vanshaj Khattar and Ming Jin · 2024
Later among the works it cites.
A cmdp-within-online framework for meta-safe reinforcement learning
Vanshaj Khattar, Yuhao Ding, Bilgehan Sel, Javad Lavaei, and Ming Jin · 2024
Later among the works it cites.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer · 2024
Later among the works it cites.
Trustworthiness and self-awareness in large language models: An exploration through the think-solve-verify framework
Zhendong Liu, Changhong Xia, Wei He, and Chongjun Wang · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Later among the works it cites.
Zero-day attack detection in digital substations using in-context learning
Faizan Manzoor, Vanshaj Khattar, Chen-Ching Liu, and Ming Jin · 2024
Later among the works it cites.
Refiner: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2024
Later among the works it cites.
On the self-verification limitations of large language models on reasoning and planning tasks
Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati · 2024
Later among the works it cites.
Enhancing distribution system resilience: A first-order meta-rl algorithm for critical load restoration
Zain ul Abdeen, Xiangyu Zhang, Waris Gill, and Ming Jin · 2024
Later among the works it cites.
Leveraging reinforcement learning in red teaming for advanced ransomware attack simulations
Cheng Wang, Christopher Redino, Ryan Clark, Abdul Rahman, Sal Aguinaga, Sathvik Murli, Dhruv Nandakumar, Roland Rao, Lanxiao Huang, Daniel Radke, et al · 2024
Later among the works it cites.
Self-evaluation guided beam search for reasoning
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, James Xu Zhao, Min-Yen Kan, Junxian He, and Michael Xie · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Safe and balanced: A framework for constrained multi-objective reinforcement learning
Shangding Gu, Bilgehan Sel, Yuhao Ding, Lu Wang, Qingwei Lin, Alois Knoll, and Ming Jin · 2025
Closest in time.