Fetching the paper…
Reading the bibliography…
Large language models (LLMs) excel at rapid generation of text and multimodal content, yet they falter on transaction-style planning that demands ACID-like guarantees and real-time disruption recovery.
On formally undecidable propositions of Principia Mathematica and related systems i
Kurt Gödel · 1967
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C. Daniel Gelatt, and Mario P. Vecchi · 1983
Earlier work this paper cites.
The Traveling Salesman Problem: A Guided Tour of Combinatorial Optimization
Eugene L. Lawler, Jan Karel Lenstra, Alexander H.G. Rinnooy Kan, and David B. Shmoys · 1985
Earlier work this paper cites.
The shifting bottleneck procedure for job shop scheduling
Joseph Adams, Egon Balas, and Daniel Zawack · 1988
Earlier work this paper cites.
Simulated Annealing: an Introduction
Emile HL Aarts and Peter JM Van Laarhoven · 1989
Earlier work this paper cites.
A generalized permutation approach to job shop scheduling with genetic algorithms
Christian Bierwirth · 1995
Earlier work this paper cites.
A fast taboo search algorithm for the job shop problem
Ewa Nowicki and Czesław Smutnicki · 1996
Earlier work this paper cites.
A computational study of shifting bottleneck procedures for shop scheduling problems
E. Demirkol, S. Mehta, and R. Uzsoy · 1998
Earlier work this paper cites.
Benchmarks for shop scheduling problems
Ebru Demirkol, Sanjay Mehta, and Reha Uzsoy · 1998
Earlier work this paper cites.
An Introduction to MultiAgent Systems
Michael Wooldridge · 2009
Earlier work this paper cites.
Scheduling: Theory, Algorithms, and Systems
M.L. Pinedo · 2012
Earlier work this paper cites.
Reactive Scheduling in a Job Shop Where Jobs Arrive Over Time
Li Nie, Liang Gao, Peigen Li, and Xinyu Shao · 2013
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, and more · 2017
Earlier work this paper cites.
Boosting binary optimization via binary classification: A case study of job shop scheduling
Oleg V. Shylo and Hesam Shams · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
Genetic Programming with Multi-tree Representation for Dynamic Flexible Job Shop Scheduling
Fangfang Zhang, Yi Mei, and Mengjie Zhang · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes Du, and Yejin Choi · 2020
Earlier work this paper cites.
Learning to dispatch for job shop scheduling via deep reinforcement learning
Cong Zhang, Wen Song, Zhiguang Cao, Jie Zhang, Puay Siew Tan, and Chi Xu · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, and Dieter Fox · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models, 2022
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, and more · 2022
Earlier work this paper cites.
Pathways to autonomous machine intelligence
Yann LeCun · 2022
Earlier work this paper cites.
Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango
Aman Madaan and Amir Yazdanbakhsh · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Jan Leike, Owain Evans, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou · 2022
Earlier work this paper cites.
A survey of job shop scheduling problem: The types and models
Hegen Xiong, Shuangyuan Shi, Danni Ren, and Jinjin Hu · 2022
Earlier work this paper cites.
CRIT: Prompting Large Language Models With the Socratic Method
Edward Y. Chang · 2023
Earlier work this paper cites.
Examining GPT-4’s Capabilities and Enhancement with SocraSynth
Edward Y Chang · 2023
Cited alongside, same era.
A deep reinforcement learning framework based on an attention mechanism and disjunctive graph embedding for the job-shop scheduling problem
Ruiqi Chen, Wenxin Li, and Hongbing Yang · 2023
Cited alongside, same era.
Flexible job shop scheduling problem under Industry 5.0: A survey on human reintegration, environmental consideration and resilience improvement
Candice Destouet, Houda Tlahig, Belgacem Bettayeb, and Bélahcène Mazari · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
S. Hong, X. Zheng, J. Chen, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, et al · 2023
Found in the Middle: Calibrating Positional Attention Bias Improves Long Context Utilization
Cheng-Yu Hsieh, Yung-Sung Chuang, Chun-Liang Li, Zifeng Wang, Long T. Le, and more · 2024
Later among the works it cites.
Automated Design of Agentic Systems
Shengran Hu, Cong Lu, and Jeff Clune · 2024
Later among the works it cites.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2024
Later among the works it cites.
Jin Huang, Xinyu Li, Liang Gao, Qihao Liu, and Yue Teng · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Cited alongside, same era.
Dissecting Chain-of-Thought: Compositionality through In-Context Filtering and Learning
Yingcong Li, Kartik Sreenivasan, Angeliki Giannou, Dimitris Papailiopoulos, and Samet Oymak · 2023
Cited alongside, same era.
Context-faithful prompting for large language models
Seungjae Park, Jamin Sen, Masato Inoue, Zhiqi Jiang, Rohan Mathur, John Z Zhang, Jieyu Zhang, Nathaniel Fernandez, Kuan Zhou, James Landon, et al · 2023
Cited alongside, same era.
Why think step by step? reasoning emerges from the locality of experience
Bart Prystawski, M. Y. Li, and Noah Goodman · 2023
Cited alongside, same era.
Chatdev: Communicative agents for software development
C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, J. Xu, D. Li, Z. Liu, and M. Sun · 2023
Cited alongside, same era.
Planbench: An extensible benchmark for evaluating large language models on planning and reasoning about change
Karthik Valmeekam, Alberto Olmo, Matthew Marquez, Sarath Sreedharan, and Subbarao Kambhampati · 2023
Cited alongside, same era.
Efficient large language models: A survey
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, et al · 2023
Cited alongside, same era.
Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, and more · 2024
Later among the works it cites.
Evaluating Creative Short Story Generation in Humans and Large Language Models
Mete Ismayilzada, Claire Stevenson, and Lonneke van der Plas · 2024
Later among the works it cites.
Self-[in]correct: Llms struggle with refining self-generated responses
Dongwei Jiang, Jingyu Zhang, Orion Weller, Nathaniel Weir, Benjamin Van Durme, and Daniel Khashabi · 2024
Later among the works it cites.
Langgraph: Building structured applications with llms
LangChain AI · 2024
Later among the works it cites.
Dynamic job-shop scheduling via graph attention networks and deep reinforcement learning
Chien-Liang Liu, Chun-Jan Tseng, and Po-Hao Weng · 2024
Later among the works it cites.
Lost in the middle: How language models use long contexts
Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2024
Later among the works it cites.
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of LLMs
Nisarg Patel, Mohith Kulkarni, Mihir Parmar, and Aashna Budhiraja · 2024
Later among the works it cites.
Hyperagent: Generalist software engineering agents to solve coding tasks at scale
H. N. Phan, T. N. Nguyen, P. X. Nguyen, and N. D. Bui · 2024
Later among the works it cites.
Chain of thoughtlessness? an analysis of cot in planning
Kaya Stechly, Karthik Valmeekam, and Subbarao Kambhampati · 2024
Later among the works it cites.
Appworld: A controllable world of apps and people for benchmarking interactive coding agents
H. Trivedi, T. Khot, M. Hartmann, R. Manku, V. Dong, E. Li, S. Gupta, A. Sabharwal, and N. Balasubramanian · 2024
Later among the works it cites.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, and Chi Wang · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks, 2024
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2024
Later among the works it cites.
Large language models can learn temporal reasoning
Siheng Xiong, Ali Payani, Ramana Kompella, and Faramarz Fekri · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shinnosuke Yao, Dong Yu, Jianfeng Zhao, Izhak Shafran, Thomas Griffiths, Yanjun Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Aflow: Automating agentic workflow generation, 2024
Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, Fengwei Teng, Xionghui Chen, Jiaqi Chen, Mingchen Zhuge, Xin Cheng, Sirui Hong, Jinlin Wang, Bingnan Zheng, Bang Liu, Yuyu Luo, and Chenglin Wu · 2024
Later among the works it cites.
Starjob: Dataset for LLM-Driven Job Shop Scheduling
Henrik Abgaryan, Tristan Cazenave, and Ararat Harutyunyan · 2025
Closest in time.
Why do multi-agent llm systems fail?, 2025
Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, and Ion Stoica · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI, Daya Guo, Dejian Yang, and more · 2025
Closest in time.
Gemini 2.5
Koray Kavukcuoglu · 2025
Closest in time.
Large language model agent: A survey on methodology, applications and challenges, 03 2025
Junyu Luo, Weizhi Zhang, Ye Yuan, Yusheng Zhao, Junwei Yang, Yiyang Gu, and more · 2025
Closest in time.
A Survey on Large Language Models with some Insights on their Capabilities and Limitations
Andrea Matarazzo and Riccardo Torlone · 2025
Closest in time.
Large Language Models: A Survey
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao · 2025
Closest in time.
Hello GPT-4o, 2024
OpenAI · 2025
Closest in time.