Fetching the paper…
Reading the bibliography…
Achieving human-level intelligence requires refining the transition from the fast, intuitive System 1 to the slower, more deliberate System 2 reasoning.
1909
Earlier work this paper cites.
C. I. Lewis, C. H. Langford, and P. Lamprecht, Symbolic logic . Dover publications New York, 1959, vol. 170
1959
Earlier work this paper cites.
M. Minsky et al. , “A framework for representing knowledge,” 1974
1974
Earlier work this paper cites.
J. McCarthy, “History of LISP,” in History of programming languages , 1978, pp. 173–185
1978
Earlier work this paper cites.
J. S. B. Evans, “Heuristic and analytic processes in reasoning,” British Journal of Psychology , vol. 75, no. 4, pp. 451–468, 1984
1984
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” nature , vol. 323, no. 6088, pp. 533–536, 1986
1986
Earlier work this paper cites.
R. G. Jeroslow, “Computation-oriented reductions of predicate to propositional logic,” Decision Support Systems , vol. 4, no. 2, pp. 183–197, 1988
1988
Earlier work this paper cites.
A. Colmerauer, “An introduction to Prolog III,” Communications of the ACM , vol. 33, no. 7, pp. 69–90, 1990
1990
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning , vol. 8, pp. 279–292, 1992
1992
Earlier work this paper cites.
Y. LeCun, Y. Bengio et al. , “Convolutional networks for images, speech, and time series,” The handbook of brain theory and neural networks , vol. 3361, no. 10, p. 1995, 1995
1995
Earlier work this paper cites.
S. Hochreiter, “Long Short-term Memory,” Neural Computation MIT-Press , 1997
1997
Earlier work this paper cites.
K. R. Apt et al. , From logic programming to Prolog . Prentice Hall London, 1997, vol. 362
1997
Earlier work this paper cites.
R. S. Sutton, A. G. Barto et al. , Reinforcement learning: An introduction . MIT press Cambridge, 1998, vol. 1, no. 1
1998
Earlier work this paper cites.
M. P. Singh, A. S. Rao, and M. P. Georgeff, Formal methods in DAI: Logic-based representation and reasoning . MIT Press Cambridge, 1999, vol. 8
1999
Earlier work this paper cites.
A. Y. Ng, D. Harada, and S. Russell, “Policy invariance under reward transformations: Theory and application to reward shaping,” in Icml , vol. 99. Citeseer, 1999, pp. 278–287
1999
Earlier work this paper cites.
L. Bachmair and H. Ganzinger, “Resolution Theorem Proving.” Handbook of automated reasoning , vol. 1, no. 02, 2001
2001
Earlier work this paper cites.
D. Kahneman, “Maps of bounded rationality: Psychology for behavioral economics,” American economic review , vol. 93, no. 5, pp. 1449–1475, 2003
2003
Earlier work this paper cites.
W. F. Clocksin and C. S. Mellish, Programming in PROLOG . Springer Science & Business Media, 2003
2003
Earlier work this paper cites.
G. E. Hinton, S. Osindero, and Y.-W. Teh, “A fast learning algorithm for deep belief nets,” Neural computation , vol. 18, no. 7, pp. 1527–1554, 2006
2006
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
S. Gelly and D. Silver, “Monte-Carlo tree search and rapid action value estimation in computer Go,” Artificial Intelligence , vol. 175, no. 11, pp. 1856–1875, 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” Advances in neural information processing systems , vol. 25, 2012
2012
Earlier work this paper cites.
R. Carnap, Introduction to symbolic logic and its applications . Courier Corporation, 2012
2012
Earlier work this paper cites.
C. B. Browne, E. Powley, D. Whitehouse, S. M. Lucas, P. I. Cowling, P. Rohlfshagen, S. Tavener, D. Perez, S. Samothrakis, and S. Colton, “A survey of monte carlo tree search methods,” IEEE Transactions on Computational Intelligence and AI in games , vol. 4, no. 1, pp. 1–43, 2012
2012
Earlier work this paper cites.
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, October 25-29, 2014, Doha, Qatar, A meeting of SIGDAT, a Special Interest Group of the ACL , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of Go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger, “Densely connected convolutional networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 4700–4708
2017
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton et al. , “Mastering the game of go without human knowledge,” nature , vol. 550, no. 7676, pp. 354–359, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
R. R. Torrado, P. Bontrager, J. Togelius, J. Liu, and D. Perez-Liebana, “Deep reinforcement learning for general video game ai,” in 2018 IEEE Conference on Computational Intelligence and Games (CIG) . IEEE, 2018, pp. 1–8
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev et al. , “Grandmaster level in StarCraft II using multi-agent reinforcement learning,” nature , vol. 575, no. 7782, pp. 350–354, 2019
2019
Earlier work this paper cites.
F. Chollet, “On the measure of intelligence,” arXiv preprint arXiv:1911.01547 , 2019
2019
Earlier work this paper cites.
T. Nandy, M. Y. I. B. Idris, R. M. Noor, L. M. Kiah, L. S. Lun, N. B. A. Juma’at, I. Ahmedy, N. A. Ghani, and S. Bhattacharyya, “Review on security of internet of things authentication mechanism,” IEEE Access , vol. 7, pp. 151 054–151 089, 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learning . Pmlr, 2021, pp. 8821–8831
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
B. Lester, R. Al-Rfou, and N. Constant, “The Power of Scale for Parameter-Efficient Prompt Tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 3045–3059
2021
Earlier work this paper cites.
D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits, “What disease does this patient have? a large-scale open domain question answering dataset from medical exams,” Applied Sciences , vol. 11, no. 14, p. 6421, 2021
2021
Earlier work this paper cites.
W. Hua and Y. Zhang, “System 1+ system 2= better world: Neural-symbolic chain of logic reasoning,” in Findings of the Association for Computational Linguistics: EMNLP 2022 , 2022, pp. 601–612
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Advances in neural information processing systems , vol. 35, pp. 22 199–22 213, 2022
2022
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International conference on machine learning . PMLR, 2022, pp. 5547–5569
2022
Earlier work this paper cites.
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer, “Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , 2022, pp. 11 048–11 064
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman, “STaR: Bootstrapping Reasoning With Reasoning,” in Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022 , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago et al. , “Competition-level code generation with alphacode,” Science , vol. 378, no. 6624, pp. 1092–1097, 2022
2022
Earlier work this paper cites.
S. Yao, H. Chen, J. Yang, and K. Narasimhan, “Webshop: Towards scalable real-world web interaction with grounded language agents,” Advances in Neural Information Processing Systems , vol. 35, pp. 20 744–20 757, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan, “Learn to explain: Multimodal reasoning via thought chains for science question answering,” Advances in Neural Information Processing Systems , vol. 35, pp. 2507–2521, 2022
2022
Earlier work this paper cites.
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-Consistency Improves Chain of Thought Reasoning in Language Models,” in The Eleventh International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. V. Le et al. , “Least-to-Most Prompting Enables Complex Reasoning in Large Language Models,” in The Eleventh International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
J. Huang and K. C.-C. Chang, “Towards Reasoning in Large Language Models: A Survey,” in Findings of the Association for Computational Linguistics: ACL 2023 , 2023, pp. 1049–1065
2023
Earlier work this paper cites.
S. Qiao, Y. Ou, N. Zhang, X. Chen, Y. Yao, S. Deng, C. Tan, F. Huang, and H. Chen, “Reasoning with Language Model Prompting: A Survey,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 5368–5393
2023
Earlier work this paper cites.
B. Wang, S. Min, X. Deng, J. Shen, Y. Wu, L. Zettlemoyer, and H. Sun, “Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 2717–2739
2023
Earlier work this paper cites.
O. Shaikh, H. Zhang, W. Held, M. Bernstein, and D. Yang, “On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 4454–4470
2023
Earlier work this paper cites.
Z. Zhang, A. Zhang, M. Li, and A. Smola, “Automatic Chain of Thought Prompting in Large Language Models,” in The Eleventh International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
S. Hao, Y. Gu, H. Ma, J. Hong, Z. Wang, D. Wang, and Z. Hu, “Reasoning with Language Model is Planning with World Model,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 8154–8173
2023
Earlier work this paper cites.
Y. Zhang, “Meta prompting for agi systems,” arXiv preprint arXiv:2311.11482 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual Instruction Tuning,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Zhu, J. Wang, L. Zhang, Y. Zhang, Y. Huang, R. Gan, J. Zhang, and Y. Yang, “Solving Math Word Problems via Cooperative Reasoning induced Language Models,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 4471–4485
2023
Earlier work this paper cites.
P. Lu, L. Qiu, K.-W. Chang, Y. N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan, “Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning,” in The Eleventh International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. You, R. Sun, Z. Wang, L. Chen, G. Wang, H. Ayyubi, K.-W. Chang, and S.-F. Chang, “IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023, pp. 11 289–11 303
2023
Earlier work this paper cites.
OpenAI, “GPT-4 Technical Report,” 2023
2023
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” in International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , 2023, pp. 19 730–19 742
2023
Earlier work this paper cites.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. C. H. Hoi, “InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning,” in Thirty-seventh Conference on Neural Information Processing Systems , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Świechowski, K. Godlewski, B. Sawicki, and J. Mańdziuk, “Monte Carlo tree search: A review of recent modifications and applications,” Artificial Intelligence Review , vol. 56, no. 3, pp. 2497–2562, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Zhao, W. S. Lee, and D. Hsu, “Large Language Models as Commonsense Knowledge for Large-Scale Task Planning,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s verify step by step,” in The Twelfth International Conference on Learning Representations , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Li, Z. Lin, S. Zhang, Q. Fu, B. Chen, J.-G. Lou, and W. Chen, “Making language models better reasoners with step-aware verifier,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 5315–5333
2023
Earlier work this paper cites.
Z. Wu, Y. Hu, W. Shi, N. Dziri, A. Suhr, P. Ammanabrolu, N. A. Smith, M. Ostendorf, and H. Hajishirzi, “Fine-grained human feedback gives better rewards for language model training,” Advances in Neural Information Processing Systems , vol. 36, pp. 59 008–59 033, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Dubois, C. X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. Liang, and T. B. Hashimoto, “AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , 2023
2023
Earlier work this paper cites.
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, S. Gupta, B. P. Majumder, K. Hermann, S. Welleck, A. Yazdanbakhsh, and P. Clark, “Self-Refine: Iterative Refinement with Self-Feedback,” in Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 , 2023
2023
Earlier work this paper cites.
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, S. Liu, B. Sun, K. Liu, and J. Zhao, “Large Language Models are Better Reasoners with Self-Verification,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023 . Association for Computational Linguistics, 2023, pp. 2550–2575
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig, “Pal: Program-aided language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 10 764–10 799
2023
Earlier work this paper cites.
H. Lee, S. Phatale, H. Mansoor, K. R. Lu, T. Mesnard, J. Ferret, C. Bishop, E. Hall, V. Carbune, and A. Rastogi, “Rlaif: Scaling reinforcement learning from human feedback with ai feedback,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M.-L. Zhang, F. Yin, and C.-L. Liu, “A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from Diagram,” in IJCAI , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
O. Contributors, “OpenCompass: A Universal Evaluation Platform for Foundation Models,” 2023. [Online]. Available: https://github.com/open-compass/opencompass
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning , 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
X. Ning, Z. Lin, Z. Zhou, Z. Wang, H. Yang, and Y. Wang, “Skeleton-of-thought: Large language models can do parallel decoding,” Proceedings ENLSP-III , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman, “STaR: Self-taught reasoner bootstrapping reasoning with reasoning,” in Proc. the 36th International Conference on Neural Information Processing Systems , vol. 1126, 2024
2024
Earlier work this paper cites.
H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li, “Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning,” in The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2024
2024
Earlier work this paper cites.
OpenAI, “Hello GPT-4o,” May 2024. [Online]. Available: https://openai.com/index/hello-gpt-4o/
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
D. Zhang, Y. Yu, J. Dong, C. Li, D. Su, C. Chu, and D. Yu, “MM-LLMs: Recent Advances in MultiModal Large Language Models,” in Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024 . Association for Computational Linguistics, 2024, pp. 12 401–12 430
2024
Earlier work this paper cites.
OpenAI, “Learning to reason with LLMs,” September 2024. [Online]. Available: https://openai.com/index/learning-to-reason-with-llms/
2024
Earlier work this paper cites.
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe, “Let’s Verify Step by Step,” in The Twelfth International Conference on Learning Representations , 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk et al. , “Graph of thoughts: Solving elaborate problems with large language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 17 682–17 690
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
P. Wu and S. Xie, “V?: Guided Visual Search as a Core Mechanism in Multimodal LLMs,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 084–13 094
2024
Earlier work this paper cites.
Z. Chen, R. Sun, W. Liu, Y. Hong, and C. Gan, “GENOME: Generative Neuro-Symbolic Visual Reasoning by Growing and Reusing Modules,” in International Conference on Learning Representations , 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu et al. , “DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 1280–1297
2024
Earlier work this paper cites.
Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, B. Chang et al. , “A survey on in-context learning,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 1107–1128
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Q. Team, “QwQ: Reflect Deeply on the Boundaries of the Unknown,” November 2024. [Online]. Available: https://qwenlm.github.io/blog/qwq-32b-preview/
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
Z. Wan, X. Feng, M. Wen, S. M. McAleer, Y. Wen, W. Zhang, and J. Wang, “Alphazero-like tree-search can guide large language model decoding and training,” in Forty-first International Conference on Machine Learning , 2024
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
J. Liu, A. Cohen, R. Pasunuru, Y. Choi, H. Hajishirzi, and A. Celikyilmaz, “Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding,” in First Conference on Language Modeling , 2024. [Online]. Available: https://openreview.net/forum?id=kh9Zt2Ldmn
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
M. Kemmerling, D. Lütticke, and R. H. Schmitt, “Beyond games: a systematic review of neural Monte Carlo tree search applications,” Appl. Intell. , vol. 54, no. 11-12, pp. 1020–1046, 2024
2024
Cited alongside, same era.
F. Yu, A. Gao, and B. Wang, “OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning,” in Findings of the Association for Computational Linguistics: NAACL 2024 , 2024, pp. 858–875
2024
Cited alongside, same era.
P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui, “Math-shepherd: Verify and reinforce llms step-by-step without human annotations,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 9426–9439
2024
Cited alongside, same era.
J. Lu, Z. Dou, W. Hongru, Z. Cao, J. Dai, Y. Feng, and Z. Guo, “Autopsv: Automated process-supervised verifier,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Leng, J. Wang, J. Li, H. Zhang, Z. Hu, B. Zhang, H. Zhang, Y. Jiang, X. Li, D. Zhao, F. Wang, Y. Rong, A. Sun, and S. Lu, “Mmr1: Advancing the frontiers of multimodal reasoning,” https://github.com/LengSicong/MMR1 , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Zheng, J. Lu, S. Wang, and Y. Xiong, “EasyR1: An Efficient, Scalable, Multi-Modality RL Training Framework,” 2025. [Online]. Available: https://github.com/hiyouga/EasyR1
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
M. Luo, S. Tan, J. Wong, X. Shi, W. Tang, M. Roongta, C. Cai, J. Luo, T. Zhang, E. Li, R. A. Popa, and I. Stoica, “DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL,” 2025, notion Blog
2025
Closest in time.
J. Cheng, L. Li, G. Xiong, J. Shao, and Y. Lv, “Stop gamma decay: Min-form credit assignment is all process reward model needs for reasoning,” 2025, notion Blog
2025
Closest in time.
H. Team, “Open r1: A fully open reproduction of deepseek-r1.” 2025, github Project. [Online]. Available: https://github.com/huggingface/open-r1
2025
Closest in time.
J. Pan, J. Zhang, X. Wang, L. Yuan, H. Peng, and A. Suhr, “TinyZero,” 2025, accessed: 2025-01-24. [Online]. Available: https://github.com/Jiayi-Pan/TinyZero
2025
Closest in time.
Z. Liu, C. Chen, W. Li, T. Pang, C. Du, and M. Lin, “There may not be aha moment in r1-zero-like training — a pilot study,” 2025, notion Blog. [Online]. Available: https://oatllm.notion.site/oat-zero
2025
Closest in time.
Z. Liu, C. Chen, C. Du, W. S. Lee, and M. Lin, “Oat: A research-friendly framework for llm online alignment,” 2025. [Online]. Available: https://github.com/sail-sg/oat
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. Zhang, J. Yao, C. Ye, W. Xiong, and T. Zhang, “Online-dpo-r1: Unlocking effective reasoning without the ppo overhead,” 2025, notion Blog
2025
Closest in time.
J. Hu, Y. Zhang, Q. Han, D. Jiang, and H.-Y. S. Xiangyu Zhang, “Open-reasoner-zero: An open source approach to scaling reinforcement learning on the base model,” 2025. [Online]. Available: https://github.com/Open-Reasoner-Zero/Open-Reasoner-Zero
2025
Closest in time.
2025
Closest in time.
L. Chen, L. Li, H. Zhao, Y. Song, and Vinci, “R1-V: Reinforcing Super Generalization Ability in Vision-Language Models with Less Than $3,” 2025, accessed: 2025-02-02. [Online]. Available: https://github.com/Deep-Agent/R1-V
2025
Closest in time.
H. Shen, Z. Zhang, Q. Zhang, R. Xu, and T. Zhao, “Vlm-r1: A stable and generalizable r1-style large vision-language model,” 2025, accessed: 2025-02-15. [Online]. Available: https://github.com/om-ai-lab/VLM-R1
2025
Closest in time.
Y. Peng, G. Zhang, X. Geng, and X. Yang, “Lmm-r1,” 2025, accessed: 2025-02-13. [Online]. Available: https://github.com/TideDra/lmm-r1
2025
Closest in time.
X. Wang and P. Peng, “Open-r1-video,” 2025. [Online]. Available: https://github.com/Wang-Xiaodong1899/Open-R1-Video
2025
Closest in time.
B. Bizhe, W. Shao, and Q. Zhang, “Efficient-r1-vllm: Efficient rl-tuned moe vision-language model for reasoning,” 2025. [Online]. Available: https://github.com/baibizhe/Efficient-R1-VLLM
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Peng, Chris, X. Wang, Y. Wei, J. Pei, W. Qiu, A. Jian, Y. Hao, J. Pan, T. Xie, L. Ge, R. Zhuang, X. Song, Y. Liu, and Y. Zhou, “Skywork r1v: Pioneering multimodal reasoning with chain-of-thought,” 2025. [Online]. Available: https://huggingface.co/Skywork/Skywork-R1V-38B
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Liu, C. Chen, W. Li, T. Pang, C. Du, and M. Lin, “There May Not be Aha Moment in R1-Zero-like Training — A Pilot Study,” 2025, notion Blog. [Online]. Available: https://oatllm.notion.site/oat-zero
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
R. Shao, S. S. Li, R. Xin, S. Geng, Y. Wang, S. Oh, S. S. Du, N. Lambert, S. Min, R. Krishna, Y. Tsvetkov, H. Hajishirzi, P. W. Koh, and L. Zettlemoyer, “Spurious rewards: Rethinking training signals in rlvr,” https://rethink-rlvr.notion.site/Spurious-Rewards-Rethinking-Training-Signals-in-RLVR-1f4df34dac1880948858f95aeb88872f , 2025, notion Blog
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
C. White, S. Dooley, M. Roberts, A. Pal, B. Feuer, S. Jain, R. Shwartz-Ziv, N. Jain, K. Saifullah, S. Dey, Shubh-Agrawal, S. S. Sandha, S. V. Naidu, C. Hegde, Y. LeCun, T. Goldstein, W. Neiswanger, and M. Goldblum, “Livebench: A challenging, contamination-limited LLM benchmark,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=sKYHBTAxVa
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Cheng, Q. Chen, J. Zhang, H. Fei, X. Feng, W. Che, M. Li, and L. Qin, “Comt: A novel benchmark for chain of multi-modal thought on large vision-language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 22, 2025, pp. 23 678–23 686
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
G. DeepMind, “Gemini 2.5 Pro,” March 2025. [Online]. Available: https://deepmind.google/technologies/gemini/pro/
2025
Closest in time.
——, “Claude 3.7 Sonnet,” February 2025. [Online]. Available: https://www.anthropic.com/news/claude-3-7-sonnet
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Doubao, “Doubao-1.5-pro,” January 2025. [Online]. Available: https://team.doubao.com/en/special/doubao_1_5_pro
2025
Closest in time.
OpenAI, “OpenAI o4-mini,” April 2025. [Online]. Available: https://openai.com/index/openai-o3-mini/
2025
Closest in time.
J. Ouyang, R. Yan, Y. Luo, M. Cheng, Q. Liu, Z. Liu, S. Yu, and D. Wang, “Training powerful llm agents with end-to-end reinforcement learning,” GitHub, 2025. [Online]. Available: https://github.com/0russwest0/Agent-R1
2025
Closest in time.
2025
Closest in time.
Volcengine, “Verl: Volcano engine reinforcement learning,” n.d., accessed: 2025-04-21. [Online]. Available: https://github.com/volcengine/verl
2025
Closest in time.
B. Seed, “Seed-thinking v1.5,” 2023, accessed: 2025-04-21. [Online]. Available: https://github.com/ByteDance-Seed/Seed-Thinking-v1.5
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
X. Liang, J. Xiang, Z. Yu, J. Zhang, S. Hong, S. Fan, and X. Tang, “Openmanus: An open-source framework for building general ai agents,” 2025. [Online]. Available: https://doi.org/10.5281/zenodo.15186407
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.