Fetching the paper…
Reading the bibliography…
The current paradigm of test-time scaling relies on generating long reasoning traces ("thinking" more) before producing a response.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Emergence of pragmatics from referential game between theory of mind agents, 2021a
Luyao Yuan, Zipeng Fu, Jingyue Shen, Lu Xu, Junhong Shen, and Song-Chun Zhu · 2001
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Tingting Zhao, Hirotaka Hachiya, Gang Niu, and Masashi Sugiyama · 2011
Earlier work this paper cites.
Policy gradients with variance related risk criteria
Dotan Di Castro, Aviv Tamar, and Shie Mannor · 2012
Earlier work this paper cites.
Teacher–student curriculum learning
Tambet Matiisen, Avital Oliver, Taco Cohen, and John Schulman · 2017
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Dwight Crow, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Joshua Tobin, P. Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M. Bayen, Sham M. Kakade, Igor Mordatch, and P. Abbeel · 2018
Earlier work this paper cites.
Curriculum learning for reinforcement learning domains: A framework and survey
Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E. Taylor, and Peter Stone · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Iterative teacher-aware learning
Luyao Yuan, Dongruo Zhou, Junhong Shen, Jingdong Gao, Jeffrey L Chen, Quanquan Gu, Ying Nian Wu, and Song-Chun Zhu · 2021
Earlier work this paper cites.
Theoretically principled deep rl acceleration via nearest neighbor function approximation
Junhong Shen and Lin F. Yang · 2021
Earlier work this paper cites.
Mathematical reconstruction of patient-specific vascular networks based on clinical images and global optimization
Junhong Shen, Abdul Hannan Faruqi, Yifan Jiang, and Nima Maftoon · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mo Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
A survey on curriculum learning
Xin Wang, Yudong Chen, and Wenwu Zhu · 2021
Earlier work this paper cites.
Efficient architecture search for diverse tasks
Junhong Shen, Mikhail Khodak, and Ameet Talwalkar · 2022
Earlier work this paper cites.
NAS-bench-360: Benchmarking neural architecture search on diverse tasks
Renbo Tu, Nicholas Roberts, Mikhail Khodak, Junhong Shen, Frederic Sala, and Ameet Talwalkar · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman · 2022
Earlier work this paper cites.
Cogagent: A visual language model for gui agents, 2023
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang · 2023
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Samuel Stevens, Boshi Wang, Huan Sun, and Yu Su · 2023
Earlier work this paper cites.
Cross-modal fine-tuning: align then refine
Junhong Shen, Liam Li, Lucio M. Dery, Corey Staten, Mikhail Khodak, Graham Neubig, and Ameet Talwalkar · 2023
Earlier work this paper cites.
Webglm: Towards an efficient web-enhanced question answering system with human preferences, 2023
Xiao Liu, Hanyu Lai, Hao Yu, Yifan Xu, Aohan Zeng, Zhengxiao Du, Peng Zhang, Yuxiao Dong, and Jie Tang · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models, 2023
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Webvoyager: Building an end-to-end web agent with large multimodal models, 2024
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu · 2024
Earlier work this paper cites.
Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 2024
Claude · 2024
Earlier work this paper cites.
Browser use: Enable ai to control your browser, 2024
Magnus Müller and Gregor Žunič · 2024
Cited alongside, same era.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Yuexiang Zhai, Hao Bai, Zipeng Lin, Jiayi Pan, Shengbang Tong, Yifei Zhou, Alane Suhr, Saining Xie, Yann LeCun, Yi Ma, and Sergey Levine · 2024
Cited alongside, same era.
Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Hao Bai, Yifei Zhou, Mert Cemri, Jiayi Pan, Alane Suhr, Sergey Levine, and Aviral Kumar · 2024
Cited alongside, same era.
Inference-aware fine-tuning for best-of-n sampling in large language models
Yinlam Chow, Guy Tennenholtz, Izzeddin Gur, Vincent Zhuang, Bo Dai, Sridhar Thiagarajan, Craig Boutilier, Rishabh Agarwal, Aviral Kumar, and Aleksandra Faust · 2024
Cited alongside, same era.
Multimodal web navigation with instruction-finetuned foundation models, 2024
Awa 1.5 achieves breakthrough performance on webarena benchmark, 2024
JaceAI · 2024
Later among the works it cites.
pi0.5: a vision-language-action model with open-world generalization
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Allen Z. Ren, Lucy Xiaoyang Shi, Laura Smith, Jost Tobias Springenberg, Kyle Stachowicz, James Tanner, Quan Vuong, Homer Rich Walke, Anna Walling, Haohuan Wang, Lili Yu, and Ury Zhilinsky · 2025
Closest in time.
Mixture-of-mamba: Enhancing multi-modal state-space models with modality-aware sparsity, 2025
Weixin Liang, Junhong Shen, Genghan Zhang, Ning Dong, Luke Zettlemoyer, and Lili Yu · 2025
Closest in time.
Hi robot: Open-ended instruction following with hierarchical vision-language-action models
Lucy Xiaoyang Shi, Brian Ichter, Michael Equi, Liyiming Ke, Karl Pertsch, Quan Vuong, James Tanner, Anna Walling, Haohuan Wang, Niccolo Fusai, Adrian Li-Bell, Danny Driess, Lachy Groom, Sergey Levine, and Chelsea Finn · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur · 2024
Cited alongside, same era.
Agentoccam: A simple yet strong baseline for llm-based web agents, 2024
Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, and Huzefa Rangwala · 2024
Cited alongside, same era.
Beyond browsing: Api-based web agents, 2024
Yueqi Song, Frank Xu, Shuyan Zhou, and Graham Neubig · 2024
Cited alongside, same era.
Autonomous evaluation and refinement of digital agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr · 2024
Cited alongside, same era.
Yao Fu, Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, and Honglak Lee · 2024
Cited alongside, same era.
OpenAI · 2024
Cited alongside, same era.
Introducing the next generation of claude, 2024
Anthropic · 2024
Cited alongside, same era.
Step: Stacked llm policies for web actions
Paloma Sodhi, S. R. K. Branavan, Yoav Artzi, and Ryan McDonald · 2024
Cited alongside, same era.
Symbiotic cooperation for web agents: Harnessing complementary strengths of large and small llms
Ruichen Zhang, Mufan Qiu, Zhen Tan, Mohan Zhang, Vincent Lu, Jie Peng, Kaidi Xu, Leandro Z. Agudelo, Peter Qian, and Tianlong Chen · 2025
Closest in time.
Digi-q: Learning q-value functions for training device-control agents
Hao Bai, Yifei Zhou, Erran L. Li, Sergey Levine, and Aviral Kumar · 2025
Closest in time.
Webrl: Training llm web agents via self-evolving online curriculum reinforcement learning, 2025
Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Wenyi Zhao, Yu Yang, Xinyue Yang, Jiadai Sun, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong · 2025
Closest in time.
Gemini deep research, 2025
Google Gemini · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Fei-Fei Li, Hanna Hajishirzi, Luke S. Zettlemoyer, Percy Liang, Emmanuel J. Candes, and Tatsunori Hashimoto · 2025
Closest in time.
Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving
Yangzhen Wu, Zhiqing Sun, Shanda Li, Sean Welleck, and Yiming Yang · 2025
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning
Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2025
Closest in time.
Cat: Content-adaptive image tokenization, 2025
Junhong Shen, Kushal Tirumala, Michihiro Yasunaga, Ishan Misra, Luke Zettlemoyer, Lili Yu, and Chunting Zhou · 2025
Closest in time.
Codepde: An inference framework for llm-driven pde solver generation, 2025
Shanda Li, Tanya Marwah, Junhong Shen, Weiwei Sun, Andrej Risteski, Yiming Yang, and Ameet Talwalkar · 2025
Closest in time.
Skillweaver: Web agents can self-improve by discovering and honing skills
Boyuan Zheng, Michael Y. Fatemi, Xiaolong Jin, Zora Zhiruo Wang, Apurva Gandhi, Yueqi Song, Yu Gu, Jayanth Srinivasa, Gaowen Liu, Graham Neubig, and Yu Su · 2025
Closest in time.
Plan-and-act: Improving planning of agents for long-horizon tasks, 2025
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami · 2025
Closest in time.
Towards internet-scale training for agents
Brandon Trabucco, Gunnar Sigurdsson, Robinson Piramuthu, and Ruslan Salakhutdinov · 2025
Closest in time.
Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments
Hongjin Su, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan Ö. Arik · 2025
Closest in time.
Scaling test-time compute without verification or rl is suboptimal
Amrith Rajagopal Setlur, Nived Rajaraman, Sergey Levine, and Aviral Kumar · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Optimizing test-time compute via meta reinforcement fine-tuning
Yuxiao Qu, Matthew Y. R. Yang, Amrith Rajagopal Setlur, Lewis Tunstall, Edward Beeching, Ruslan Salakhutdinov, and Aviral Kumar · 2025
Closest in time.
Two heads are better than one: Test-time scaling of multi-agent collaborative reasoning
Can Jin, Hongwu Peng, Qixin Zhang, Yujin Tang, Dimitris N. Metaxas, and Tong Che · 2025
Closest in time.
Research: Learning to reason with search for llms via reinforcement learning, 2025
Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen · 2025
Closest in time.
Think twice, act once: A co-evolution framework of llm and rl for large-scale decision making, 2025
Xu Wan, Wenyue Xu, Chao Yang, and Mingyang Sun · 2025
Closest in time.
Gemma Team · 2025
Closest in time.
Ui-tars: Pioneering automated gui interaction with native agents
Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, et al · 2025
Closest in time.
Cognitive behaviors that enable self-improving reasoners, or, four habits of highly effective stars
Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, nathan lile, and Noah D. Goodman · 2025
Closest in time.
Towards enterprise-ready computer using generalist agent
Sami Marreed, Alon Oved, Avi Yaeli, Segev Shlomov, Ido Levy, Aviad Sela, Asaf Adi, and Nir Mashkif · 2025
Closest in time.