Fetching the paper…
Reading the bibliography…
The emergence of large language models (LLMs) has sparked discussion on Artificial Superintelligence (ASI), a hypothetical AI system that surpasses human intelligence.
Abram Demski and Scott Garrabrant · 1902
Earlier work this paper cites.
An analysis of transformations revisited
Peter J. Bickel and Kjell A. Doksum · 1981
Earlier work this paper cites.
Evaluation may be easier than generation (extended abstract)
Moni Naor · 1996
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Co-training and expansion: towards bridging theory and practice
Maria-Florina Balcan, Avrim Blum, and Ke Yang · 2004
Earlier work this paper cites.
Spectral Analysis of Large Dimensional Random Matrices
Zhidong Bai and Jack W. Silverstein · 2010
Earlier work this paper cites.
Theoretical analysis of self-training with deep networks on unlabeled data, 2022a
Colin Wei, Kendrick Shen, Yining Chen, and Tengyu Ma · 2010
Earlier work this paper cites.
Artificial general intelligence: concept, state of the art, and future prospects
Ben Goertzel · 2014
Earlier work this paper cites.
Superintelligence: Paths, dangers, strategies
Bostrom Nick · 2014
Earlier work this paper cites.
Artificial superintelligence: Extinction or nirvana?
Jens Pohl · 2015
Earlier work this paper cites.
Concrete problems in ai safety, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
Data programming: creating large training sets, quickly
Alexander Ratner, Christopher De Sa, Sen Wu, Daniel Selsam, and Christopher Ré · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Thinking fast and slow with deep learning and tree search
Thomas Anthony, Zheng Tian, and David Barber · 2017
Earlier work this paper cites.
Artificial intelligence in life extension: from deep learning to superintelligence
Mikhail Batin, Alexey Turchin, Markov Sergey, Alisa Zhila, and David Denkenberger · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts, 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Growth, degrowth, and the challenge of artificial superintelligence
Salvador Pueyo · 2018
Earlier work this paper cites.
Reframing superintelligence: Comprehensive ai services as general intelligence
K Eric Drexler · 2019
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, et al · 2020
Earlier work this paper cites.
Fast and three-rious: speeding up weak supervision with triplet methods
Daniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper, Kayvon Fatahalian, and Christopher Ré · 2020
Earlier work this paper cites.
On the theory of transfer learning: the importance of task diversity
Nilesh Tripuraneni, Michael I. Jordan, and Chi Jin · 2020
Earlier work this paper cites.
To trust or to think: Cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos · 2021
Earlier work this paper cites.
A theory of label propagation for subpopulation shift
Tianle Cai, Ruiqi Gao, Jason Lee, and Qi Lei · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, et al · 2021
Earlier work this paper cites.
Comatch: Semi-supervised learning with contrastive graph regularization
Junnan Li, Caiming Xiong, and Steven C. H. Hoi · 2021
Earlier work this paper cites.
Anchoring bias affects mental model formation and user reliance in explainable ai systems
Mahsan Nourani, Chiradeep Roy, Jeremy E Block, Donald R Honeycutt, Tahrima Rahman, Eric Ragan, and Vibhav Gogate · 2021
Earlier work this paper cites.
Asymptotics of ridge(less) regression under general source condition
Dominic Richards, Jaouad Mourtada, and Lorenzo Rosasco · 2021
Earlier work this paper cites.
On the utility of prediction sets in human-ai teams
Varun Babbar, Umang Bhatt, and Adrian Weller · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, et al · 2022
Earlier work this paper cites.
Role of human-ai interaction in selective prediction
Elizabeth Bondi, Raphael Koster, Hannah Sheahan, Martin Chadwick, Yoram Bachrach, Taylan Cemgil, Ulrich Paquet, and Krishnamurthy Dvijotham · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models, 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, et al · 2022
Earlier work this paper cites.
Formalizing the presumption of independence, 2022
Paul Christiano, Eric Neyman, and Mark Xu · 2022
Earlier work this paper cites.
Towards artificial general intelligence via a multimodal foundation model
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, et al · 2022
Earlier work this paper cites.
Who goes first? influences of human-ai workflow on decision making in clinical imaging
Riccardo Fogliato, Shreya Chappidi, Matthew Lungren, Paul Fisher, Diane Wilson, Michael Fitzke, Mark Parkinson, Eric Horvitz, Kori Inkpen, and Besmira Nushi · 2022
Earlier work this paper cites.
Optimal weak to strong learning, 2022
Kasper Green Larsen and Martin Ritzert · 2022
Earlier work this paper cites.
Human-ai collaboration in decision-making: Beyond learning to defer, 2022
Diogo Leitão, Pedro Saleiro, Mário A. T. Figueiredo, and Pedro Bizarro · 2022
Earlier work this paper cites.
Active evaluation: Efficient NLG evaluation with few pairwise comparisons
Akash Kumar Mohankumar and Mitesh Khapra · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Quality: Question answering with long input texts, yes!, 2022
Richard Yuanzhe Pang, Alicia Parrish, Nitish Joshi, Nikita Nangia, Jason Phang, Angelica Chen, Vishakh Padmakumar, Johnny Ma, Jana Thompson, He He, and Samuel R. Bowman · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators, 2022
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners, 2022
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei · 2022
Earlier work this paper cites.
Star: Bootstrapping reasoning with reasoning, 2022
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman · 2022
Earlier work this paper cites.
Scalable ai safety via doubly-efficient debate, 2023
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision, 2023
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeff Wu · 2023
Earlier work this paper cites.
DuNST: Dual noisy self training for semi-supervised controllable text generation
Yuxi Feng, Xiaoyuan Yi, Xiting Wang, Laks Lakshmanan, V.S., and Xing Xie · 2023
Earlier work this paper cites.
The capacity for moral self-correction in large language models, 2023
Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I. Liao, Kamilė Lukošiūtė, Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, Dawn Drain, Dustin Li, Eli Tran-Johnson, Ethan Perez, et al · 2023
Earlier work this paper cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2023
Cited alongside, same era.
Reinforced self-training (rest) for language modeling, 2023
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas · 2023
Cited alongside, same era.
Learning to defer with limited expert predictions
Patrick Hemmer, Lukas Thede, Michael Vössing, Johannes Jakubik, and Niklas Kühl · 2023
Cited alongside, same era.
Advancing human-ai complementarity: The impact of user expertise and algorithmic tuning on joint decision making
Kori Inkpen, Shreya Chappidi, Keri Mallari, Besmira Nushi, Divya Ramesh, Pietro Michelucci, Vani Mandava, Libuše Hannah Vepřek, and Gabrielle Quinn · 2023
Cited alongside, same era.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis · 2023
Improving weak-to-strong generalization with scalable oversight and ensemble learning, 2024
Jitao Sang, Yuhang Wang, Jing Zhang, Yanxu Zhu, Chao Kong, Junhong Ye, Shuyu Wei, and Jinlin Xiao · 2024
Closest in time.
A critical evaluation of AI feedback for aligning large language models
Archit Sharma, Sedrick Keh, Eric Mitchell, Chelsea Finn, Kushal Arora, and Thomas Kollar · 2024
Closest in time.
The curse of recursion: Training on generated data makes models forget, 2024
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot, and Ross Anderson · 2024
Closest in time.
Easy-to-hard generalization: Scalable alignment beyond human supervision, 2024
Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang, Sean Welleck, and Chuang Gan · 2024
Closest in time.
Asking the right question at the right time: Human and model uncertainty guidance to ask clarification questions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Critique ability of large language models, 2023
Liangchen Luo, Zi Lin, Yinxiao Liu, Lei Shu, Yun Zhu, Jingbo Shang, and Lei Meng · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2023
Cited alongside, same era.
Who should predict? exact algorithms for learning to defer to humans, 2023
Hussein Mozannar, Hunter Lang, Dennis Wei, Prasanna Sattigeri, Subhro Das, and David Sontag · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Investigating factored cognition in large language models for answering ethically nuanced questions
Benjamin Sturgeon, Brian Muhia, and Jonas Kgomo · 2023
Cited alongside, same era.
Explanations can reduce overreliance on ai systems during decision-making
Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna · 2023
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Cited alongside, same era.
Alberto Testoni and Raquel Fernández · 2024
Closest in time.
On the essence and prospect: An investigation of alignment approaches for big models
Xinpeng Wang, Shitong Duan, Xiaoyuan Yi, Jing Yao, Shanlin Zhou, Zhihua Wei, Peng Zhang, Dongkuan Xu, Maosong Sun, and Xing Xie · 2024
Closest in time.
Self-preference bias in LLM-as-a-judge
Koki Wataoka, Tsubasa Takahashi, and Ryokan Ri · 2024
Closest in time.
Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness, 2024
Tianyu Yu, Haoye Zhang, Yuan Yao, Yunkai Dang, Da Chen, Xiaoman Lu, Ganqu Cui, Taiwen He, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2024
Closest in time.
Self-rewarding language models, 2024
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston · 2024
Closest in time.
Ensemw2s: Enhancing weak-to-strong generalization with large language model ensembles, 2025
Aakriti Agrawal, Mucong Ding, Zora Che, Chenghao Deng, Anirudh Satheesh, Bang An, Bayan Bruss, John Langford, and Furong Huang · 2025
Closest in time.
Ai deception: Risks, dynamics, and controls, 2025
Boyuan Chen, Sitong Fang, Jiaming Ji, Yanxu Zhu, Pengcheng Wen, Jinzhou Wu, Yingshui Tan, Boren Zheng, Mengying Yuan, Wenqi Chen, Donghai Hong, Alex Qiu, Xin Chen, Jiayi Zhou, Kaile Wang, Juntao Dai, Borong Zhang, Tianzhuo Yang, et al · 2025
Closest in time.
Post-completion learning for language models, 2025
Xiang Fei, Siqi Wang, Shu Wei, Yuxiang Nie, Wei Shi, Hao Feng, Chao Feng, and Can Huang · 2025
Closest in time.
Heshan Fernando, Han Shen, Parikshit Ram, Yi Zhou, Horst Samulowitz, Nathalie Baracaldo, and Tianyi Chen · 2025
Closest in time.
Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, Sharon Li, Andy Zou, Lionel Levine, Bo Han, Jie Fu, Ziwei Liu, Jinwoo Shin, Kimin Lee, Mantas Mazeika, Long Phan, George Ingebretsen, Adam Khoja, Cihang Xie, Olawale Salaudeen, Matthias Hein, Kevin Zhao, Alexander Pan, David Duvenaud, Bo Li, Steve Omohundro, Gabriel Alfour, Max Tegmark, Kevin McGrew, Gary Marcus, Jaan Tallinn, Eric Schmidt, and Yoshua Bengio · 2025
Closest in time.
M. Emrullah Ildiz, Halil Alperen Gozeten, Ege Onur Taga, Marco Mondelli, and Samet Oymak · 2025
Closest in time.
Weak-to-strong generalization under distribution shifts
Myeongho Jeon, Jan Sobotka, Suhwan Choi, and Maria Brbic · 2025
Closest in time.
ReFLAIR: Enhancing multimodal reasoning via structured reflection and reward-guided learning
Jiazhou Ji and Xinru Lu · 2025
Closest in time.
Contrastive weak-to-strong generalization, 2025
Houcheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang, Chen Gao, Yong Li, Xiang Wang, Xiangnan He, and Yang Deng · 2025
Closest in time.
Multiple LLM agents debate for equitable cultural alignment
Dayeon Ki, Rachel Rudinger, Tianyi Zhou, and Marine Carpuat · 2025
Closest in time.
Safework-r1: Coevolving safety and intelligence under the ai-45 ∘
Shanghai AI Lab · 2025
Closest in time.
Selective weak-to-strong generalization, 2025
Hao Lang, Fei Huang, and Yongbin Li · 2025
Closest in time.
Curriculum-rlaif: Curriculum alignment with reinforcement learning from ai feedback, 2025
Mengdi Li, Jiaye Lin, Xufeng Zhao, Wenhao Lu, Peilin Zhao, Stefan Wermter, and Di Wang · 2025
Closest in time.
MACPO: Weak-to-strong alignment via multi-agent contrastive preference optimization
Yougang Lyu, Lingyong Yan, Zihan Wang, Dawei Yin, Pengjie Ren, Maarten de Rijke, and Zhaochun Ren · 2025
Closest in time.
On the mechanisms of weak-to-strong generalization: A theoretical perspective
Behrad Moniri and Hamed Hassani · 2025
Closest in time.
Relating misfit to gain in weak-to-strong generalization beyond the squared loss
Abhijeet Mulgund and Chirag Pabbaraju · 2025
Closest in time.
Aligning artificial superintelligence via a multi-box protocol
Avraham Yair Negozio · 2025
Closest in time.
Emergence of human-like polarization among large language model agents, 2025
Jinghua Piao, Zhihong Lu, Chen Gao, Fengli Xu, Qinghua Hu, Fernando P. Santos, Yong Li, and James Evans · 2025
Closest in time.
Towards understanding sycophancy in language models, 2025
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez · 2025
Closest in time.
How to mitigate overfitting in weak-to-strong generalization?
Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng, Qipeng Guo, and Xipeng Qiu · 2025
Closest in time.
Toshiyuki Shigemura · 2025
Closest in time.
Weak-to-strong generalization through the data-centric lens
Changho Shin, John Cooper, and Frederic Sala · 2025
Closest in time.
CortexDebate: Debating sparsely and equally for multi-agent debate
Yiliu Sun, Zicheng Zhao, Sheng Wan, and Chen Gong · 2025
Closest in time.
Learning task decomposition to assist humans in competitive programming, 2025
Jiaxin Wen, Ruiqi Zhong, Pei Ke, Zhihong Shao, Hongning Wang, and Minlie Huang · 2025
Closest in time.
Alice: Proactive learning with teacher’s demonstrations for weak-to-strong generalization, 2025
Shujin Wu, Cheng Qian, Yi R. Fung, Paul Pu Liang, and Heng Ji · 2025
Closest in time.
Online iterative self-alignment for radiology report generation
Ting Xiao, Lei Shi, Yang Zhang, HaoFeng Yang, Zhe Wang, and Chenjia Bai · 2025
Closest in time.
On the emergence of weak-to-strong generalization: A bias-variance perspective, 2025
Gengze Xu, Wei Yao, Ziqiao Wang, and Yong Liu · 2025
Closest in time.
Yihao Xue, Jiping Li, and Baharan Mirzasoleiman · 2025
Closest in time.
Diverse AI feedback for large language model alignment
Tianshu Yu, Ting-En Lin, Yuchuan Wu, Min Yang, Fei Huang, and Yongbin Li · 2025
Closest in time.
Llama 4: Advancing multimodal intelligence
Meta AI · 2026
Closest in time.
Claude 4.5: Our most intelligent model yet
Anthropic · 2026
Closest in time.
Codecontests-o: Powering llms via feedback-driven iterative test case generation, 2026
Jianfeng Cai, Jinhua Zhu, Ruopei Sun, Kangwen Zhao, Dongyun Xue, Mingxiao Feng, Wengang Zhou, and Houqiang Li · 2026
Closest in time.
Towards scalable oversight with collaborative multi-agent debate in error detection, 2026
Yongqiang Chen, Gang Niu, Yu Mao, James Cheng, Bo Han, and Masashi Sugiyama · 2026
Closest in time.
The three kinds of ai based on capabilities
IBM Data and AI Team · 2026
Closest in time.
Gemini 3.1 pro model card
Google DeepMind · 2026
Closest in time.
Why self-rewarding works: Theoretical guarantees for iterative alignment of language models, 2026
Shi Fu, Yingjie Wang, Shengchao Hu, Peng Wang, and Dacheng Tao · 2026
Closest in time.
HyunJin Kim, Xiaoyuan Yi, Jing Yao, Muhua Huang, JinYeong Bak, James Evans, and Xing Xie · 2026
Closest in time.
Introducing gpt-5.4
OpenAI · 2026
Closest in time.
A benchmark of expert-level academic questions to assess ai capabilities
Long Phan, Alice Gatti, Nathaniel Li, others, and HLE Contributors Consortium · 2026
Closest in time.
Qwen3.5: Towards native multimodal agents
Qwen Team · 2026
Closest in time.
Scalable oversight for superhuman ai via recursive self-critiquing, 2026
Xueru Wen, Jie Lou, Xinyu Lu, Junjie Yang, Yanjiang Liu, Yaojie Lu, Debing Zhang, and Xing Yu · 2026
Closest in time.
Ruimeng Ye, Zihan Wang, Yang Xiao, Zinan Ling, Manling Li, and Bo Hui · 2026
Closest in time.
Knowledge divergence and the value of debate for scalable oversight, 2026
Robin Young · 2026
Closest in time.