Fetching the paper…
Reading the bibliography…
AI model alignment is crucial due to inadvertent biases in training data and the underspecified machine learning pipeline, where models with excellent test metrics may not meet end-user requirements.
Statistical inference under order restrictions : the theory and application of isotonic regression
Hugh D. Brunk, Richard E. Barlow, David J. Bartholomew, and Joan M. Bremner · 1973
Earlier work this paper cites.
Self-testing/correcting with applications to numerical problems
Manuel Blum, Michael Luby, and Ronitt Rubinfeld · 1993
Earlier work this paper cites.
Microeconomic Theory
Andreu Mas-Colell, Michael D. Whinston, and Jerry R. Green · 1995
Earlier work this paper cites.
Monotonicity hints
Joseph Sill and Yaser Abu-Mostafa · 1996
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Algorithmic Learning in a Random World, Second Edition
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer · 2005
Earlier work this paper cites.
Testing polynomials over general fields
Tali Kaufman and Dana Ron · 2006
Earlier work this paper cites.
Distribution-free property-testing
Shirley Halevy and Eyal Kushilevitz · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
On proximity-oblivious testing
Oded Goldreich and Dana Ron · 2008
Earlier work this paper cites.
Property testing: A learning theory perspective
Dana Ron · 2008
Earlier work this paper cites.
Optimal testing of reed-muller codes
Arnab Bhattacharyya, Swastik Kopparty, Grant Robert Schoenebeck, Madhu Sudan, and David Zuckerman · 2009
Earlier work this paper cites.
Testing juntas nearly optimally
Eric Blais · 2009
Earlier work this paper cites.
Nonparametric Estimation under Shape Constraints: Estimators, Algorithms and Asymptotics
Piet Groeneboom and Geurt Jongbloed · 2014
Earlier work this paper cites.
Combined Cycle Power Plant
Pnar Tfekci and Heysem Kaya · 2014
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Introduction to Property Testing
Oded Goldreich · 2017
Earlier work this paper cites.
Deep lattice networks and partial monotonic functions, 2017
Seungil You, David Ding, Kevin Canini, Jan Pfeifer, and Maya Gupta · 2017
Cited alongside, same era.
Prediction rule reshaping, 2018
Matt Bonakdarpour, Sabyasachi Chatterjee, Rina Foygel Barber, and John Lafferty · 2018
Cited alongside, same era.
Monotonic classification: an overview on algorithms, performance measures and data sets, 2018
José-Ramón Cano, Pedro Antonio Gutiérrez, Bartosz Krawczyk, Michał Woźniak, and Salvador García · 2018
Cited alongside, same era.
Supervising strong learners by amplifying weak experts, 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Cited alongside, same era.
Agi safety literature review, 2018
Tom Everitt, Gary Lea, and Marcus Hutter · 2018
Cited alongside, same era.
Ai safety via debate, 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Later among the works it cites.
Discovering language model behaviors with model-written evaluations, 2022
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan · 2022
Later among the works it cites.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals, 2022
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scalable agent alignment via reward modeling: a research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Cited alongside, same era.
Distribution-free testing of linear functions on ℝ n \mathbb{R}^{n}
Noah Fleming and Yuichi Yoshida · 2020
Cited alongside, same era.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Deontological ethics by monotonicity shape constraints, 2020
Serena Wang and Maya Gupta · 2020
Cited alongside, same era.
A survey on artificial intelligence assurance
Feras A. Batarseh, Laura Freeman, and Chih-Hao Huang · 2021
Cited alongside, same era.
Distribution-free, risk-controlling prediction sets, 2021
Stephen Bates, Anastasios Angelopoulos, Lihua Lei, Jitendra Malik, and Michael I. Jordan · 2021
Cited alongside, same era.
J.J. Vallon, N. Panjwani, X. Ling, S. Vij, E. Pollom, H.P. Bagshaw, S. Srinivas, J. Leppert, M. Bayati, and M.K. Buyyounouski · 2022
Later among the works it cites.
Low Degree Testing over the Reals , pages 738–792
Vipul Arora, Arnab Bhattacharyya, Noah Fleming, Esty Kelman, and Yuichi Yoshida · 2023
Later among the works it cites.
Uci machine learning repository, 2023
Dheeru Dua and Casey Graff · 2023
Later among the works it cites.
Aligning ai with shared human values, 2023
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2023
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Later among the works it cites.
Six Lectures on Linearized Neural Networks
Theodor Misiakiewicz and Andrea Montanari · 2023
Later among the works it cites.
Constrained monotonic neural networks, 2023
Davor Runje and Sharath M. Shankaranarayana · 2023
Later among the works it cites.
Model evaluation for extreme risks, 2023
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, and Allan Dafoe · 2023
Later among the works it cites.
Conformal risk control
Anastasios Nikolas Angelopoulos, Stephen Bates, Adam Fisch, Lihua Lei, and Tal Schuster · 2024
Closest in time.
Fairness in machine learning: A survey
Simon Caton and Christian Haas · 2024
Closest in time.
AI Safety, Ethics, and Society
Dan Hendrycks · 2024
Closest in time.
Ai alignment: A comprehensive survey, 2024
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Kwan Yee Ng, Juntao Dai, Xuehai Pan, Aidan O’Gara, Yingshan Lei, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao · 2024
Closest in time.
The alignment problem from a deep learning perspective, 2024
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2024
Closest in time.
On aligning prediction models with clinical experiential learning: A prostate cancer case study
Jacqueline J. Vallon, William Overman, Wanqiao Xu, Neil Panjwani, Xi Ling, Sushmita Vij, Hilary P. Bagshaw, John T. Leppert, Sumit Shah, Geoffrey Sonn, Sandy Srinivas, Erqi Pollom, Mark K. Buyyounouski, and Mohsen Bayati · 2024
Closest in time.
Mitigating llm hallucinations via conformal abstention, 2024
Yasin Abbasi Yadkori, Ilja Kuzborskij, David Stutz, András György, Adam Fisch, Arnaud Doucet, Iuliya Beloshapka, Wei-Hung Weng, Yao-Yuan Yang, Csaba Szepesvári, Ali Taylan Cemgil, and Nenad Tomasev · 2024
Closest in time.