Challenges of real-world reinforcement learning
Original
Gabriel Dulac-Arnold, Daniel J. Mankowitz, and Todd Hester. 2019 · 1904
Earlier work this paper cites.
Fine-tuning language models from human preferences
Original
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
A law of comparative judgement
Louis Leon Thurstone. 1927 · 1927
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Paul R. Rosenbaum and Donald B. Rubin. 1983 · 1983
Earlier work this paper cites.
A note on importance sampling using standardized weights
Augustine Kong. 1992 · 1992
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh. 2000 · 2000
Earlier work this paper cites.
Evaluation of machine translation and its evaluation
Joseph P Turian, Luke Shea, and I Dan Melamed. 2003 · 2003
Earlier work this paper cites.
Safe exploration for reinforcement learning
Alexander Hans, Daniel Schneegaß, Anton Maximilian Schäfer, and Steffen Udluft. 2008 · 2008
Earlier work this paper cites.
Exploration scavenging
John Langford, Alexander Strehl, and Jennifer Wortman. 2008 · 2008
Earlier work this paper cites.
Fast, cheap, and creative: Evaluating translation quality using Amazon’s Mechanical Turk
Chris Callison-Burch. 2009 · 2009
Earlier work this paper cites.
Learning dense models of query similarity from user click logs
Fabio De Bona, Stefan Riezler, Keith Hall, Massimiliano Ciaramita, Amaç Herdaǧdelen, and Maria Holmqvist. 2010 · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
Alexander L. Strehl, John Langford, Lihong Li, and Sham M. Kakade. 2010 · 2010
Earlier work this paper cites.
The process of post-editing: a pilot study
Michael Carl, Barbara Dragsted, Jakob Elming, Daniel Hardt, and Arnt Lykke Jakobsen. 2011 · 2011
Earlier work this paper cites.
An empirical evaluation of Thompson sampling
Olivier Chapelle and Lihong Li. 2011 · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudik, John Langford, and Lihong Li. 2011 · 2011
Earlier work this paper cites.
Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipanakar Ray, Patrice Simard, and Ed Snelson. 2013 · 2013
Earlier work this paper cites.