Fetching the paper…
Reading the bibliography…
The goal in offline data-driven decision-making is synthesize decisions that optimize a black-box utility function, using a previously-collected static dataset, with no active interaction.
Conditioning by adaptive sampling for robust design
David H Brookes, Hahnbeom Park, and Jennifer Listgarten · 1901
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Population-based black-box optimization for biological sequence design
Christof Angermueller, David Belanger, Andreea Gane, Zelda Mariet, David Dohan, Kevin Murphy, Lucy Colwell, and D Sculley · 2006
Earlier work this paper cites.
The CMA evolution strategy: A comparing review
Nikolaus Hansen · 2006
Earlier work this paper cites.
On integral probability metrics, \ \backslash phi-divergences and binary classification
Bharath K Sriperumbudur, Kenji Fukumizu, Arthur Gretton, Bernhard Schölkopf, and Gert RG Lanckriet · 2009
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Earlier work this paper cites.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P. Adams · 2012
Earlier work this paper cites.
Auto-encoding variational bayes, 2013
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation
Yaroslav Ganin and Victor Lempitsky · 2015
Earlier work this paper cites.
Scalable bayesian optimization using deep neural networks
Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Mostofa Patwary, Mr Prabhat, and Ryan Adams · 2015
Earlier work this paper cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Learning representations for counterfactual inference
Fredrik Johansson, Uri Shalit, and David Sontag · 2016
Earlier work this paper cites.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Rémi Tachet des Combes · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
The reparameterization trick for acquisition functions
James T. Wilson, Riccardo Moriconi, Frank Hutter, and Marc Peter Deisenroth · 2017
Cited alongside, same era.
A data-driven statistical model for predicting the critical temperature of a superconductor
Kam Hamidieh · 2018
Cited alongside, same era.
Stable baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu · 2018
Cited alongside, same era.
Distributionally robust bayesian optimization
Johannes Kirschner, Ilija Bogunovic, Stefanie Jegelka, and Andreas Krause · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Chip placement with deep reinforcement learning
Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, Joe Jiang, Ebrahim Songhori, Shen Wang, Young-Joon Lee, Eric Johnson, Omkar Pathak, Sungmin Bae, et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning weighted representations for generalization across designs
Fredrik D Johansson, Nathan Kallus, Uri Shalit, and David Sontag · 2018
Cited alongside, same era.
Representation balancing mdps for off-policy policy evaluation
Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo Faisal, Finale Doshi-Velez, and Emma Brunskill · 2018
Cited alongside, same era.
Representation learning for treatment effect estimation from observational data
Liuyi Yao, Sheng Li, Yaliang Li, Mengdi Huai, Jing Gao, and Aidong Zhang · 2018
Cited alongside, same era.
Model inversion networks for model-based optimization
Aviral Kumar and Sergey Levine · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Safe policy improvement with an estimated baseline policy
Thiago D Simão, Romain Laroche, and Rémi Tachet des Combes · 2019
Cited alongside, same era.
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C Courville, and Marc Bellemare · 2021
Later among the works it cites.
Low-n protein engineering with data-efficient deep learning
Surojit Biswas, Grigory Khimulya, Ethan C Alley, Kevin M Esvelt, and George M Church · 2021
Later among the works it cites.
Offline model-based optimization via normalized maximum likelihood estimation
Justin Fu and Sergey Levine · 2021
Later among the works it cites.
Bias-robust bayesian optimization via dueling bandits
Johannes Kirschner and Andreas Krause · 2021
Later among the works it cites.
Data-driven offline optimization for architecting hardware accelerators
Aviral Kumar, Amir Yazdanbakhsh, Milad Hashemi, Kevin Swersky, and Sergey Levine · 2021
Later among the works it cites.
Representation balancing offline model-based reinforcement learning
Byung-Jun Lee, Jongmin Lee, and Kee-Eung Kim · 2021
Later among the works it cites.
Uncertainty quantification using martingales for misspecified gps
Willie Neiswanger and Aaditya Ramdas · 2021
Later among the works it cites.
Value-at-risk optimization with gaussian processes
Quoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, and Patrick Jaillet · 2021
Later among the works it cites.
Offline neural contextual bandits: Pessimism, optimization and generalization
Thanh Nguyen-Tang, Sunil Gupta, A Tuan Nguyen, and Svetha Venkatesh · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch rl can be exponentially harder than online rl
Andrea Zanette · 2021
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Closest in time.