Fetching the paper…
Reading the bibliography…
Recently, self-learning methods based on user satisfaction metrics and contextual bandits have shown promising results to enable consistent improvements in conversational AI systems.
Minmax optimization: Stable limit points of gradient descent ascent are locally optimal
Chi Jin, Praneeth Netrapalli, and Michael I Jordan. 2019 · 1902
Earlier work this paper cites.
Lessons from contextual bandit learning in a customer support bot
Nikos Karampatziakis, Sebastian Kochman, Jade Huang, Paul Mineiro, Kathy Osborne, and Weizhu Chen. 2019 · 1905
Earlier work this paper cites.
Linear stochastic bandits under safety constraints
Sanae Amani, Mahnoosh Alizadeh, and Christos Thrampoulidis. 2019 · 1908
Earlier work this paper cites.
Thompson sampling for contextual bandit problems with auxiliary safety constraints
Samuel Daulton, Shaun Singh, Vashist Avadhanula, Drew Dimmery, and Eytan Bakshy. 2019 · 1911
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard S Sutton. 1992 · 1992
Earlier work this paper cites.
Multi-armed bandits with limited exploration
Sudipto Guha and Kamesh Munagala. 2007 · 2007
Earlier work this paper cites.
Sunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal, Sungjin Lee, Young-Bum Kim, Spyros Matsoukas, and Ruhi Sarikaya. 2020 · 2010
Earlier work this paper cites.
On correlation and budget constraints in model-based bandit optimization with application to automatic machine learning
Matthew Hoffman, Bobak Shahriari, and Nando Freitas. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Off-policy evaluation for slate recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miroslav Dudík, John Langford, Damien Jose, and Imed Zitouni. 2016 · 2016
Cited alongside, same era.
The technology behind personal digital assistants: An overview of the system architecture and key components
Ruhi Sarikaya. 2017 · 2017
Cited alongside, same era.
Using contextual bandits with behavioral constraints for constrained online movie recommendation
Avinash Balakrishnan, Djallel Bouneffouf, Nicholas Mattei, and Francesca Rossi. 2018 · 2018
A primal dual formulation for deep learning with constraints
Yatin Nandwani, Abhishek Pathak, Parag Singla, et al. 2019 · 2019
Later among the works it cites.
Safe exploration for optimizing contextual bandits
Rolf Jagerman, Ilya Markov, and Maarten De Rijke. 2020 · 2020
Later among the works it cites.
Self-supervised contrastive learning for efficient user satisfaction prediction in conversational agents
Mohammad Kachuee, Hao Yuan, Young-Bum Kim, and Sungjin Lee. 2021 · 2021
Later among the works it cites.
Han Li, Sunghyun Park, Aswarth Dara, Jinseok Nam, Sungjin Lee, Young-Bum Kim, Spyros Matsoukas, and Ruhi Sarikaya. 2021 · 2021
Later among the works it cites.
Learning from extreme bandit feedback
Romain Lopez, Inderjit S Dhillon, and Michael I Jordan. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning with logged bandit feedback
Thorsten Joachims, Adith Swaminathan, and Maarten de Rijke. 2018 · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado van Hasselt, and David Silver. 2018 · 2018
Cited alongside, same era.
Scalable and robust self-learning for skill routing in large-scale conversational ai systems
Mohammad Kachuee, Jinseok Nam, Sarthak Ahuja, Jin-Myung Won, and Sungjin Lee. 2022 · 2022
Closest in time.