2018

Improving a Neural Semantic Parser by Counterfactual Learning from Human Bandit Feedback

Lawrence, Carolin, Riezler, Stefan

Understand

Counterfactual learning from human bandit feedback describes a scenario where user feedback on the quality of outputs of a historic system is logged and used to improve a target system.

  • We show how to apply this learning framework to neural semantic parsing.
  • From a machine learning perspective, the key challenge lies in a proper reweighting of the estimator so as to avoid known degeneracies in counterfactual learning, while still being applicable to stochastic gradient optimization.
  • To conduct experiments with human users, we devise an easy-to-use interface to collect human feedback on semantic parses.

Reading the bibliography…