Fetching the paper…
Reading the bibliography…
Document summarisation can be formulated as a sequential decision-making problem, which can be solved by Reinforcement Learning (RL) algorithms.
Simple statistical gradient-following
Ronald J. Williams · 1992
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y. Ng · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
A scalable global model for summarization
Dan Gillick and Benoit Favre · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S. Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Maximum margin ranking algorithms for information retrieval
Shivani Agarwal and Michael Collins · 2010
Earlier work this paper cites.
A short introduction to learning to rank
Hang Li · 2011
Earlier work this paper cites.
Framework of automatic text summarization using reinforcement learning
Seonggi Ryang and Takeshi Abekawa · 2012
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim · 2014
Earlier work this paper cites.
Fear the REAPER: A system for automatic multi-document
Cody Rioux, Sadid A. Hasan, and Yllias Chali · 2014
Earlier work this paper cites.
Learning summary prior representation for extractive summarization
Ziqiang Cao, Furu Wei, Sujian Li, Wenjie Li, Ming Zhou, and Houfeng Wang · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Cited alongside, same era.
Improving multi-document summarization via text classification
Ziqiang Cao, Wenjie Li, Sujian Li, and Furu Wei · 2017
Cited alongside, same era.
Just sort it! A simple and effective approach to active preference learning
Lucas Maystre and Matthias Grossglauser · 2017
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
APRIL: Interactively learning to summarise by combining active preference learning and reinforcement learning
Yang Gao, Christian M. Meyer, and Iryna Gurevych · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Later among the works it cites.
Reliability and learnability of human bandit feedback for sequence-to-sequence reinforcement learning
Julia Kreutzer, Joshua Uyheng, and Stefan Riezler · 2018
Later among the works it cites.
Improving abstraction in text summarization
Wojciech Kryscinski, Romain Paulus, Caiming Xiong, and Richard Socher · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning for bandit neural machine translation with simulated human feedback
Khanh Nguyen, Hal Daumé III, and Jordan L. Boyd-Graber · 2017
Cited alongside, same era.
A principled framework for evaluating summarizers: Comparing models of summary quality against human judgments
Maxime Peyrard and Judith Eckle-Kohler · 2017
Cited alongside, same era.
Sentence simplification with deep reinforcement learning
Xingxing Zhang and Mirella Lapata · 2017
Cited alongside, same era.
Discourse-aware neural rewards for coherent text generation
Antoine Bosselut, Asli Çelikyilmaz, Xiaodong He, Jianfeng Gao, Po-Sen Huang, and Yejin Choi · 2018
Cited alongside, same era.
The price of debiasing automatic metrics in natural language evalaution
Arun Tejasvi Chaganty, Stephen Mussmann, and Percy Liang · 2018
Cited alongside, same era.
Fast policy learning through imitation and reinforcement
Ching-An Cheng, Xinyan Yan, Nolan Wagener, and Byron Boots · 2018
Cited alongside, same era.
Shisha Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Later among the works it cites.
Objective function learning to match human judgements for optimization-based summarization
Maxime Peyrard and Iryna Gurevych · 2018
Later among the works it cites.
Sentence relations for extractive summarization with deep neural networks
Pengjie Ren, Zhumin Chen, Zhaochun Ren, Furu Wei, Liqiang Nie, Jun Ma, and Maarten de Rijke · 2018
Later among the works it cites.
Learning to extract coherent summary via deep reinforcement learning
Yuxiang Wu and Baotian Hu · 2018
Later among the works it cites.
Deep reinforcement learning for extractive document summarization
Kaichun Yao, Libo Zhang, Tiejian Luo, and Yanjun Wu · 2018
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.