Fetching the paper…
Reading the bibliography…
In recent years some researchers have explored the use of reinforcement learning (RL) algorithms as key components in the solution of various natural language processing tasks.
Translation
W. Weaver · 1955
Earlier work this paper cites.
On Certain Formal Properties of Grammars
N. Chomsky · 1959
Earlier work this paper cites.
Aspects of the Theory of Syntax
N. Chomsky · 1965
Earlier work this paper cites.
Pseudogradient Adaptation and Training Algorithms
B. T. Poljak · 1973
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
A Statistical Approach to Machine Translation
P. F. Brown, J. Cocke, S. A. D. Pietra, V. J. D. Pietra, F. Jelinek, J. D. Lafferty, R. L. Mercer, and P. S. Roossin · 1990
Earlier work this paper cites.
An Introduction to Machine Translation
W. J. Hutchins and H. L. Somers · 1992
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
A Stochastic Model of Human-Machine Interaction for Learning Dialog Strategies
E. Levin, R. Pieraccini, and W. Eckert · 2000
Earlier work this paper cites.
Automatic Optimization of Dialogue Management
D. J. Litman, M. S. Kearns, S. P. Singh, and M. A. Walker · 2000
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Empirical Evaluation of a Reinforcement Learning Spoken Dialogue System
S. Singh, M. Kearns, D. J. Litman, and M. A. Walker · 2000
Earlier work this paper cites.
An Application of Reinforcement Learning to Dialogue Strategy Selection in a Spoken Dialogue System for Email
M. A. Walker · 2000
Earlier work this paper cites.
Probabilistic Methods in Spoken-Dialogue Systems
S. J. Young · 2000
Earlier work this paper cites.
Simulating the Evolution of Language
A. Cangelosi and D. Parisi, editors · 2002
Earlier work this paper cites.
Bleu: A Method for Automatic Evaluation of Machine Translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Optimizing Dialogue Management with Reinforcement Learning: Experiments with the NJFun System
S. P. Singh, D. Litman, M. Kearns, and M. Walker · 2002
Earlier work this paper cites.
Statistical Phrase-Based Translation
P. Koehn, F. J. Och, and D. Marcu · 2003
Earlier work this paper cites.
Minimum Error Rate Training in Statistical Machine Translation
F. J. Och · 2003
Earlier work this paper cites.
Simultaneous Translation of Lectures and Speeches
C. Fügen, A. Waibel, and M. Kolss · 2007
Earlier work this paper cites.
The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits
J. Langford and T. Zhang · 2007
Earlier work this paper cites.
Reinforcement Learning in Continuous Action Spaces
H. van Hasselt and M. A. Wiering · 2007
Earlier work this paper cites.
Partially Observable Markov Decision Processes for Spoken Dialog Systems
J. D. Williams and S. Young · 2007
Earlier work this paper cites.
Hybrid Reinforcement/Supervised Learning of Dialogue Policies from Fixed Data Sets
J. Henderson, O. Lemon, and K. Georgila · 2008
Earlier work this paper cites.
Dependency Parsing
S. Kübler, R. McDonald, and J. Nivre · 2008
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Search-Based Structured Prediction
H. Daumé III, J. Langford, and D. Marcu · 2009
Earlier work this paper cites.
Statistical Machine Translation
P. Koehn · 2009
Earlier work this paper cites.
Training Parsers by Inverse Reinforcement Learning
G. Neu and C. Szepesvári · 2009
Earlier work this paper cites.
The Hidden Agenda User Simulation Model
J. Schatzmann and S. Young · 2009
Earlier work this paper cites.
Dependency Parsing with Energy-Based Reinforcement Learning
L. Zhang and K. P. Chan · 2009
Earlier work this paper cites.
Natural Belief-Critic: A Reinforcement Algorithm for Parameter Estimation in Statistical Spoken Dialogue Systems
F. Jurcicek, B. Thomson, S. Keizer, F. Mairesse, M. Gasic, K. Yu, and S. J. Young · 2010
Earlier work this paper cites.
A Contextual-Bandit Approach to Personalized News Article Recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
S. Russell and P. Norvig · 2010
Earlier work this paper cites.
Bayesian Update of Dialogue State: A POMDP Framework for Spoken Dialogue Systems
B. Thomson and S. Young · 2010
Earlier work this paper cites.
Learning to Follow Navigational Directions
A. Vogel and D. Jurafsky · 2010
Earlier work this paper cites.
The Hidden Information State Model: A Practical Framework for POMDP-Based Spoken Dialogue Management
S. Young, M. Gašić, S. Keizer, F. Mairesse, J. Schatzmann, B. Thomson, and K. Yu · 2010
Earlier work this paper cites.
Combining Hierarchical Reinforcement Learning and Bayesian Networks for Natural Language Generation in Situated Dialogue
N. Dethlefs and H. Cuayáhuitl · 2011
Earlier work this paper cites.
Hierarchical Reinforcement Learning and Hidden Markov Models for Task-Oriented Natural Language Generation
N. Dethlefs and H. Cuayáhuitl · 2011
Earlier work this paper cites.
Learning What to Say and How to Say It: Joint Optimisation of Spoken Dialogue Management and Natural Language Generation
O. Lemon · 2011
Earlier work this paper cites.
Learning to Win by Reading Manuals in a Monte-Carlo Framework
S. R. K. Branavan, D. Silver, and R. Barzilay · 2012
Earlier work this paper cites.
Learned Prioritization for Trading Off Accuracy and Speed
J. Jiang, A. Teichert, J. Eisner, and H. Daumé III · 2012
Earlier work this paper cites.
Recurrent Continuous Translation Models
N. Kalchbrenner and P. Blunsom · 2013
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Introduction to the Theory of Computation
M. Sipser · 2013
Earlier work this paper cites.
POMDP-Based Statistical Spoken Dialog Systems: A Review
S. Young, M. Gašić, B. Thomson, and J. D. Williams · 2013
Earlier work this paper cites.
Learning Phrase Representations Using RNN Encoder–Decoder for Statistical Machine Translation
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Real User Evaluation of a POMDP Spoken Dialogue System Using Automatic Belief Compression
P. A. Crook, S. Keizer, Z. Wang, W. Tang, and O. Lemon · 2014
Earlier work this paper cites.
Nonstrict Hierarchical Reinforcement Learning for Interactive Systems and Robots
H. Cuayáhuitl, I. Kruijff-Korbayová, and N. Dethlefs · 2014
Earlier work this paper cites.
Fast and Robust Neural Network Joint Models for Statistical Machine Translation
J. Devlin, R. Zbib, Z. Huang, T. Lamar, R. Schwartz, and J. Makhoul · 2014
Earlier work this paper cites.
Generative Adversarial Nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Don’t Until the Final Verb Wait: Reinforcement Learning for Simultaneous Machine Translation
A. Grissom II, H. He, J. Boyd-Graber, J. Morgan, and H. Daumé III · 2014
Cited alongside, same era.
Distributed Representations of Sentences and Documents
Q. Le and T. Mikolov · 2014
Cited alongside, same era.
GloVe: Global Vectors for Word Representation
J. Pennington, R. Socher, and C. Manning · 2014
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Sequence to Sequence Learning with Neural Networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Cited alongside, same era.
SeqGAN: Sequence Generative Adversarial Nets with Policy Gradient
L. Yu, W. Zhang, J. Wang, and Y. Yu · 2017
Later among the works it cites.
Learning Conversational Systems that Interleave Task and Non-Task Content
Z. Yu, A. Rudnicky, and A. Black · 2017
Later among the works it cites.
D. Cer, Y. Yang, S.-y. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, Y.-H. Sung, B. Strope, and R. Kurzweil · 2018
Later among the works it cites.
Improving Interactive Reinforcement Learning: What Makes a Good Teacher?
F. Cruz, S. Magg, Y. Nagai, and S. Wermter · 2018
Later among the works it cites.
Multi-modal Feedback for Affordance-driven Interactive Reinforcement Learning
F. Cruz, G. I. Parisi, and S. Wermter · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural Machine Translation by Jointly Learning to Align and Translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Cited alongside, same era.
Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer · 2015
Cited alongside, same era.
Generating Text with Deep Reinforcement Learning
H. Guo · 2015
Cited alongside, same era.
Fatal or Not? Finding Errors That Lead to Dialogue Breakdowns in Chat-Oriented Dialogue Systems
R. Higashinaka, M. Mizukami, K. Funakoshi, M. Araki, H. Tsukahara, and Y. Kobayashi · 2015
Cited alongside, same era.
Skip-Thought Vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
Deep Learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Go for a Walk and Arrive at the Answer: Reasoning Over Paths in Knowledge Bases Using Reinforcement Learning
R. Das, S. Dhuliawala, M. Zaheer, L. Vilnis, I. Durugkar, A. Krishnamurthy, A. Smola, and A. McCallum · 2018
Later among the works it cites.
Neural Approaches to Conversational AI
J. Gao, M. Galley, and L. Li · 2018
Later among the works it cites.
Long Text Generation Via Adversarial Training with Leaked Information
J. Guo, S. Lu, H. Cai, W. Zhang, Y. Yu, and J. Wang · 2018
Later among the works it cites.
Achieving Human Parity on Automatic Chinese to English News Translation
H. Hassan, A. Aue, C. Chen, V. Chowdhary, J. Clark, C. Federmann, X. Huang, M. Junczys-Dowmunt, W. Lewis, M. Li, S. Liu, T.-Y. Liu, R. Luo, A. Menezes, T. Qin, F. Seide, X. Tan, F. Tian, L. Wu, S. Wu, Y. Xia, D. Zhang, Z. Zhang, and M. Zhou · 2018
Later among the works it cites.
Paraphrase Generation with Deep Reinforcement Learning
Z. Li, X. Jiang, L. Shang, and H. Li · 2018
Later among the works it cites.
Emergence of Grounded Compositional Language in Multi-Agent Populations
I. Mordatch and P. Abbeel · 2018
Later among the works it cites.
Grounding Language for Transfer in Deep Reinforcement Learning
K. Narasimhan, R. Barzilay, and T. Jaakkola · 2018
Later among the works it cites.
Deep Contextualized Word Representations
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Later among the works it cites.
Concatenated Power Mean Word Embeddings as Universal Cross-Lingual Sentence Representations
A. Rücklé, S. Eger, M. Peyrard, and I. Gurevych · 2018
Later among the works it cites.
A Survey of Available Corpora For Building Data-Driven Dialogue Systems: The Journal Version
I. V. Serban, R. Lowe, P. Henderson, L. Charlin, and J. Pineau · 2018
Later among the works it cites.
Toward Diverse Text Generation with Inverse Reinforcement Learning
Z. Shi, X. Chen, X. Qiu, and X. Huang · 2018
Later among the works it cites.
From Eliza to XiaoIce: Challenges and Opportunities with Social Chatbots
H.-y. Shum, X.-d. He, and D. Li · 2018
Later among the works it cites.
Reward Estimation for Dialogue Policy Optimisation
P.-H. Su, M. Gašić, and S. Young · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Quality Expectations of Machine Translation
A. Way · 2018
Later among the works it cites.
A Study of Reinforcement Learning for Neural Machine Translation
L. Wu, F. Tian, T. Qin, J. Lai, and T.-Y. Liu · 2018
Later among the works it cites.
A Bi-directional Multiple Timescales LSTM Model for Grounding of Actions and Verbs
A. Antunes, A. Laflaquiere, T. Ogata, and A. Cangelosi · 2019
Later among the works it cites.
Semantic Parsing with Dual Learning
R. Cao, S. Zhu, C. Liu, J. Li, and K. Yu · 2019
Later among the works it cites.
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Later among the works it cites.
From Semantics to Execution: Integrating Action Planning With Reinforcement Learning for Robotic Causal Problem-Solving
M. Eppe, P. D. H. Nguyen, and S. Wermter · 2019
Later among the works it cites.
Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation
Y. Gao, C. Meyer, M. Mesgar, and I. Gurevych · 2019
Later among the works it cites.
Deep Intrinsically Motivated Continuous Actor-Critic for Efficient Robotic Visuomotor Skill Learning
M. B. Hafez, C. Weber, M. Kerzel, and S. Wermter · 2019
Later among the works it cites.
Interactive-Predictive Neural Machine Translation Through Reinforcement and Imitation
T. K. Lam, S. Schamoni, and S. Riezler · 2019
Later among the works it cites.
Goal-Oriented Dialogue Policy Learning from Failures
K. Lu, S. Zhang, and X. Chen · 2019
Later among the works it cites.
A Survey of Reinforcement Learning Informed by Natural Language
J. Luketina, N. Nardelli, G. Farquhar, J. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel · 2019
Later among the works it cites.
Collaborative Multi-Agent Dialogue Model Training Via Reinforcement Learning
A. Papangelis, Y.-C. Wang, P. Molino, and G. Tur · 2019
Later among the works it cites.
Deep Reinforcement Learning for Modeling Chit-Chat Dialog with Discrete Attributes
C. Sankar and S. Ravi · 2019
Later among the works it cites.
Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models
T. Zhao, K. Xie, and M. Eskenazi · 2019
Later among the works it cites.
Language Models Are Few-Shot Learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Later among the works it cites.
Distributed Structured Actor-Critic Reinforcement Learning for Universal Dialogue Management
Z. Chen, L. Chen, X. Liu, and K. Yu · 2020
Later among the works it cites.
MQA: Answering the Question via Robotic Manipulation
Y. Deng, X. Guo, N. Zhang, D. Guo, H. Liu, and F. Sun · 2020
Later among the works it cites.
Improving Robot Dual-System Motor Learning with Intrinsically Motivated Meta-Control and Latent-Space Experience Imagination
M. B. Hafez, C. Weber, M. Kerzel, and S. Wermter · 2020
Later among the works it cites.
Crossmodal Language Grounding in an Embodied Neurocognitive Model
S. Heinrich, Y. Yao, T. Hinz, Z. Liu, T. Hummel, M. Kerzel, C. Weber, and S. Wermter · 2020
Later among the works it cites.
Deep Reinforcement Learning for Sequence-to-Sequence Models
Y. Keneshloo, T. Shi, N. Ramakrishnan, and C. K. Reddy · 2020
Later among the works it cites.
Document-Editing Assistants and Model-Based Reinforcement Learning as a Path to Conversational AI
K. Kudashkina, P. M. Pilarski, and R. S. Sutton · 2020
Later among the works it cites.
You Impress Me: Dialogue Generation Via Mutual Persona Perception
Q. Liu, Y. Chen, B. Chen, J.-G. Lou, Z. Chen, B. Zhou, and D. Zhang · 2020
Later among the works it cites.
Plato Dialogue System: A Flexible Conversational AI Research Platform
A. Papangelis, M. Namazifar, C. Khatri, Y.-C. Wang, P. Molino, and G. Tur · 2020
Later among the works it cites.
Curious Hierarchical Actor-Critic Reinforcement Learning
F. Röder, M. Eppe, P. D. H. Nguyen, and S. Wermter · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. Lillicrap, and D. Silver · 2020
Later among the works it cites.
Neural Machine Translation: A Review
F. Stahlberg · 2020
Later among the works it cites.
Towards Embodied Scene Description
S. Tan and H. Liu · 2020
Later among the works it cites.
Dual Learning for Semi-Supervised Natural Language Understanding
S. Zhu, R. Cao, and K. Yu · 2020
Later among the works it cites.
Generalization in Multimodal Language Learning from Simulation
A. Eisermann, J. H. Lee, C. Weber, and S. Wermter · 2021
Closest in time.
Improving Factual Consistency Between a Response and Persona Facts
M. Mesgar, E. Simpson, and I. Gurevych · 2021
Closest in time.
Multitask Learning and Reinforcement Learning for Personalized Dialog Generation: An Empirical Study
M. Yang, W. Huang, W. Tu, Q. Qu, Y. Shen, and K. Lei · 2021
Closest in time.