Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms describe how an agent can learn an optimal action policy in a sequential decision process, through repeated experience.
Quant gans: deep generation of financial time series
M. Wiese, R. Knobloch, R. Korn, and P. Kretschmer · 1907
Earlier work this paper cites.
Animal Intelligence
E. L. Thorndike · 1911
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
W. R. Thompson · 1933
Earlier work this paper cites.
The behavior of organisms: An experimental analysis
B. F. Skinner · 1938
Earlier work this paper cites.
Cognitive maps in rats and men
E. C. Tolman · 1948
Earlier work this paper cites.
Iterative solutions of games by fictitious play
G. W. Brown · 1951
Earlier work this paper cites.
An iterative method of solving a game
J. Robinson · 1951
Earlier work this paper cites.
Some aspects of the sequential design of experiments
H. Robbins · 1952
Earlier work this paper cites.
The Travelling Salesman Problem
M. M. Flood · 1956
Earlier work this paper cites.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
A method for solving traveling-salesman problems
G. A. Croes · 1958
Earlier work this paper cites.
Dynamic Programming and Markov Processes
R. A. Howard · 1960
Earlier work this paper cites.
Steps toward artificial intelligence
M. Minsky · 1961
Earlier work this paper cites.
Some topics in two-person games
L. Shapley · 1964
Earlier work this paper cites.
Theories of bounded rationality
H. A. Simon · 1972
Earlier work this paper cites.
Sequential models in economic dynamics
M. F. Hellwig · 1973
Earlier work this paper cites.
Rational expectations and bayesian analysis
R. M. Cyert and M. H. DeGroot · 1974
Earlier work this paper cites.
Approximate algorithms for the traveling salesperson problem
D. J. Rosenkrantz, R. E. Stearns, and P. M. Lewis · 1974
Earlier work this paper cites.
A two-armed bandit theory of market pricing
M. Rothschild · 1974
Earlier work this paper cites.
Worst-case analysis of a new heuristic for the travelling salesman problem
N. Christofides · 1976
Earlier work this paper cites.
Animal learning and behavior theory
H. M. Jenkins · 1979
Earlier work this paper cites.
Aspects of the reinforcer learned in second-order Pavlovian conditioning
R. A. Rescorla · 1979
Earlier work this paper cites.
Optimal search for the best alternative
M. L. Weitzman · 1979
Earlier work this paper cites.
The nature of learning explanations
J. Garcia · 1981
Earlier work this paper cites.
Toward a modern theory of adaptive networks: Expectation and prediction
R. S. Sutton and A. G. Barto · 1981
Earlier work this paper cites.
Selection and the evolution of industry
B. Jovanovic · 1982
Earlier work this paper cites.
Optimization Over Time , volume 1
P. Whittle · 1983
Earlier work this paper cites.
Rationalizable strategic behavior
B. D. Bernheim · 1984
Earlier work this paper cites.
Price dispersion and incomplete learning in the long run
A. McLennan · 1984
Earlier work this paper cites.
Job matching and occupational choice
R. A. Miller · 1984
Earlier work this paper cites.
The Rate of Obsolescence of Patents, Research Gestation Lags, and the Private Rate of Return to Research Resources , pages 73–88
A. Pakes and M. Schankerman · 1984
Earlier work this paper cites.
Rationalizable strategic behavior and the problem of perfection
D. G. Pearce · 1984
Earlier work this paper cites.
An estimable dynamic stochastic model of fertility and child mortality
K. I. Wolpin · 1984
Earlier work this paper cites.
Bandits Problems Sequential Allocation of Experiments. — (Monographs on statistics and applied probability)
D. A. Berry and B. Fristedt · 1985
Earlier work this paper cites.
Minimal Rationality
C. Cherniak · 1986
Earlier work this paper cites.
Escaping brittleness: The possibilities of general-purpose learning algorithms applied to parallel rule-based systems
J. H. Holland · 1986
Earlier work this paper cites.
Patents as options: Some estimates of the value of holding european patent stocks
A. Pakes · 1986
Earlier work this paper cites.
Bayesian learning and convergence to rational expectations
M. Feldman · 1987
Earlier work this paper cites.
Optimal replacement of gmc bus engines: An empirical model of harold zurcher
J. Rust · 1987
Earlier work this paper cites.
Bandit Processes and Dynamic Allocation Indices
J. Gittins · 1989
Earlier work this paper cites.
On money as a medium of exchange
N. Kiyotaki and R. Wright · 1989
Earlier work this paper cites.
Recursive Methods in Economic Dynamics
N. L. Stokey, R. E. Lucas, and E. C. Prescott · 1989
Earlier work this paper cites.
Learning from delayed reward
C. J. Watkins · 1989
Earlier work this paper cites.
Designing economic agents that act like human agents: A behavioral approach to bounded rationality
W. B. Arthur · 1991
Earlier work this paper cites.
On the computational economics of reinforcement learning
A. G. Barto and S. P. Singh · 1991
Earlier work this paper cites.
q q -learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
On the gittins index for multiarmed bandits
R. Weber · 1992
Earlier work this paper cites.
Conditional choice probabilities and the estimation of dynamic models
V. J. Hotz and R. A. Miller · 1993
Earlier work this paper cites.
Bounded Rationality in Macroeconomics
T. Sargent · 1993
Earlier work this paper cites.
Inductive reasoning and bounded rationality
W. B. Arthur · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Cited alongside, same era.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Cited alongside, same era.
Markov-perfect industry dynamics: A framework for empirical work
R. Ericson and A. Pakes · 1995
Cited alongside, same era.
Ant-Q: A reinforcement learning approach to the traveling salesman problem
L. M. Gambardella and M. Dorigo · 1995
Cited alongside, same era.
Provably bounded-optimal agents
S. J. Russell and D. Subramanian · 1995
Cited alongside, same era.
Learning and strategic pricing
D. Bergemann and J. Välimäki · 1996
Cited alongside, same era.
Neuro-Dynamic Programming
Inverse reinforcement learning through structured classification
E. Klein, M. Geist, B. Piot, and O. Pietquin · 2012
Later among the works it cites.
Constrained optimization approaches to estimation of structural models
C.-L. Su and K. L. Judd · 2012
Later among the works it cites.
Equilibrium analysis of dynamic models of imperfect competition
J. F. Escobar · 2013
Later among the works it cites.
Recursive Models of Dynamic Linear Economies
L. P. Hansen and T. J. Sargent · 2013
Later among the works it cites.
A Sparsity-Based model of bounded rationality
X. Gabaix · 2014
Later among the works it cites.
Applying reinforcement learning to economic problems
N. Hughes · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Bertsekas and J. Tsitsiklis · 1996
Cited alongside, same era.
Ant colonies for the traveling salesman problem
M. Dorigo and L. M. Gambardella · 1996
Cited alongside, same era.
Reasoning the fast and frugal way: models of bounded rationality
G. Gigerenzer and D. Goldstein · 1996
Cited alongside, same era.
Learning from demonstration
S. Schaal · 1996
Cited alongside, same era.
Rationality and bounded rationality
R. J. Aumann · 1997
Cited alongside, same era.
Evolutionary games and equilibrium selection
L. Samuelson · 1997
Cited alongside, same era.
Online portfolio selection: A survey
B. Li and S. C. Hoi · 2014
Later among the works it cites.
Optimal real-time bidding for display advertising
W. Zhang, S. Yuan, and J. Wang · 2014
Later among the works it cites.
Computational rationality: A converging paradigm for intelligence in brains, minds, and machines
S. J. Gershman, E. J. Horvitz, and J. B. Tenenbaum · 2015
Later among the works it cites.
The econometrics of randomized experiments
S. Athey and G. W. Imbens · 2016
Later among the works it cites.
Neural combinatorial optimization with reinforcement learning
I. Bello, H. Pham, Q. V. Le, M. Norouzi, and S. Bengio · 2016
Later among the works it cites.
Deep direct reinforcement learning for financial signal representation and trading
Y. Deng, F. Bao, Y. Kong, Z. Ren, and Q. Dai · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
Sequential markets, market power, and arbitrage
K. Ito and M. Reguant · 2016
Later among the works it cites.
An adaptive portfolio trading system: A risk-return portfolio optimization using recurrent reinforcement learning with expected maximum drawdown
S. Almahdi and S. Y. Yang · 2017
Later among the works it cites.
Real-time bidding by reinforcement learning in display advertising
H. Cai, K. Ren, W. Zhang, K. Malialis, J. Wang, Y. Yu, and D. Guo · 2017
Later among the works it cites.
Learning combinatorial optimization algorithms over graphs
H. Dai, E. B. Khalil, Y. Zhang, B. Dilkina, and L. Song · 2017
Later among the works it cites.
Optimal auctions through deep learning, 2017
P. Dütting, Z. Feng, H. Narasimhan, D. C. Parkes, and S. S. Ravindranath · 2017
Later among the works it cites.
Optimal Transport Methods in Economics
A. Galichon · 2017
Later among the works it cites.
M. Igami · 2017
Later among the works it cites.
Machine learning: An applied econometric approach
S. Mullainathan and J. Spiess · 2017
Later among the works it cites.
Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations
E. Weinan, J. Han, and A. Jentzen · 2017
Later among the works it cites.
Econometrics and machine learning
A. Charpentier, E. Flachaire, and A. Ly · 2018
Later among the works it cites.
Learning heuristics for the tsp by policy gradient
M. Deudon, P. Cournut, A. Lacoste, Y. Adulyasak, and L.-M. Rousseau · 2018
Later among the works it cites.
Deep learning for revenue-optimal auctions with budgets
Z. Feng, H. Narasimhan, and D. C. Parkes · 2018
Later among the works it cites.
Recursive Macroeconomic Theory
L. Ljungqvist and T. J. Sargent · 2018
Later among the works it cites.
Actor-critic fictitious play in simultaneous move multistage games
J. Perolat, B. Piot, and O. Pietquin · 2018
Later among the works it cites.
Machine learning for dynamic discrete choice
V. Semenova · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis · 2018
Later among the works it cites.
Market making via reinforcement learning
T. Spooner, J. Fearnley, R. Savani, and A. Koukorinis · 2018
Later among the works it cites.
Deep reinforcement learning for sponsored search real-time bidding
J. Zhao, G. Qiu, Z. Guan, W. Zhao, and X. He · 2018
Later among the works it cites.
Concepts in Bounded Rationality: Perspectives from Reinforcement Learning
D. Abel · 2019
Later among the works it cites.
Machine learning methods that economists should know about
S. Athey and G. W. Imbens · 2019
Later among the works it cites.
B. Baldacci, I. Manziuk, T. Mastrolia, and M. Rosenbaum · 2019
Later among the works it cites.
Deep hedging
H. Buehler, L. Gonon, J. Teichmann, and B. Wood · 2019
Later among the works it cites.
Controlling an autonomous vehicle with deep reinforcement learning
A. Folkers, M. Rick, and C. Buskens · 2019
Later among the works it cites.
Risk management with machine-learning-based algorithms
S. Fécamp, J. Mikael, and X. Warin · 2019
Later among the works it cites.
Reinforcement learning for market making in a multi-agent dealer market
S. Ganesh, N. Vadori, M. Xu, H. Zheng, P. Reddy, and M. Veloso · 2019
Later among the works it cites.
Adaptive treatment assignment in experiments for policy choice
M. Kasy and A. Sautmann · 2019
Later among the works it cites.
Learning leads to bounded rationality and the evolution of cognitive bias in public goods games
O. Leimar and J. McNamara · 2019
Later among the works it cites.
Dynamic online pricing with incomplete information using multiarmed bandit experiments
K. Misra, E. M. Schwartz, and J. Abernethy · 2019
Later among the works it cites.
The seven tools of causal inference, with reflections on machine learning
J. Pearl · 2019
Later among the works it cites.
Algorithms, Machine Learning, and Collusion
U. Schwalbe · 2019
Later among the works it cites.
S. Vyetrenko and S. Xu · 2019
Later among the works it cites.
Continuous-time mean-variance portfolio optimization via reinforcement learning
H. Wang and X. Y. Zhou · 2019
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms, 2019
K. Zhang, Z. Yang, and T. Başar · 2019
Later among the works it cites.
On the convergence of model free learning in mean field games
R. Elie, J. Perolat, M. Laurière, M. Geist, and O. Pietquin · 2020
Closest in time.
Deep reinforcement learning for market making in corporate bonds: beating the curse of dimensionality
O. Guéant and I. Manziuk · 2020
Closest in time.
Deep reinforcement learning for autonomous driving: A survey
B. R. Kiran, I. Sobh, V. Talpaert, P. Mannion, A. A. A. Sallab, S. Yogamani, and P. Pérez · 2020
Closest in time.