Fetching the paper…
Reading the bibliography…
LLMs are increasingly used to design reward functions based on human preferences in Reinforcement Learning (RL).
Restless bandits: Activity allocation in a changing world
P. Whittle · 1988
Earlier work this paper cites.
The complexity of optimal queueing network control
C. H. Papadimitriou and J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Fair division and collective welfare
H. Moulin · 2004
Earlier work this paper cites.
Deployed armor protection: the application of a game theoretic model for security at the los angeles international airport
J. Pita, M. Jain, J. Marecki, F. Ordóñez, C. Portway, M. Tambe, C. Western, P. Paruchuri, and S. Kraus · 2008
Earlier work this paper cites.
Iris-a tool for strategic security allocation in transportation networks
J. Tsai, S. Rathi, C. Kiekintveld, F. Ordonez, and M. Tambe · 2009
Earlier work this paper cites.
Handbook of social choice and welfare
K. J. Arrow, A. Sen, and K. Suzumura · 2010
Earlier work this paper cites.
Security and game theory: algorithms, deployed systems, lessons learned
M. Tambe · 2011
Earlier work this paper cites.
Hypervolume-based multi-objective reinforcement learning
K. V. Moffaert, M. M. Drugan, and A. Nowé · 2013
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
D. M. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley · 2013
Earlier work this paper cites.
Traveling towards disease: transportation barriers to health care access
S. T. Syed, B. S. Gerber, and L. K. Sharp · 2013
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
K. Van Moffaert and A. Nowé · 2014
Earlier work this paper cites.
Disparities in the use of a mhealth medication adherence promotion intervention for low-income adults with type 2 diabetes
L. A. Nelson, S. A. Mulvaney, T. Gebretsadik, Y.-X. Ho, K. B. Johnson, and C. Y. Osborn · 2016
Earlier work this paper cites.
Restless poachers: Handling exploration-exploitation tradeoffs in security domains
Y. Qian, C. Zhang, B. Krishnamachari, and M. Tambe · 2016
Earlier work this paper cites.
Restless poachers: Handling exploration-exploitation tradeoffs in security domains
Y. Qian, C. Zhang, B. Krishnamachari, and M. Tambe · 2016
Earlier work this paper cites.
A theory of justice
J. Rawls · 2017
Cited alongside, same era.
Strategies to improve treatment coverage in community-based public health programs: a systematic review of the literature
K. V. Deardorff, A. Rubin Means, K. H. Ásbjörnsdóttir, and J. Walson · 2018
Cited alongside, same era.
Group maintenance: A restless bandits approach
A. Abbou and V. Makis · 2019
Cited alongside, same era.
Prioritizing hepatitis c treatment in us prisons
T. Ayer, C. Zhang, A. Bonifonte, A. C. Spaulding, and J. Chhatwal · 2019
Cited alongside, same era.
Using natural language for reward shaping in reinforcement learning
P. Goyal, S. Niekum, and R. J. Mooney · 2019
Cited alongside, same era.
Learning fairness in multi-agent systems
J. Jiang and Z. Lu · 2019
Markovian restless bandits and index policies: A review
J. Niño-Mora · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao · 2023
Later among the works it cites.
Expanding impact of mobile health programs: SAHELI for maternal and child care
S. Verma, G. Singh, A. Mate, P. Verma, S. Gorantla, N. Madhiwalla, A. Hegde, D. Thakkar, M. Jain, M. Tambe, and A. Taneja · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
W. Yu, N. Gileadi, C. Fu, S. Kirmani, K. Lee, M. G. Arenas, H. L. Chiang, T. Erez, L. Hasenclever, J. Humplik, B. Ichter, T. Xiao, P. Xu, A. Zeng, T. Zhang, N. Heess, D. Sadigh, J. Tan, Y. Tassa, and F. Xia · 2023
Later among the works it cites.
Welfare and fairness in multi-objective reinforcement learning
F. Zimeng, P. Nianli, T. Muhang, and F. Brandon · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey of inverse reinforcement learning: Challenges, methods and progress
S. Arora and P. Doshi · 2021
Cited alongside, same era.
Ella: Exploration through learned language abstraction
S. Mirchandani, S. Karamcheti, and D. Sadigh · 2021
Cited alongside, same era.
Learning fair policies in decentralized cooperative multi-agent reinforcement learning
M. Zimmer, C. Glanois, U. Siddique, and P. Weng · 2021
Cited alongside, same era.
Eager: Asking and answering questions for automatic reward shaping in language-guided rl
T. Carta, P.-Y. Oudeyer, O. Sigaud, and S. Lamprier · 2022
Cited alongside, same era.
Field study in deploying restless multi-armed bandits: Assisting non-profits in improving maternal and child health
A. Mate, L. Madaan, A. Taneja, N. Madhiwalla, S. Verma, G. Singh, A. Hegde, P. Varakantham, and M. Tambe · 2022
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
R. Anil and et al · 2023
Cited alongside, same era.
ARMMAN · 2024
Closest in time.
A decision-language model (DLM) for dynamic restless multi-armed bandit tasks in public health
N. Behari, E. Zhang, Y. Zhao, A. Taneja, D. Mysore Nagaraj, and M. Tambe · 2024
Closest in time.
Survey on large language model-enhanced reinforcement learning: Concept, taxonomy, and methods
Y. Cao, H. Zhao, Y. Cheng, T. Shu, G. Liu, G. Liang, J. Zhao, and Y. Li · 2024
Closest in time.
Revolve: Reward evolution with large language models for autonomous driving
R. Hazra, A. Sygkounas, A. Persson, A. Loutfi, and P. Z. D. Martires · 2024
Closest in time.
Irl for restless multi-armed bandits with applications in maternal and child health
G. Jain, P. Varakantham, H. Xu, A. Taneja, P. Doshi, and M. Tambe · 2024
Closest in time.
Dreureka: Language model guided sim-to-real transfer
Y. J. Ma, W. Liang, H. Wang, S. Wang, Y. Zhu, L. Fan, O. Bastani, and D. Jayaraman · 2024
Closest in time.
Group fairness in predict-then-optimize settings for restless bandits
S. Verma, Y. Zhao, N. B. Sanket Shah, A. Taneja, and M. Tambe · 2024
Closest in time.
Text2reward: Reward shaping with language models for reinforcement learning
T. Xie, S. Zhao, C. H. Wu, Y. Liu, Q. Luo, V. Zhong, Y. Yang, and T. Yu · 2024
Closest in time.
Navigating the social welfare frontier: Portfolios for multi-objective reinforcement learning
C. W. Kim, J. Moondra, S. Verma, M. Pollack, L. Kong, M. Tambe, and S. Gupta · 2025
Closest in time.