Fetching the paper…
Reading the bibliography…
User simulators can rapidly generate a large volume of timely user behavior data, providing a testing platform for reinforcement learning-based recommender systems, thus accelerating their iteration and optimization.
Recsim: A configurable simulation platform for recommender systems
Ie, E.; Hsu, C.-w.; Mladenov, M.; Jain, V.; Narvekar, S.; Wang, J.; Wu, R.; and Boutilier, C. 2019 · 1909
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Markov decision processes
Puterman, M. L. 1990 · 1990
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J. 2018 · 2018
Earlier work this paper cites.
Self-attentive sequential recommendation
Kang, W.-C.; and McAuley, J. 2018 · 2018
Earlier work this paper cites.
Rohde, D.; Bonner, S.; Dunlop, T.; Vasile, F.; and Karatzoglou, A. 2018 · 2018
Earlier work this paper cites.
Generative adversarial user model for reinforcement learning based recommendation system
Chen, X.; Li, S.; Li, H.; Jiang, S.; Qi, Y.; and Song, L. 2019 · 2019
Earlier work this paper cites.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning
Shi, J.-C.; Yu, Y.; Da, Q.; Chen, S.-Y.; and Zeng, A.-X. 2019 · 2019
Cited alongside, same era.
Deep reinforcement learning for search, recommendation, and online advertising: a survey
Zhao, X.; Xia, L.; Tang, J.; and Yin, D. 2019 · 2019
Cited alongside, same era.
Deep reinforcement learning for information retrieval: Fundamentals and advances
Zhang, W.; Zhao, X.; Zhao, L.; Yin, D.; Yang, G. H.; and Beutel, A. 2020 · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. 2021 · 2021
Cited alongside, same era.
Usersim: User simulation via supervised generativeadversarial network
Zhao, X.; Xia, L.; Zou, L.; Liu, H.; Yin, D.; and Tang, J. 2021b · 2021
Cited alongside, same era.
MILL: Mutual Verification with Large Language Models for Zero-Shot Query Expansion
Jia, P.; Liu, Y.; Zhao, X.; Li, X.; Hao, C.; Wang, S.; and Yin, D. 2023 · 2023
Later among the works it cites.
Automlp: Automated mlp for sequential recommendations
Li, M.; Zhang, Z.; Zhao, X.; Wang, W.; Zhao, M.; Wu, R.; and Guo, R. 2023 · 2023
Later among the works it cites.
Multi-task recommendations with reinforcement learning
Liu, Z.; Tian, J.; Cai, Q.; Zhao, X.; Gao, J.; Liu, S.; Chen, D.; He, T.; Zheng, D.; Jiang, P.; et al. 2023b · 2023
Later among the works it cites.
Llm lies: Hallucinations are not bugs, but features as adversarial examples
Yao, J.-Y.; Ning, K.-P.; Liu, Z.-H.; Ning, M.-N.; and Yuan, L. 2023 · 2023
Later among the works it cites.
Spatio-Temporal Digraph Convolutional Network-Based Taxi Pickup Location Recommendation
Zhang, Y.; Shen, G.; Han, X.; Wang, W.; and Kong, X. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023 · 2023
Cited alongside, same era.
A survey of chain of thought reasoning: Advances, frontiers and future
Chu, Z.; Chen, J.; Chen, Q.; Yu, W.; He, T.; Wang, H.; Peng, W.; Liu, M.; Qin, B.; and Liu, T. 2023 · 2023
Cited alongside, same era.
A unified framework for multi-domain ctr prediction via large language models
Fu, Z.; Li, X.; Wu, C.; Wang, Y.; Dong, K.; Zhao, X.; Zhao, M.; Guo, H.; and Tang, R. 2023 · 2023
Cited alongside, same era.
Towards mitigating LLM hallucination via self reflection
Ji, Z.; Yu, T.; Xu, Y.; Lee, N.; Ishii, E.; and Fung, P. 2023 · 2023
Cited alongside, same era.
Linrec: Linear attention mechanism for long-term sequential recommender systems
Liu, L.; Cai, L.; Zhang, C.; Zhao, X.; Gao, J.; Wang, W.; Lv, Y.; Fan, W.; Wang, Y.; He, M.; et al. 2023a
Cited in the paper.
Large language model distilling medication recommendation model
Liu, Q.; Wu, X.; Zhao, X.; Zhu, Y.; Zhang, Z.; Tian, F.; and Zheng, Y. 2024a
Cited in the paper.
User retention-oriented recommendation with decision transformer
Zhao, K.; Zou, L.; Zhao, X.; Wang, M.; and Yin, D. 2023b · 2023
Later among the works it cites.
SUBER: An RL Environment with Simulated Human Behavior for Recommender Systems
Corecco, N.; Piatti, G.; Lanzendörfer, L. A.; Fan, F. X.; and Wattenhofer, R. 2024 · 2024
Closest in time.
Adapting Job Recommendations to User Preference Drift with Behavioral-Semantic Fusion Learning
Han, X.; Zhu, C.; Hu, X.; Qin, C.; Zhao, X.; and Zhu, H. 2024 · 2024
Closest in time.
Large Language Model Interaction Simulator for Cold-Start Item Recommendation
Huang, F.; Yang, Z.; Jiang, J.; Bei, Y.; Zhang, Y.; and Chen, H. 2024 · 2024
Closest in time.
Large Multimodal Model Compression via Iterative Efficient Pruning and Distillation
Wang, M.; Zhao, Y.; Liu, J.; Chen, J.; Zhuang, C.; Gu, J.; Guo, R.; and Zhao, X. 2024b · 2024
Closest in time.