Fetching the paper…
Reading the bibliography…
Deep Generative AI has been a long-standing essential topic in the machine learning community, which can impact a number of application areas like text generation and computer vision.
1907
Earlier work this paper cites.
V. Mnih et al., “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning . PMLR, 2016, pp. 1928–1937
1937
Earlier work this paper cites.
R. Bellman et al., “The theory of dynamic programming,” Bulletin of the American Mathematical Society , vol. 60, no. 6, pp. 503–515, 1954
1954
Earlier work this paper cites.
R. S. Sutton et al., Temporal credit assignment in reinforcement learning . University of Massachusetts Amherst, 1984
1984
Earlier work this paper cites.
R. S. Sutton et al., “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine learning proceedings 1990 . Elsevier, 1990, pp. 216–224
1990
Earlier work this paper cites.
R. J. Williams et al., “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Reinforcement learning , pp. 5–32, 1992
1992
Earlier work this paper cites.
K. Papineni et al., “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004, pp. 74–81
2004
Earlier work this paper cites.
M. A. Carreira-Perpinan et al., “On contrastive divergence learning,” in International workshop on artificial intelligence and statistics . PMLR, 2005, pp. 33–40
2005
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
Y. LeCun, S. Chopra, R. Hadsell, M. Ranzato, F. Huang et al. , “A tutorial on energy-based learning,” Predicting structured data , vol. 1, no. 0, 2006
2006
Earlier work this paper cites.
E. Todorov et al., “Linearly-solvable markov decision problems,” Advances in neural information processing systems , vol. 19, 2006
2006
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning , 2008, pp. 160–167
2008
Earlier work this paper cites.
S. Ross et al., “A reduction of imitation learning and structured prediction to no-regret online learning,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 2011, pp. 627–635
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
D. J. Rezende, S. Mohamed, and D. Wierstra, “Stochastic backpropagation and approximate inference in deep generative models,” in International conference on machine learning . PMLR, 2014, pp. 1278–1286
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. Silver et al., “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, E. P. Xing and T. Jebara, Eds., vol. 32, no. 1. Bejing, China: PMLR, 22–24 Jun 2014, pp. 387–395. [Online]. Available: https://proceedings.mlr.press/v32/silver14.html
2014
Earlier work this paper cites.
D. Rezende and S. Mohamed, “Variational inference with normalizing flows,” in International conference on machine learning . PMLR, 2015, pp. 1530–1538
2015
Earlier work this paper cites.
V. et al. Mnih, “Human-level control through deep reinforcement learning,” Nat. , vol. 518, no. 7540, pp. 529–533, 2015. [Online]. Available: https://doi.org/10.1038/nature14236
2015
Earlier work this paper cites.
J. Schulman et al., “Trust region policy optimization,” in International conference on machine learning . PMLR, 2015, pp. 1889–1897
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 4566–4575
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Van Den Oord, N. Kalchbrenner, and K. Kavukcuoglu, “Pixel recurrent neural networks,” in International conference on machine learning . PMLR, 2016, pp. 1747–1756
2016
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
——, “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
M. Norouzi et al., “Reward augmented maximum likelihood for neural structured prediction,” Advances In Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
K. He et al., “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , June 2016
2016
Earlier work this paper cites.
G. Brockman et al., “Openai gym,” arXiv preprint arXiv:1606.01540 , 2016
2016
Earlier work this paper cites.
H. Chen et al., “A survey on dialogue systems: Recent advances and new frontiers,” Acm Sigkdd Explorations Newsletter , vol. 19, no. 2, pp. 25–35, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
T. Haarnoja et al., “Reinforcement learning with deep energy-based policies,” in International conference on machine learning . PMLR, 2017, pp. 1352–1361
2017
Earlier work this paper cites.
——, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347 , 2017
2017
Earlier work this paper cites.
——, “Mastering the game of go without human knowledge,” nature , vol. 550, no. 7676, pp. 354–359, 2017
2017
Earlier work this paper cites.
R. Devon Hjelm et al., “Boundary-seeking generative adversarial networks,” arXiv e-prints , pp. arXiv–1702, 2017
2017
Earlier work this paper cites.
L. Yu et al., “Seqgan: Sequence generative adversarial nets with policy gradient,” in Proceedings of the AAAI conference on artificial intelligence , vol. 31, no. 1, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Lin et al., “Adversarial ranking for language generation,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
N. Jaques et al., “Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control,” in International Conference on Machine Learning . PMLR, 2017, pp. 1645–1654
2017
Earlier work this paper cites.
J. Li et al., “Adversarial learning for neural dialogue generation,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Copenhagen, Denmark: Association for Computational Linguistics, September 2017, pp. 2157–2169. [Online]. Available: https://aclanthology.org/D17-1230
2017
Earlier work this paper cites.
S. J. Rennie et al., “Self-critical sequence training for image captioning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 7008–7024
2017
Earlier work this paper cites.
M. Popović, “chrf++: words helping character n-grams,” in Proceedings of the second conference on machine translation , 2017, pp. 612–618
2017
Earlier work this paper cites.
K. Nguyen et al., “Reinforcement learning for bandit neural machine translation with simulated human feedback,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Copenhagen, Denmark: Association for Computational Linguistics, September 2017, pp. 1464–1474. [Online]. Available: https://aclanthology.org/D17-1153
2017
Earlier work this paper cites.
B. et al. Zoph, “Neural architecture search with reinforcement learning,” in International Conference on Learning Representations , 2017. [Online]. Available: https://openreview.net/forum?id=r1Ue8Hcxg
2017
Earlier work this paper cites.
B. et al. Baker, “Designing neural network architectures using reinforcement learning,” in International Conference on Learning Representations , 2017. [Online]. Available: https://openreview.net/forum?id=S1c2cvqee
2017
Earlier work this paper cites.
X. Zhang et al., “Sentence simplification with deep reinforcement learning,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Copenhagen, Denmark: Association for Computational Linguistics, September 2017, pp. 584–594. [Online]. Available: https://aclanthology.org/D17-1062
2017
Earlier work this paper cites.
I. V. Serban et al., “A deep reinforcement learning chatbot,” arXiv preprint arXiv:1709.02349 , 2017
2017
Earlier work this paper cites.
M. Yang et al., “Personalized response generation via domain adaptation,” in Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2017, pp. 1021–1024
2017
Earlier work this paper cites.
J. D. Williams et al., “Hybrid code networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Vancouver, Canada: Association for Computational Linguistics, July 2017, pp. 665–677. [Online]. Available: https://aclanthology.org/P17-1062
2017
Earlier work this paper cites.
M. Lewis et al., “Deal or no deal? end-to-end learning of negotiation dialogues,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . Copenhagen, Denmark: Association for Computational Linguistics, September 2017, pp. 2443–2453. [Online]. Available: https://aclanthology.org/D17-1259
2017
Earlier work this paper cites.
X. Li et al., “End-to-end task-completion neural dialogue systems,” in Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Taipei, Taiwan: Asian Federation of Natural Language Processing, November 2017, pp. 733–743. [Online]. Available: https://aclanthology.org/I17-1074
2017
Earlier work this paper cites.
Z. Ren et al., “Deep reinforcement learning-based image captioning with embedding reward,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 290–298
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. De Vries et al., “Guesswhat?! visual object discovery through multi-modal dialogue,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5503–5512
2017
Earlier work this paper cites.
Z. Zhou et al., “Optimizing chemical reactions with deep reinforcement learning,” ACS central science , vol. 3, no. 12, pp. 1337–1344, 2017
2017
Earlier work this paper cites.
B. Sanchez-Lengeling et al., “Optimizing distributions over molecular space. an objective-reinforced generative adversarial network for inverse-design chemistry (organic),” 2017
2017
Earlier work this paper cites.
M. Olivecrona et al., “Molecular de-novo design through deep reinforcement learning,” Journal of cheminformatics , vol. 9, no. 1, pp. 1–14, 2017
2017
Earlier work this paper cites.
A. S. Vezhnevets et al., “Feudal networks for hierarchical reinforcement learning,” in International Conference on Machine Learning . PMLR, 2017, pp. 3540–3549
2017
Earlier work this paper cites.
C. Finn et al., “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning . PMLR, 2017, pp. 1126–1135
2017
Earlier work this paper cites.
R. S. Sutton et al., Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
T. Haarnoja et al., “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning . PMLR, 2018, pp. 1861–1870
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Wan et al., “Improving automatic source code summarization via deep reinforcement learning,” in Proceedings of the 33rd ACM/IEEE international conference on automated software engineering , 2018, pp. 397–407
2018
Earlier work this paper cites.
K. Mo et al., “Personalizing a dialogue system with transfer reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
Z. et al. Shi, “Toward diverse text generation with inverse reinforcement learning,” in Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, IJCAI 2018, July 13-19, 2018, Stockholm, Sweden , J. Lang, Ed. ijcai.org, 2018, pp. 4361–4367. [Online]. Available: https://doi.org/10.24963/ijcai.2018/606
2018
Earlier work this paper cites.
S. Narayan et al., “Ranking sentences for extractive summarization with reinforcement learning,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) . New Orleans, Louisiana: Association for Computational Linguistics, June 2018, pp. 1747–1759. [Online]. Available: https://aclanthology.org/N18-1158
2018
Earlier work this paper cites.
Y.-C. Chen et al., “Fast abstractive summarization with reinforce-selected sentence rewriting,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Melbourne, Australia: Association for Computational Linguistics, July 2018, pp. 675–686. [Online]. Available: https://aclanthology.org/P18-1063
2018
Earlier work this paper cites.
H. Pham et al., “Efficient neural architecture search via parameters sharing,” in International conference on machine learning . PMLR, 2018, pp. 4095–4104
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
N.-Q. Pham et al., “Towards one-shot learning for rare-word translation with external experts,” in Proceedings of the 2nd Workshop on Neural Machine Translation and Generation . Melbourne, Australia: Association for Computational Linguistics, July 2018, pp. 100–109. [Online]. Available: https://aclanthology.org/W18-2712
2018
Earlier work this paper cites.
Z. Zhong et al., “Practical block-wise neural network architecture generation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 2423–2432
2018
Earlier work this paper cites.
R. et al. Paulus, “A deep reinforced model for abstractive summarization,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=HkAClQgA-
2018
Earlier work this paper cites.
L. Wang et al., “A reinforced topic-aware convolutional sequence-to-sequence model for abstractive text summarization,” in Proceedings of the 27th International Joint Conference on Artificial Intelligence , ser. IJCAI’18. AAAI Press, 2018, p. 4453–4460
2018
Earlier work this paper cites.
Y. Wu et al., “Learning to extract coherent summary via deep reinforcement learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
L. Wu et al., “A study of reinforcement learning for neural machine translation,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . Brussels, Belgium: Association for Computational Linguistics, October-November 2018, pp. 3612–3621. [Online]. Available: https://aclanthology.org/D18-1397
2018
Earlier work this paper cites.
T. K. Lam et al., “A reinforcement learning approach to interactive-predictive neural machine translation,” in Proceedings of the 21st Annual Conference of the European Association for Machine Translation , Alicante, Spain, May 2018, pp. 189–198. [Online]. Available: https://aclanthology.org/2018.eamt-main.17
2018
Earlier work this paper cites.
H. Aghakhani et al., “Detecting deceptive reviews using generative adversarial networks,” in 2018 IEEE Security and Privacy Workshops (SPW) . IEEE, 2018, pp. 89–95
2018
Earlier work this paper cites.
Y. Li et al., “A generative model for category text generation,” Information Sciences , vol. 450, pp. 301–315, 2018
2018
Earlier work this paper cites.
J. Xu et al., “Diversity-promoting gan: A cross-entropy based generative adversarial network for diversified text generation,” in Proceedings of the 2018 conference on empirical methods in natural language processing , 2018, pp. 3940–3949
2018
Earlier work this paper cites.
D. Yarats et al., “Hierarchical text generation and planning for strategic dialogue,” in International Conference on Machine Learning . PMLR, 2018, pp. 5591–5599
2018
Earlier work this paper cites.
J. Zhang et al., “Goal-oriented visual question generation via intermediate rewards,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 186–201
2018
Earlier work this paper cites.
J. Zhang et al., “Sch-gan: Semi-supervised cross-modal hashing by generative adversarial network,” IEEE transactions on cybernetics , vol. 50, no. 2, pp. 489–502, 2018
2018
Earlier work this paper cites.
J. Guo et al., “Long text generation via adversarial training with leaked information,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
E. Putin et al., “Reinforced adversarial neural computer for de novo molecular design,” Journal of chemical information and modeling , vol. 58, no. 6, pp. 1194–1204, 2018
2018
Earlier work this paper cites.
M. Popova et al., “Deep reinforcement learning for de novo drug design,” Science advances , vol. 4, no. 7, p. eaap7885, 2018
2018
Earlier work this paper cites.
E. Putin et al., “Adversarial threshold neural computer for molecular de novo design,” Molecular pharmaceutics , vol. 15, no. 10, pp. 4386–4397, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Z. Hossain et al., “A comprehensive survey of deep learning for image captioning,” ACM Computing Surveys (CsUR) , vol. 51, no. 6, pp. 1–36, 2019
2019
Earlier work this paper cites.
Y. Jaafra et al., “Reinforcement learning for neural architecture search: A review,” Image and Vision Computing , vol. 89, pp. 57–66, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Sarmad et al., “Rl-gan-net: A reinforcement learning agent controlled gan network for real-time point cloud shape completion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Ke et al., “ARAML: A stable adversarial training framework for text generation,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . Hong Kong, China: Association for Computational Linguistics, November 2019, pp. 4271–4281. [Online]. Available: https://aclanthology.org/D19-1436
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Guo et al., “Irlas: Inverse reinforcement learning for architecture search,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 9021–9029
2019
Earlier work this paper cites.
L. Gui et al., “Neural topic model with reinforcement learning,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2019, pp. 3478–3483
2019
Cited alongside, same era.
Z. Yao et al., “Coacor: Code annotation for code retrieval with reinforcement learning,” in The world wide web conference , 2019, pp. 2203–2214
2019
Cited alongside, same era.
N. Brown et al., “Guacamol: benchmarking models for de novo molecular design,” Journal of chemical information and modeling , vol. 59, no. 3, pp. 1096–1108, 2019
2019
Cited alongside, same era.
D. Ha et al., “Reinforcement learning for improving agent design,” Artificial life , vol. 25, no. 4, pp. 352–365, 2019
2019
Cited alongside, same era.
J. Ho et al., “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Q. Wu et al., “Automatic math word problem generation with topic-expression co-attention mechanism and reinforcement learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1061–1072, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
I. Kobyzev, S. J. Prince, and M. A. Brubaker, “Normalizing flows: An introduction and review of current methods,” IEEE transactions on pattern analysis and machine intelligence , vol. 43, no. 11, pp. 3964–3979, 2020
2020
Cited alongside, same era.
H. Fan et al., “Recurrent attention network with reinforced generator for visual dialog,” ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM) , vol. 16, no. 3, pp. 1–16, 2020
2020
Cited alongside, same era.
Q. Wang et al., “Grl: Knowledge graph completion with gan-based reinforcement learning,” Knowledge-Based Systems , vol. 209, p. 106421, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. Scialom et al., “Coldgans: Taming language gans with cautious sampling strategies,” Advances in Neural Information Processing Systems , vol. 33, pp. 18 978–18 989, 2020
2020
Cited alongside, same era.
N. Stiennon et al., “Learning to summarize with human feedback,” Advances in Neural Information Processing Systems , vol. 33, pp. 3008–3021, 2020
2020
Cited alongside, same era.
Y. Gao et al., “Preference-based interactive multi-document summarisation,” Information Retrieval Journal , vol. 23, pp. 555–585, 2020
2020
Cited alongside, same era.
2022
Later among the works it cites.
C. Wang et al., “Enriching query semantics for code search with reinforcement learning,” Neural Networks , vol. 145, pp. 22–32, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
L. Zhang et al., “Learnedsqlgen: Constraint-aware sql generation using reinforcement learning,” in Proceedings of the 2022 International Conference on Management of Data , ser. SIGMOD ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 945–958. [Online]. Available: https://doi.org/10.1145/3514221.3526155
2022
Later among the works it cites.
H. Le et al., “Coderl: Mastering code generation through pretrained models and deep reinforcement learning,” Advances in Neural Information Processing Systems , vol. 35, pp. 21 314–21 328, 2022
2022
Later among the works it cites.
A. Ostonov et al., “Rlss: A deep reinforcement learning algorithm for sequential scene generation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2022, pp. 2219–2228
2022
Later among the works it cites.
R. Zhang et al., “Qinet: Decision surface learning and adversarial enhancement for quasi-immune completion of diverse corrupted point clouds,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–14, 2022
2022
Later among the works it cites.
L. Siyao, W. Yu, T. Gu, C. Lin, Q. Wang, C. Qian, C. C. Loy, and Z. Liu, “Bailando: 3d dance generation by actor-critic gpt with choreographic memory,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 11 050–11 059
2022
Later among the works it cites.
D. Deng, A. Wu, H. Qu, and Y. Wu, “Dashbot: Insight-driven dashboard generation based on deep reinforcement learning,” IEEE Transactions on Visualization and Computer Graphics , vol. 29, no. 1, pp. 690–700, 2022
2022
Later among the works it cites.
T. Liu, Q. Meng, J.-J. Huang, A. Vlontzos, D. Rueckert, and B. Kainz, “Video summarization through reinforcement learning with a 3d spatio-temporal u-net,” IEEE transactions on image processing , vol. 31, pp. 1573–1586, 2022
2022
Later among the works it cites.
J. Gibson et al., “A reinforcement learning approach to speech coding,” Information , vol. 13, no. 7, p. 331, 2022
2022
Later among the works it cites.
M. Thomas et al., “Augmented hill-climb increases reinforcement learning efficiency for language-based de novo molecule generation,” Journal of Cheminformatics , vol. 14, no. 1, pp. 1–22, 2022
2022
Later among the works it cites.
R. Ishitani et al., “Molecular design method using a reversible tree representation of chemical compounds and deep reinforcement learning,” Journal of Chemical Information and Modeling , vol. 62, no. 17, pp. 4032–4048, 2022
2022
Later among the works it cites.
T. Fu et al., “Reinforced genetic algorithm for structure-based drug design,” Advances in Neural Information Processing Systems , vol. 35, pp. 12 325–12 338, 2022
2022
Later among the works it cites.
M. Sun et al., “Molsearch: Search-based multi-objective molecular generation and property optimization,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , 2022, pp. 4724–4732
2022
Later among the works it cites.
P. C. Nguyen, N. N. Vlassis, B. Bahmani, W. Sun, H. Udaykumar, and S. S. Baek, “Synthesizing controlled microstructures of porous media using generative adversarial networks and reinforcement learning,” Scientific reports , vol. 12, no. 1, p. 9034, 2022
2022
Later among the works it cites.
S. R. Atance et al., “De novo drug design using reinforcement learning with graph-based deep generative models,” Journal of Chemical Information and Modeling , vol. 62, no. 20, pp. 4863–4872, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Guo et al., “Efficient (soft) q-learning for text generation with limited good data,” in Findings of the Association for Computational Linguistics: EMNLP 2022 , 2022, pp. 6969–6991
2022
Later among the works it cites.
M. Korshunova et al., “Generative and reinforcement learning approaches for the automated de novo design of bioactive compounds,” Communications Chemistry , vol. 5, no. 1, p. 129, 2022
2022
Later among the works it cites.
“Chatgpt,” https://openai.com/chatgpt
2023
Closest in time.
Z. Lin et al., “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science , vol. 379, no. 6637, pp. 1123–1130, 2023
2023
Closest in time.
M. Mohebbi Moghaddam et al., “Games of gans: Game-theoretical models for generative adversarial networks,” Artificial Intelligence Review , pp. 1–37, 2023
2023
Closest in time.
V. Uc-Cetina et al., “Survey on reinforcement learning for language processing,” Artificial Intelligence Review , vol. 56, no. 2, pp. 1543–1575, 2023
2023
Closest in time.
J. Ni et al., “Recent advances in deep learning based dialogue systems: A systematic survey,” Artificial intelligence review , vol. 56, no. 4, pp. 3055–3155, 2023
2023
Closest in time.
2023
Closest in time.
J. C. Fromer et al., “Computer-aided multi-objective optimization in small molecule discovery,” Patterns , vol. 4, no. 2, 2023
2023
Closest in time.
S. Luukkonen et al., “Artificial intelligence in multi-objective drug design,” Current Opinion in Structural Biology , vol. 79, p. 102537, 2023
2023
Closest in time.
Anna, “Exclusive: Chatgpt traffic slips again for third month in a row,” The Reuters . [Online]. Available: https://www.reuters.com/technology/chatgpt-traffic-slips-again-third-month-row-2023-09-07/
2023
Closest in time.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” Proceedings of Machine Learning and Systems , vol. 5, pp. 606–624, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Srivastava et al., “Response-act guided reinforced dialogue generation for mental health counseling,” in Proceedings of the ACM Web Conference 2023 , 2023, pp. 1118–1129
2023
Closest in time.
Z. Wu, Y. Hu, W. Shi, N. Dziri, A. Suhr, P. Ammanabrolu, N. A. Smith, M. Ostendorf, and H. Hajishirzi, “Fine-grained human feedback gives better rewards for language model training,” Advances in Neural Information Processing Systems , vol. 36, pp. 59 008–59 033, 2023
2023
Closest in time.
Z. Li, T. Xu, Y. Zhang, Z. Lin, Y. Yu, R. Sun, and Z.-Q. Luo, “Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models,” in Forty-first International Conference on Machine Learning , 2023
2023
Closest in time.
R. et al. Rafailov, “Direct preference optimization: Your language model is secretly a reward model,” 2023
2023
Closest in time.
2023
Closest in time.
T. Korbak, K. Shi, A. Chen, R. V. Bhalerao, C. Buckley, J. Phang, S. R. Bowman, and E. Perez, “Pretraining language models with human preferences,” in International Conference on Machine Learning . PMLR, 2023, pp. 17 506–17 533
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Havrilla, M. Zhuravinskyi, D. Phung, A. Tiwari, J. Tow, S. Biderman, Q. Anthony, and L. Castricato, “trlx: A framework for large scale reinforcement learning from human feedback,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 8578–8595
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Dessì, M. Bevilacqua, E. Gualdoni, N. C. Rakotonirina, F. Franzon, and M. Baroni, “Cross-domain image captioning with discriminative finetuning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6935–6944
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Z. Zhang et al., “Point cloud scene completion with joint color and semantic estimation from single rgb-d image,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Closest in time.
K. Zhao, Y. Zhang, S. Wang, T. Beeler, and S. Tang, “Synthesizing diverse human motions in 3d indoor scenes,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 14 738–14 749
2023
Closest in time.
Y. Yu, J. Chung, H. Yun, J. Hessel, J. S. Park, X. Lu, R. Zellers, P. Ammanabrolu, R. Le Bras, G. Kim et al. , “Fusing pre-trained language models with multimodal prompts through reinforcement learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 10 845–10 856
2023
Closest in time.
P. Hu et al., “De novo drug design based on stack-rnn with multi-objective reward-weighted sum and reinforcement learning,” Journal of Molecular Modeling , vol. 29, no. 4, pp. 1–12, 2023
2023
Closest in time.
X. Liu et al., “Drugex v3: scaffold-constrained drug design with graph transformer-based reinforcement learning,” Journal of Cheminformatics , vol. 15, no. 1, p. 24, 2023
2023
Closest in time.
A. A. Volk, R. W. Epps, D. T. Yonemoto, B. S. Masters, F. N. Castellano, K. G. Reyes, and M. Abolhasani, “Alphaflow: autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by reinforcement learning,” Nature Communications , vol. 14, no. 1, p. 1403, 2023
2023
Closest in time.
S. et al. Wang, “Causal decision transformer for recommender systems via offline reinforcement learning,” 2023
2023
Closest in time.
W. et al. Huang, “Voxposer: Composable 3d value maps for robotic manipulation with language models,” 2023
2023
Closest in time.
G. et al. Wang, “Voyager: An open-ended embodied agent with large language models,” 2023
2023
Closest in time.
U. Honda et al., “Switching to discriminative image captioning by relieving a bottleneck of reinforcement learning,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 1124–1134
2023
Closest in time.
T. Zhao et al., “A multi-scenario text generation method based on meta reinforcement learning,” Pattern Recognition Letters , vol. 165, pp. 47–54, 2023
2023
Closest in time.
H. et al. Liu, “Evaluating the logical reasoning ability of chatgpt and gpt-4,” 2023
2023
Closest in time.
G. Franceschelli and M. Musolesi, “Reinforcement learning for generative ai: State of the art, opportunities and open research challenges,” Journal of Artificial Intelligence Research , vol. 79, pp. 417–446, 2024
2024
Closest in time.
I. MidJourney, “Midjourney,” https://www.midjourney.com
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Li, H. Zhang, F. Zhang, T.-W. Chang, K. Kuang, L. Chen, and J. Zhou, “Optimizing language models with fair and stable reward composition in reinforcement learning,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 10 122–10 140
2024
Closest in time.
2024
Closest in time.
M. Wulfmeier, M. Bloesch, N. Vieillard, A. Ahuja, J. Bornschein, S. Huang, A. Sokolov, M. Barnes, G. Desjardins, A. Bewley et al. , “Imitating language via scalable inverse reinforcement learning,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems , 2024
2024
Closest in time.
2024
Closest in time.
T. R. McIntosh, T. Susnjak, T. Liu, P. Watters, and M. N. Halgamuge, “The inadequacy of reinforcement learning from human feedback-radicalizing large language models via semantic vulnerabilities,” IEEE Transactions on Cognitive and Developmental Systems , 2024
2024
Closest in time.
J. Wang, J. Wu, M. Chen, Y. Vorobeychik, and C. Xiao, “Rlhfpoison: Reward poisoning attack for reinforcement learning with human feedback in large language models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 2551–2570
2024
Closest in time.
X. Wang, J. Peng, K. Xu, H. Yao, and T. Chen, “Reinforcement learning-driven llm agent for automated attacks on llms,” in Proceedings of the Fifth Workshop on Privacy in Natural Language Processing , 2024, pp. 170–177
2024
Closest in time.
C. Wang, H. Zhou, Y. Hu, Y. Huo, B. Li, T. Liu, T. Xiao, and J. Zhu, “Esrl: Efficient sampling-based reinforcement learning for sequence generation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 19 107–19 115
2024
Closest in time.
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: Language agents with verbal reinforcement learning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Dai, Z. Zhu, H. Hu, G. Tang, L. Liu, and S. Xu, “Enhancing e-commerce query rewriting: A large language model approach with domain-specific pre-training and reinforcement learning,” in Proceedings of the 33rd ACM International Conference on Information and Knowledge Management , 2024, pp. 4439–4445
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Khan, V. K. BG, S. Schulter, Y. Fu, and M. Chandraker, “Self-training large language models for improved visual program synthesis with visual reinforcement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 344–14 353
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhao, W. S. Lee, and D. Hsu, “Large language models as commonsense knowledge for large-scale task planning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.