Fetching the paper…
Reading the bibliography…
Large language models (LLMs) like ChatGPT and GPT-4 have attracted great attention given their surprising performance on a wide range of NLP tasks.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
A package for automatic evaluation of summaries
L. C. ROUGE · 2004
Earlier work this paper cites.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska · 2012
Earlier work this paper cites.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Learning-based single-document summarization with compression and anaphoricity constraints
G. Durrett, T. Berg-Kirkpatrick, and D. Klein · 2016
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
J. Li, W. Monroe, A. Ritter, D. Jurafsky, M. Galley, and J. Gao · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath · 2017
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Reinforcement learning for bandit neural machine translation with simulated human feedback
K. Nguyen, H. Daumé III, and J. Boyd-Graber · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
End-to-end offline goal-oriented dialog policy learning via policy gradient
L. Zhou, K. Small, O. Rokhlenko, and C. Elkan · 2017
Earlier work this paper cites.
Controllable abstractive summarization
A. Fan, D. Grangier, and M. Auli · 2018
Earlier work this paper cites.
Controlling length in abstractive summarization using a convolutional neural network
Y. Liu, Z. Luo, and K. Zhu · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
R. Paulus, C. Xiong, and R. Socher · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Cited alongside, same era.
Global optimization under length constraint for neural text summarization
T. Makino, T. Iwakura, H. Takamura, and M. Okumura · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Ernie: Enhanced representation through knowledge integration
Y. Sun, S. Wang, Y. Li, S. Feng, X. Chen, H. Zhang, X. Tian, D. Zhu, H. Tian, and H. Wu · 2019
Cited alongside, same era.
Positional encoding to control output sequence length
S. Takase and N. Okazaki · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Later among the works it cites.
Lenatten: An effective length controlling unit for text summarization
Z. Yu, Z. Wu, H. Zheng, Z. XuanYuan, J. Fong, and W. Su · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
A survey for in-context learning
Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, and Z. Sui · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Human-centric dialog training via offline reinforcement learning
N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2020
Cited alongside, same era.
Asking questions the human way: Scalable question-answer generation from text corpus
B. Liu, H. Wei, D. Niu, H. Chen, and Y. He · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
T. Goyal, J. J. Li, and G. Durrett · 2022
Later among the works it cites.
Length control in abstractive summarization by pretraining information selection
Y. Liu, Q. Jia, and K. Zhu · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Deep reinforcement learning: a survey
X. Wang, S. Wang, X. Liang, D. Zhao, J. Huang, X. Xu, B. Dai, and Q. Miao · 2022
Later among the works it cites.
Latent prompt tuning for text summarization
Y. Zhang, X. Zhang, X. Wang, S.-q. Chen, and F. Wei · 2022
Later among the works it cites.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2023
Closest in time.
Pangu- Σ \Sigma : Towards trillion parameter language model with sparse heterogeneous computing, 2023
X. Ren, P. Zhou, X. Meng, X. Huang, Y. Wang, W. Wang, P. Li, X. Zhang, A. Podolskiy, G. Arshinov, A. Bout, I. Piontkovskaya, J. Wei, X. Jiang, T. Su, Q. Liu, and J. Yao · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
A complete survey on generative ai (aigc): Is chatgpt from gpt-4 to gpt-5 all you need?
C. Zhang, C. Zhang, S. Zheng, Y. Qiao, C. Li, M. Zhang, S. K. Dam, C. M. Thwal, Y. L. Tun, L. L. Huy, et al · 2023
Closest in time.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.