Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have attracted great attention given their strong performance on a wide range of NLP tasks.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Earlier work this paper cites.
A package for automatic evaluation of summaries
L. C. ROUGE · 2004
Earlier work this paper cites.
A survey of actor-critic reinforcement learning: Standard and natural policy gradients
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska · 2012
Earlier work this paper cites.
Teaching machines to read and comprehend
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom · 2015
Earlier work this paper cites.
Learning-based single-document summarization with compression and anaphoricity constraints
G. Durrett, T. Berg-Kirkpatrick, and D. Klein · 2016
Earlier work this paper cites.
Deep reinforcement learning for dialogue generation
J. Li, W. Monroe, A. Ritter, D. Jurafsky, M. Galley, and J. Gao · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
D. Bahdanau, P. Brakel, K. Xu, A. Goyal, R. Lowe, J. Pineau, A. Courville, and Y. Bengio · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
End-to-end offline goal-oriented dialog policy learning via policy gradient
L. Zhou, K. Small, O. Rokhlenko, and C. Elkan · 2017
Earlier work this paper cites.
Controllable abstractive summarization
A. Fan, D. Grangier, and M. Auli · 2018
Earlier work this paper cites.
Controlling length in abstractive summarization using a convolutional neural network
Y. Liu, Z. Luo, and K. Zhu · 2018
Earlier work this paper cites.
A deep reinforced model for abstractive summarization
R. Paulus, C. Xiong, and R. Socher · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Global optimization under length constraint for neural text summarization
T. Makino, T. Iwakura, H. Takamura, and M. Okumura · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Positional encoding to control output sequence length
S. Takase and N. Okazaki · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Human-centric dialog training via offline reinforcement learning
N. Jaques, J. H. Shen, A. Ghandeharioun, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2020
Cited alongside, same era.
A survey for in-context learning
Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Xu, and Z. Sui · 2022
Later among the works it cites.
News summarization and evaluation in the era of gpt-3
T. Goyal, J. J. Li, and G. Durrett · 2022
Later among the works it cites.
Length control in abstractive summarization by pretraining information selection
Y. Liu, Q. Jia, and K. Zhu · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Deep reinforcement learning: a survey
X. Wang, S. Wang, X. Liang, D. Zhao, J. Huang, X. Xu, B. Dai, and Q. Miao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Liu, H. Wei, D. Niu, H. Chen, and Y. He · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Ernie 2.0: A continual pre-training framework for language understanding
Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu, and H. Wang · 2020
Cited alongside, same era.
Controlling dialogue generation with semantic exemplars
P. Gupta, J. P. Bigham, Y. Tsvetkov, and A. Pavel · 2021
Cited alongside, same era.
Conquest: Contextual question paraphrasing through answer-aware synthetic question generation
M. Mirshekari, J. Gu, and A. Sisto · 2021
Cited alongside, same era.
Text generation by learning from demonstrations
R. Y. Pang and H. He · 2021
Cited alongside, same era.
Y. Zhang, X. Zhang, X. Wang, S.-q. Chen, and F. Wei · 2022
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, et al · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Is reinforcement learning (not) for natural language processing?: Benchmarks, baselines, and building blocks for natural language policy optimization
R. Ramamurthy, P. Ammanabrolu, K. Brantley, J. Hessel, R. Sifa, C. Bauckhage, H. Hajishirzi, and Y. Choi · 2023
Later among the works it cites.
Pangu- Σ \Sigma : Towards trillion parameter language model with sparse heterogeneous computing, 2023
X. Ren, P. Zhou, X. Meng, X. Huang, Y. Wang, W. Wang, P. Li, X. Zhang, A. Podolskiy, G. Arshinov, A. Bout, I. Piontkovskaya, J. Wei, X. Jiang, T. Su, Q. Liu, and J. Yao · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Later among the works it cites.
A complete survey on generative ai (aigc): Is chatgpt from gpt-4 to gpt-5 all you need?
C. Zhang, C. Zhang, S. Zheng, Y. Qiao, C. Li, M. Zhang, S. K. Dam, C. M. Thwal, Y. L. Tun, L. L. Huy, et al · 2023
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Later among the works it cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tai, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. Le, and J. Wei · 2024
Closest in time.
Tinyllama: An open-source small language model
P. Zhang, G. Zeng, T. Wang, and W. Lu · 2024
Closest in time.