Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have exhibited remarkable performance across various natural language processing (NLP) tasks.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
R. Nallapati, B. Zhou, C. N. dos Santos, Ç. Gülçehre, and B. Xiang · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Overview of the IWSLT 2017 evaluation campaign
M. Cettolo, M. Federico, L. Bentivogli, J. Niehues, S. Stüker, K. Sudoh, K. Yoshino, and C. Federmann · 2017
Earlier work this paper cites.
Reinforcement learning for bandit neural machine translation with simulated human feedback
K. Nguyen, H. Daumé III, and J. Boyd-Graber · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Towards coherent and cohesive long-form text generation
W. S. Cho, P. Zhang, Y. Zhang, X. Li, M. Galley, C. Brockett, M. Wang, and J. Gao · 2018
Earlier work this paper cites.
Learning to extract coherent summary via deep reinforcement learning
Y. Wu and B. Hu · 2018
Earlier work this paper cites.
Automatic adaptation of object detectors to new domains using self-training
A. RoyChowdhury, P. Chakrabarty, A. Singh, S. Jin, H. Jiang, L. Cao, and E. Learned-Miller · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. F. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, and N. R. et al · 2020
Cited alongside, same era.
Transformer reinforcement learning X
CarperAI · 2020
Cited alongside, same era.
Evaluation of text generation: A survey
A. Celikyilmaz, E. Clark, and J. Gao · 2020
Cited alongside, same era.
Revisiting self-training for neural sequence generation
J. He, J. Gu, J. Shen, and M. Ranzato · 2020
Cited alongside, same era.
Commongen: A constrained text generation challenge for generative commonsense reasoning
B. Y. Lin, M. Shen, W. Zhou, P. Zhou, C. Bhagavatula, Y. Choi, and X. Ren · 2020
Cited alongside, same era.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, and S. L. et al · 2022
Later among the works it cites.
Gpt-critic: Offline reinforcement learning for end-to-end task-oriented dialogue systems
Y. Jang, J. Lee, and K.-E. Kim · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Later among the works it cites.
Planning with large language models via corrective re-prompting
S. S. Raman, V. Cohen, E. Rosen, I. Idrees, D. Paulius, and S. Tellex · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, and A. R. et al · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bertscore: Evaluating text generation with BERT
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2020
Cited alongside, same era.
Semi-supervised semantic segmentation with cross pseudo supervision
X. Chen, Y. Yuan, G. Zeng, and J. Wang · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
Automated news summarization using transformers
A. Gupta, D. Chugh, Anjum, and R. Katarya · 2021
Cited alongside, same era.
Generate & rank: A multi-task framework for math word problems
J. Shen, Y. Yin, L. Li, L. Shang, X. Jiang, M. Zhang, and Q. Liu · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
W. Yuan, G. Neubig, and P. Liu · 2021
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2022
Later among the works it cites.
Large language models are reasoners with self-verification
Y. Weng, M. Zhu, S. He, K. Liu, and J. Zhao · 2022
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig · 2023
Closest in time.
Is prompt all you need? no. A comprehensive and broader view of instruction learning
R. Lou, K. Zhang, and W. Yin · 2023
Closest in time.
GPT-4 technical report
OpenAI · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, and D. Zhou · 2023
Closest in time.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.