Fetching the paper…
Reading the bibliography…
LLMs have demonstrated impressive performance across various language tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2019 · 1910
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H. 1989 · 1989
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. 1992 · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A.; and Niranjan, M. 1994 · 1994
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P.; Littman, M. L.; and Moore, A. W. 1996 · 1996
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S.; McAllester, D.; Singh, S.; and Mansour, Y. 1999 · 1999
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P.; Liu, X.; Gao, J.; and Chen, W. 2020 · 2006
Earlier work this paper cites.
Reinforcement learning and markov decision processes
Van Otterlo, M.; and Wiering, M. 2012 · 2012
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P.; Hunt, J. J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; and Wierstra, D. 2015 · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Earlier work this paper cites.
Zhu, Y.; Kiros, R.; Zemel, R.; Salakhutdinov, R.; Urtasun, R.; Torralba, A.; and Fidler, S. 2015 · 2015
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G.; Dabney, W.; and Munos, R. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Earlier work this paper cites.
Deep Dyna-Q: Integrating Planning for Task-Completion Dialogue Policy Learning
Peng, B.; Li, X.; Gao, J.; Liu, J.; and Wong, K.-F. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
When to trust your model: Model-based policy optimization
Janner, M.; Fu, J.; Zhang, M.; and Levine, S. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
BLEURT: Learning Robust Metrics for Text Generation
Sellam, T.; Das, D.; and Parikh, A. 2020 · 2020
Cited alongside, same era.
Program synthesis with large language models
Austin, J.; Odena, A.; Nye, M.; Bosma, M.; Michalewski, H.; Dohan, D.; Jiang, E.; Cai, C.; Terry, M.; Le, Q.; et al. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. d. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
Lu, K.; Yuan, H.; Lin, R.; Lin, J.; Yuan, Z.; Zhou, C.; and Zhou, J. 2023 · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z.; Xu, C.; Zhao, P.; Sun, Q.; Geng, X.; Hu, W.; Tao, C.; Ma, J.; Lin, Q.; and Jiang, D. 2023 · 2023
Later among the works it cites.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Mu, Y.; Zhang, Q.; Hu, M.; Wang, W.; Ding, M.; Jin, J.; Wang, B.; Dai, J.; Qiao, Y.; and Luo, P. 2023 · 2023
Later among the works it cites.
Peng, B.; Galley, M.; He, P.; Cheng, H.; Xie, Y.; Hu, Y.; Huang, Q.; Liden, L.; Yu, Z.; Chen, W.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2021 · 2021
Cited alongside, same era.
He, P.; Gao, J.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Yuan, W.; Neubig, G.; and Liu, P. 2021 · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022 · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Cited alongside, same era.
Merging models with fisher-weighted averaging
Matena, M. S.; and Raffel, C. A. 2022 · 2022
Cited alongside, same era.
Ul2: Unifying language learning paradigms
Tay, Y.; Dehghani, M.; Tran, V. Q.; Garcia, X.; Wei, J.; Wang, X.; Chung, H. W.; Bahri, D.; Schuster, T.; Zheng, S.; et al. 2022 · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2022 · 2022
Cited alongside, same era.
Prompting large language models with answer heuristics for knowledge-based visual question answering
Shao, Z.; Yu, Z.; Wang, M.; and Yu, J. 2023 · 2023
Later among the works it cites.
Adaptation Augmented Model-based Policy Optimization
Shen, J.; Lai, H.; Liu, M.; Zhao, H.; Yu, Y.; and Zhang, W. 2023 · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single GPU
Sheng, Y.; Zheng, L.; Yuan, B.; Li, Z.; Ryabinin, M.; Chen, B.; Liang, P.; Ré, C.; Stoica, I.; and Zhang, C. 2023 · 2023
Later among the works it cites.
Stablelm: Stability ai language models
Stability-AI. 2023 · 2023
Later among the works it cites.
Dobby: A Conversational Service Robot Driven by GPT-4
Stark, C.; Chun, B.; Charleston, C.; Ravi, V.; Pabon, L.; Sunkari, S.; Mohan, T.; Stone, P.; and Hart, J. 2023 · 2023
Later among the works it cites.
MOSS: Training Conversational Language Models from Synthetic Data
Sun, T.; Zhang, X.; He, Z.; Li, P.; Cheng, Q.; Yan, H.; Liu, X.; Shao, Y.; Tang, Q.; Zhao, X.; Chen, K.; Zheng, Y.; Zhou, Z.; Li, R.; Zhan, J.; Zhou, Y.; Li, L.; Yang, X.; Wu, L.; Yin, Z.; Huang, X.; and Qiu, X. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, ly usable llms
Team, M. N.; et al. 2023 · 2023
Later among the works it cites.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Xu, C.; Guo, D.; Duan, N.; and McAuley, J. 2023 · 2023
Later among the works it cites.
Solving math word problem with problem type classification
Yao, J.; Zhou, Z.; and Wang, Q. 2023 · 2023
Later among the works it cites.
Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch
Yu, L.; Yu, B.; Yu, H.; Huang, F.; and Li, Y. 2023 · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Later among the works it cites.
Reconcile: Round-table conference improves reasoning via consensus among diverse llms
Chen, J. C.-Y.; Saha, S.; and Bansal, M. 2024 · 2024
Closest in time.
K2: A foundation language model for geoscience knowledge understanding and utilization
Deng, C.; Zhang, T.; He, Z.; Chen, Q.; Shi, Y.; Xu, Y.; Fu, L.; Zhang, W.; Wang, X.; Zhou, C.; et al. 2024 · 2024
Closest in time.
Jiang, A. Q.; Sablayrolles, A.; Roux, A.; Mensch, A.; Savary, B.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Hanna, E. B.; Bressand, F.; et al. 2024 · 2024
Closest in time.
Better Zero-Shot Reasoning with Role-Play Prompting
Kong, A.; Zhao, S.; Chen, H.; Li, Q.; Qin, Y.; Sun, R.; Zhou, X.; Wang, E.; and Dong, X. 2024 · 2024
Closest in time.
Li, J.; Zhang, Q.; Yu, Y.; Fu, Q.; and Ye, D. 2024 · 2024
Closest in time.
Merging Multi-Task Models via Weight-Ensembling Mixture of Experts
Tang, A.; Shen, L.; Luo, Y.; Yin, N.; Zhang, L.; and Tao, D. 2024 · 2024
Closest in time.
Knowledge Fusion of Large Language Models
Wan, F.; Huang, X.; Cai, D.; Quan, X.; Bi, W.; and Shi, S. 2024 · 2024
Closest in time.
PMC-LLaMA: toward building open-source language models for medicine
Wu, C.; Lin, W.; Zhang, X.; Zhang, Y.; Xie, W.; and Wang, Y. 2024 · 2024
Closest in time.