Fetching the paper…
Reading the bibliography…
We release and introduce the TigerBot family of large language models (LLMs), consisting of base and chat models, sized from 7, 13, 70 and 180 billion parameters.
Semantic parsing on freebase from question-answer pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
A. Talmor, J. Herzig, N. Lourie, and J. Berant · 2018
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. tau Yih · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. A. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia · 2021
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation
O. Press, N. A. Smith, and M. Lewis · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu · 2021
Cited alongside, same era.
Recursively summarizing books with human feedback
J. Wu, L. Ouyang, D. M. Ziegler, N. Stiennon, R. Lowe, J. Leike, and P. Christiano · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Cited alongside, same era.
Transformers
Huggingface · 2023
Closest in time.
Chatharuhi: Reviving anime character in reality via large language model
C. Li, Z. Leng, C. Yan, J. Shen, H. Wang, W. MI, Y. Fei, X. Feng, S. Yan, H. Wang, L. Zhan, Y. Jia, P. Wu, and H. Sun · 2023
Closest in time.
Megatron-deepspeed
Microsoft · 2023
Closest in time.
Tensorrt open source software
NVIDIA · 2023
Closest in time.
Meta completes research supercluster, announces next-gen datacenter
O. Peckham · 2023
Closest in time.
Yarn: Efficient context window extension of large language models
B. Peng, J. Quesnelle, H. Fan, and E. Shippole · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Yao, R. Y. Aminabadi, M. Zhang, X. Wu, C. Li, and Y. He · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai · 2023
Cited alongside, same era.
Claude 2
Anthropic · 2023
Cited alongside, same era.
Walking down the memory maze: Beyond context limit through interactive reading
H. Chen, R. Pasunuru, J. Weston, and A. Celikyilmaz · 2023
Cited alongside, same era.
Opencompass: A universal evaluation platform for foundation models
O. Contributors · 2023
Cited alongside, same era.
Efficient and effective text encoding for chinese llama and alpaca
Y. Cui, Z. Yang, and X. Yao · 2023
Cited alongside, same era.
Sentencepiece
Google · 2023
Cited alongside, same era.
Text generation inference
Huggingface · 2023
Cited alongside, same era.
An important next step on our ai journey
S. Pichai · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample · 2023
Closest in time.
Exllamav2
Turboderp · 2023
Closest in time.
Effective long-context scaling of foundation models
W. Xiong, J. Liu, I. Molybog, H. Zhang, P. Bhargava, R. Hou, L. Martin, R. Rungta, K. A. Sankararaman, B. Oguz, M. Khabsa, H. Fang, Y. Mehdad, S. Narang, K. Malik, A. Fan, S. Bhosale, S. Edunov, M. Lewis, S. Wang, and H. Ma · 2023
Closest in time.