Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have recently emerged as a focal point of research and application, driven by their unprecedented ability to understand and generate text with human-like quality.
“Language models are few-shot learners”
Tom Brown et al · 1901
Earlier work this paper cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, 2021
Alexey Dosovitskiy et al · 2010
Earlier work this paper cites.
“Microsoft COCO: Common objects in context”
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Is neural machine translation the new state of the art?”
Sheila Castilho et al · 2017
Earlier work this paper cites.
“Improving language understanding with unsupervised learning”, Technical report, OpenAI, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Think you have solved question answering? try arc, the AI2 reasoning challenge”
Peter Clark et al · 2018
Earlier work this paper cites.
“Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering”
Todor Mihaylov, Peter Clark, Tushar Khot and Ashish Sabharwal · 2018
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Generating long sequences with sparse transformers”
Rewon Child, Scott Gray, Alec Radford and Ilya Sutskever · 2019
Earlier work this paper cites.
“Hellaswag: Can a machine really finish your sentence?”
Rowan Zellers et al · 2019
Earlier work this paper cites.
“BoolQ: Exploring the surprising difficulty of natural yes/no questions”
Christopher Clark et al · 2019
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“The pile: An 800GB dataset of diverse text for language modeling”
Leo Gao et al · 2020
Earlier work this paper cites.
“RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models”
Samuel Gehman et al · 2020
Earlier work this paper cites.
“PIQA: Reasoning about physical commonsense in natural language”
Yonatan Bisk, Rowan Zellers, Jianfeng Gao and Yejin Choi · 2020
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks et al · 2020
Earlier work this paper cites.
“Review on Usage of Hidden Markov Model in Natural Language Processing”
Amrita Anandika, Smita Mishra and Madhusmita Das · 2021
Earlier work this paper cites.
“LoRA: Low-Rank Adaptation of Large Language Models”, 2021
Edward. Hu et al · 2021
Earlier work this paper cites.
“Learning transferable visual models from natural language supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“Prompt programming for large language models: Beyond the few-shot paradigm”
Laria Reynolds and Kyle McDonell · 2021
Earlier work this paper cites.
“Choosing Neural Networks over N-Gram Models for Natural Language Processing — towardsdatascience.com” [Accessed 22-Feb-2024], https://towardsdatascience.com/choosing-neural-networks-over-n-gram-models-for-natural-language-processing-156ea3a57fc , 2022
Benjamin McCloskey · 2022
Earlier work this paper cites.
“Scaling Language Models: Methods, Analysis & Insights from Training Gopher”, 2022
Jack. Rae et al · 2022
Earlier work this paper cites.
“Training Compute-Optimal Large Language Models”, 2022
Jordan Hoffmann et al · 2022
Earlier work this paper cites.
“What can transformers learn in-context? a case study of simple function classes”
Shivam Garg, Dimitris Tsipras, Percy Liang and Gregory Valiant · 2022
Earlier work this paper cites.
“A New Generation of Perspective API: Efficient Multilingual Character-level Transformers”
Alyssa Lees et al · 2022
Earlier work this paper cites.
“TruthfulQA: Measuring How Models Mimic Human Falsehoods”
Stephanie Lin, Jacob Hilton and Owain Evans · 2022
Earlier work this paper cites.
“ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope”
Partha Ray · 2023
Earlier work this paper cites.
“LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation”
Leigang Qu et al · 2023
Earlier work this paper cites.
“Information Retrieval meets Large Language Models: A strategic report from Chinese IR community”
Qingyao Ai et al · 2023
Earlier work this paper cites.
“Adaptive Machine Translation with Large Language Models”
Yasmin Moslem, Rejwanul Haque, John. Kelleher and Andy Way · 2023
Earlier work this paper cites.
“Gemini: a family of highly capable multimodal models”
Rohan Anil et al · 2023
Earlier work this paper cites.
“LLaMA: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Earlier work this paper cites.
“Large language model applications for evaluation: Opportunities and ethical implications”
Cari Head et al · 2023
Earlier work this paper cites.
“Ethical implications of large language models a multidimensional exploration of societal, economic, and technical concerns”
Kassym-Jomart Tokayev · 2023
Cited alongside, same era.
“From words to watts: Benchmarking the energy costs of large language model inference”
Siddharth Samsi et al · 2023
Cited alongside, same era.
“Risks and benefits of large language models for the environment”
Matthias Rillig et al · 2023
Cited alongside, same era.
“A brief history of language models — towardsdatascience.com” [Accessed 22-Feb-2024], https://towardsdatascience.com/a-brief-history-of-language-models-d9e4620e025b , 2023
Dorian Drost · 2023
Cited alongside, same era.
“A Survey of Text Representation and Embedding Techniques in NLP”
Rajvardhan Patil, Sorio Boit, Venkat Gudivada and Jagadeesh Nandigam · 2023
Cited alongside, same era.
“RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-Training”
Chen-Wei Xie et al · 2023
Later among the works it cites.
“Deep learning approaches on image captioning: A review”
Taraneh Ghandi, Hamidreza Pourreza and Hamidreza Mahyar · 2023
Later among the works it cites.
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Later among the works it cites.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
Wei-Lin Chiang et al · 2023
Later among the works it cites.
“Improved baselines with visual instruction tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Can Cui et al · 2023
Cited alongside, same era.
“Large language models (LLMs): A brief History, applications & challenges — blog.gopenai.com” [Accessed 22-Feb-2024], https://blog.gopenai.com/large-language-models-llms-a-brief-history-applications-challenges-c2fab10fa2e7 , 2023
Ambika · 2023
Cited alongside, same era.
“Attention Is All You Need”, 2023
Ashish Vaswani et al · 2023
Cited alongside, same era.
“Language is not all you need: Aligning perception with language models”
Shaohan Huang et al · 2023
Cited alongside, same era.
“A Short History Of ChatGPT: How We Got To Where We Are Today — forbes.com” [Accessed 22-Feb-2024], https://www.forbes.com/sites/bernardmarr/2023/05/19/a-short-history-of-chatgpt-how-we-got-to-where-we-are-today/ , 2023
Bernard Marr · 2023
Cited alongside, same era.
“GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”
Joshua Ainslie et al · 2023
Cited alongside, same era.
“What I learned from Bloomberg’s experience of building their own LLM — linkedin.com” [Accessed 21-Feb-2024], https://www.linkedin.com/pulse/what-i-learned-from-bloombergs-experience-building-own-chanen-phd/ , 2023
Ari Chanen · 2023
Cited alongside, same era.
Zhiliang Peng et al · 2023
Later among the works it cites.
“MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models”, 2023
Deyao Zhu et al · 2023
Later among the works it cites.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality” [Accessed 15-Feb-2024], https://lmsys.org/blog/2023-03-30-vicuna/ , 2023
Wei-Lin Chiang et al · 2023
Later among the works it cites.
“mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality”, 2023
Qinghao Ye et al · 2023
Later among the works it cites.
“Full Fine-Tuning, PEFT, Prompt Engineering, or RAG? — deci.ai” [Accessed 17-Feb-2024], https://deci.ai/blog/fine-tuning-peft-prompt-engineering-and-rag-which-one-is-right-for-you , 2023
Najeeb Nabwani · 2023
Later among the works it cites.
Lingling Xu et al · 2023
Later among the works it cites.
“QLoRA: Efficient Finetuning of Quantized LLMs”, 2023
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer · 2023
Later among the works it cites.
“Supervised Fine-tuning: customizing LLMs — medium.com” [Accessed 02-Mar-2024], https://medium.com/mantisnlp/supervised-fine-tuning-customizing-llms-a2c1edbf22c3 , 2023
Jose. Martinez · 2023
Later among the works it cites.
“The Ultimate Guide to LLM Fine Tuning: Best Practices & Tools — Lakera – Protecting AI teams that disrupt the world. — lakera.ai” [Accessed 02-Mar-2024], https://www.lakera.ai/blog/llm-fine-tuning-guide , 2023
Armin Norouzi · 2023
Later among the works it cites.
“Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing”
Pengfei Liu et al · 2023
Later among the works it cites.
“Prompt Prototype Learning Based on Ranking Instruction For Few-Shot Visual Tasks”
Li Sun, Liuan Wang, Jun Sun and Takayuki Okatani · 2023
Later among the works it cites.
“What is RLHF: Reinforcement Learning from Human Feedback — medium.com” [Accessed 02-Mar-2024], https://medium.com/generative-ai-insights-for-business-leaders-and/what-is-rlhf-reinforcement-learning-from-human-feedback-876da930bf16 , 2023
Now AI · 2023
Later among the works it cites.
“LLM Benchmarks: What Do They All Mean? — whytryai.com” [Accessed 02-Mar-2024], https://www.whytryai.com/p/llm-benchmarks , 2023
Daniel Nest · 2023
Later among the works it cites.
“Detecting and preventing hallucinations in large vision language models”
Anisha Gunjal, Jihan Yin and Erhan Bas · 2023
Later among the works it cites.
“LLM Benchmarks (Introduction to Benchmarks Techniques). — nageshmashette32” [Accessed 02-Mar-2024], https://medium.com/@nageshmashette32/llm-benchmarks-introduction-to-benchmarks-techniques-6518527620eb , 2023
Nagesh Mashette · 2023
Later among the works it cites.
“Harnessing the Power of LLMs in Practice: A Survey on ChatGPT and Beyond”
Jingfeng Yang et al · 2024
Closest in time.
“LLM-grounded Video Diffusion Models”
Long Lian et al · 2024
Closest in time.
“Generative AI: A systematic review using topic modelling techniques”
Priyanka Gupta, Bosheng Ding, Chong Guan and Ding Ding · 2024
Closest in time.
“Claude 2: a guide to Anthropic’s AI model and Chatbot” Accessed: 22-Feb-2024, https://www.zapier.com/blog/claude-ai/
2024
Closest in time.
“A Review on Large Language Models: Architectures, Applications, Taxonomies, Open Issues and Challenges”
Mohaimenul Raiaan et al · 2024
Closest in time.
“What Is Language Modeling? — Definition from TechTarget — techtarget.com” [Accessed 22-Feb-2024], https://www.techtarget.com/searchenterpriseai/definition/language-modeling
Nick Barney · 2024
Closest in time.
“Llama 2: The Next Revolution in AI Language Models - Complete 2024 Guide - viso.ai — viso.ai” [Accessed 07-Mar-2024], https://viso.ai/deep-learning/llama-2/
Gaudenz Boesch · 2024
Closest in time.
“Navigating the Attention Landscape: MHA, MQA, and GQA Decoded — iamshobhitagarwal.medium.com” [Accessed 23-Feb-2024], https://iamshobhitagarwal.medium.com/navigating-the-attention-landscape-mha-mqa-and-gqa-decoded-288217d0a7d1 , 2024
Shobhit Agarwal · 2024
Closest in time.
“Open source large language models: Benefits, risks and types - IBM Blog — ibm.com” [Accessed 22-Feb-2024], https://www.ibm.com/blog/open-source-large-language-models-benefits-risks-and-types/
IBM Data and AI Team · 2024
Closest in time.
“EU AI Act: first regulation on artificial intelligence” Accessed 15-Mar-2024, https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence
2024
Closest in time.
“Introducing Claude” Accessed 26-Mar-2024, https://www.anthropic.com/news/introducing-claude
Anthropic · 2024
Closest in time.
“Review: LLaMA: Open and Efficient Foundation Language Models — sh-tsang.medium.com” [Accessed 05-Mar-2024], https://sh-tsang.medium.com/review-llama-open-and-efficient-foundation-language-models-671d9284d523
Sik-Ho Tsang · 2024
Closest in time.
“What is Grouped Query Attention (GQA)?” [Accessed 07-Mar-2024], https://klu.ai/glossary/grouped-query-attention
Stephen. Walker · 2024
Closest in time.
“Common Crawl” Accessed: 22-Feb-2024, https://commoncrawl.org/
2024
Closest in time.