Fetching the paper…
Reading the bibliography…
The widespread adoption of Large Language Models (LLMs) has become commonplace, particularly with the emergence of open-source models.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Sid Black, Gao Leo, Phil Wang, et al. 2021 · 2021
Earlier work this paper cites.
How Can We Know When Language Models Know? On the Calibration of Language Models for Question Answering
Zhengbao Jiang, Jun Araki, Haibo Ding, et al. 2021 · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Ben Wang and Aran Komatsuzaki. 2021 · 2021
Earlier work this paper cites.
Massively Multilingual Natural Language Understanding 2022 (MMNLU-22) Workshop and Competition
Jack FitzGerald, Christopher Hench, Charith Peris, et al. 2022 · 2022
Earlier work this paper cites.
OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization
Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, et al. 2022 · 2022
Earlier work this paper cites.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, et al. 2023 · 2023
Earlier work this paper cites.
Performance of ChatGPT on the MCAT: The Road to Personalized and Equitable Premedical Learning
Vikas L Bommineni, Sanaea Bhagwagar, Daniel Balcarcel, et al. 2023 · 2023
Earlier work this paper cites.
Comparing ChatGPT and GPT-4 performance in usmle soft skill assessments
Dana Brin, Vera Sorin, Akhil Vaid, et al. 2023 · 2023
Earlier work this paper cites.
h2oGPT: Democratizing Large Language Models
Arno Candel, Jon McKinney, Philipp Singer, et al. 2023 · 2023
Earlier work this paper cites.
Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM
Mike Conover, Matt Hayes, Ankit Mathur, et al. 2023 · 2023
Earlier work this paper cites.
From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
Shangbin Feng, Chan Young Park, Yuhan Liu, et al. 2023 · 2023
Earlier work this paper cites.
What’s going on with the Open LLM Leaderboard?
Clémentine Fourrier, Nathan Habib, Julien Launay, , et al. 2023 · 2023
Earlier work this paper cites.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, et al. 2023 · 2023
Earlier work this paper cites.
OpenLLaMA: An Open Reproduction of LLaMA
Xinyang Geng and Hao Liu. 2023 · 2023
Earlier work this paper cites.
llama.cpp
Georgi Gerganov. 2023 · 2023
Earlier work this paper cites.
The rise of small language models - efficient & customizable
Bijit Ghosh. 2023 · 2023
Cited alongside, same era.
Akshat Gupta, Xiaoyang Song, and Gopala Anumanchipalli. 2023 · 2023
Cited alongside, same era.
ChatGPT an ENFJ, Bard an ISTJ: Empirical Study on Personalities of Large Language Models
Jen Tse Huang, Wenxuan Wang, Man Ho Lam, et al. 2023 · 2023
Cited alongside, same era.
Huggingfaceh4/zephyr-7b-alpha
HuggingFaceH4. 2023 · 2023
Cited alongside, same era.
Reliability Check: An Analysis of GPT-3’s Response to Sensitive Topics and Prompt Wording
Aisha Khatun and Daniel Brown. 2023 · 2023
Cited alongside, same era.
Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions
Pouya Pezeshkpour and Estevam Hruschka. 2023 · 2023
Later among the works it cites.
Pygmalionai/pygmalion-6b
PygmalionAI. 2023a · 2023
Later among the works it cites.
Pygmalionai/pygmalion-7b
PygmalionAI. 2023b · 2023
Later among the works it cites.
Leveraging Large Language Models for Multiple Choice Question Answering
Joshua Robinson and David Wingate. 2023 · 2023
Later among the works it cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, et al. 2023 · 2023
Later among the works it cites.
RedPajama 7B now available, instruct model outperforms all open 7B models on HELM benchmarks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinqi Lai, Wensheng Gan, Jiayang Wu, et al. 2023 · 2023
Cited alongside, same era.
MistralOrca: Mistral-7B Model Instruct-tuned on Filtered OpenOrcaV1 GPT-4 Dataset
Wing Lian, Bleys Goodson, Guan Wang, et al. 2023 · 2023
Cited alongside, same era.
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Haipeng Luo, Qingfeng Sun, Can Xu, et al. 2023 · 2023
Cited alongside, same era.
MLC LLM
MLC LLM. 2023 · 2023
Cited alongside, same era.
Introducing MPT-7B: A New Standard for Open-Source, Commercially Usable LLMs
MosaicML NLP Team. 2023 · 2023
Cited alongside, same era.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, et al. 2023 · 2023
Cited alongside, same era.
Orca: Progressive Learning from Complex Explanation Traces of GPT-4
Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, et al. 2023 · 2023
Cited alongside, same era.
Together. 2023 · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, et al. 2023 · 2023
Later among the works it cites.
Increasing probability mass on answer choices does not always improve accuracy
Sarah Wiegreffe, Matthew Finlayson, Oyvind Tafjord, et al. 2023 · 2023
Later among the works it cites.
Sean Wu, Michael Koo, Lesley Blum, et al. 2023 · 2023
Later among the works it cites.
Mini-Giants: “Small” Language Models and Open Source Win-Win
Zhengping Zhou, Lezhi Li, Xinxi Chen, et al. 2023 · 2023
Later among the works it cites.
Questioning the survey responses of large language models
Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-Dünner. 2024 · 2024
Closest in time.
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
Aisha Khatun and Daniel G. Brown. 2024 · 2024
Closest in time.
Beyond probabilities: Unveiling the misalignment in evaluating large language models
Chenyang Lyu, Minghao Wu, and Alham Aji. 2024 · 2024
Closest in time.
Xinpeng Wang, Bolei Ma, Chengzhi Hu, et al. 2024 · 2024
Closest in time.