Fetching the paper…
Reading the bibliography…
The development of generative language models that can create long and coherent textual outputs via autoregression has lead to a proliferation of uses and a corresponding sweep of analyses as researches work to determine the limitations of this new paradigm.
“Good debt or bad debt: Detecting semantic orientations in economic texts”
Pekka Malo et al · 2014
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Image transformer”
Niki Parmar et al · 2018
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“The bitter lesson”
Richard Sutton · 2019
Earlier work this paper cites.
“CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge”
Alon Talmor, Jonathan Herzig, Nicholas Lourie and Jonathan Berant · 2019
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”
Dan Hendrycks et al · 2020
Earlier work this paper cites.
“Reformer: The efficient transformer”
Nikita Kitaev, Łukasz Kaiser and Anselm Levskaya · 2020
Earlier work this paper cites.
“Multi-hop attention graph neural network”
Guangtao Wang, Rex Ying, Jing Huang and Jure Leskovec · 2020
Earlier work this paper cites.
“Incorporating residual and normalization layers into analysis of masked language models”
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi and Kentaro Inui · 2021
Earlier work this paper cites.
Josh Achiam et al · 2023
Earlier work this paper cites.
Rohan Anil et al · 2023
Earlier work this paper cites.
Albert Jiang et al · 2023
Cited alongside, same era.
“A survey on fairness in large language models”
Yingji Li et al · 2023
Cited alongside, same era.
R McCoy et al · 2023
Cited alongside, same era.
“Proving test set contamination in black box language models”
Yonatan Oren et al · 2023
Cited alongside, same era.
“Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions”
Pouya Pezeshkpour and Estevam Hruschka · 2023
Cited alongside, same era.
“When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model Leaderboards”
Norah Alzahrani et al · 2024
Closest in time.
“Make Your LLM Fully Utilize the Context”
Shengnan An et al · 2024
Closest in time.
“A Primer on the Inner Workings of Transformer-based Language Models”
Javier Ferrando, Gabriele Sarti, Arianna Bisazza and Marta Costa-jussà · 2024
Closest in time.
“A comprehensive survey on deep graph representation learning”
Wei Ju et al · 2024
Closest in time.
“Lost in the middle: How language models use long contexts”
Nelson Liu et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“RoFormer: Enhanced Transformer with Rotary Position Embedding”, 2023
Jianlin Su et al · 2023
Cited alongside, same era.
“Do llms exhibit human-like response biases? a case study in survey design”
Lindia Tjuatja et al · 2023
Cited alongside, same era.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Cited alongside, same era.
Mark.. Adian Potsawee · 2024
Cited alongside, same era.
“Llama 3 Model Card”, 2024
AI@Meta · 2024
Cited alongside, same era.
Tsendsuren Munkhdalai, Manaal Faruqui and Siddharth Gopal · 2024
Closest in time.
“In-Context Impersonation Reveals Large Language Models’ Strengths and Biases”
Leonard Salewski et al · 2024
Closest in time.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2024
Closest in time.
“Adapted large language models can outperform medical experts in clinical text summarization”
Dave Van et al · 2024
Closest in time.
“The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions”
Eric Wallace et al · 2024
Closest in time.
“Large Language Models Are Not Robust Multiple Choice Selectors”
Chujie Zheng et al · 2024
Closest in time.