Fetching the paper…
Reading the bibliography…
In-context learning (ICL) and supervised fine-tuning (SFT) are two common strategies for improving the performance of modern large language models (LLMs) on specific tasks.
A statistical method for evaluating systematic relationships
Robert R Sokal · 1958
Earlier work this paper cites.
Objective criteria for the evaluation of clustering methods
William M. Rand · 1971
Earlier work this paper cites.
Comparing partitions
Lawrence Hubert and Phipps Arabie · 1985
Earlier work this paper cites.
A study of the comparability of external criteria for hierarchical cluster analysis
Glenn W. Milligan and Martha C. Cooper · 1986
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction : with 200 Full-color Illustrations
R. Tibshirani, T. Hastie, and J.H. Friedman · 2001
Earlier work this paper cites.
Clustering by fast search and find of density peaks
Alex Rodriguez and Alessandro Laio · 2014
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2017
Guillaume Alain and Yoshua Bengio · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass · 2017
Earlier work this paper cites.
Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks
Yonatan Belinkov, Lluís Màrquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Deep rnns encode soft hierarchical syntax
T. Blevins, O. Levy, and L. Zettlemoyer · 2018
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim · 2019
Earlier work this paper cites.
Intrinsic dimension of data representations in deep neural networks
Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Earlier work this paper cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh · 2019
Earlier work this paper cites.
The intrinsic dimension of protein sequence evolution
Elena Facco, Andrea Pagnani, Elena Tea Russo, and Alessandro Laio · 2019
Earlier work this paper cites.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky · 2020
Earlier work this paper cites.
Hierarchical nucleation in deep neural networks
Diego Doimo, Aldo Glielmo, Alessio Ansuini, and Alessandro Laio · 2020
Earlier work this paper cites.
Asking without telling: Exploring latent ontologies in contextual representations
Julian Michael, Jan A. Botha, and Ian Tenney · 2020
Earlier work this paper cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing, 2021
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2021
Cited alongside, same era.
All bark and no bite: Rogue dimensions in transformer language models obscure representational quality
William Timkey and Marten van Schijndel · 2021
Cited alongside, same era.
Analysis and evaluation of language models for word sense disambiguation
Daniel Loureiro, Kiamehr Rezaee, Mohammad Taher Pilehvar, and José Camacho-Collados · 2021
Cited alongside, same era.
Automatic topography of high-dimensional data sets by non-parametric density peak clustering
Maria d’Errico, Elena Facco, Alessandro Laio, and Alex Rodriguez · 2021
Cited alongside, same era.
Isotropy in the contextual embedding space: Clusters and manifolds
Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church · 2021
Cited alongside, same era.
Prompting GPT-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Lee Boyd-Graber, and Lijuan Wang · 2023
Later among the works it cites.
Few-shot fine-tuning vs. in-context learning: A fair comparison and evaluation
Marius Mosbach, Tiago Pimentel, Shauli Ravfogel, Dietrich Klakow, and Yanai Elazar · 2023
Later among the works it cites.
Bridging information-theoretic and geometric compression in language models
Emily Cheng, Corentin Kervadec, and Marco Baroni · 2023
Later among the works it cites.
Samuel Marks and Max Tegmark · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Cited alongside, same era.
How many data points is a prompt worth?
Teven Le Scao and Alexander Rush · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
A cluster-based approach for improving isotropy in contextual embedding space
Sara Rajaee and Mohammad Taher Pilehvar · 2021
Cited alongside, same era.
Automatic topography of high-dimensional data sets by non-parametric density peak clustering
Maria d’Errico, Elena Facco, Alessandro Laio, and Alex Rodriguez · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Cited alongside, same era.
Exploring the landscape of distributional robustness for question answering models
Anas Awadalla, Mitchell Wortsman, Gabriel Ilharco, Sewon Min, Ian Magnusson, Hannaneh Hajishirzi, and Ludwig Schmidt · 2022
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao · 2023
Later among the works it cites.
DyLoRA: Parameter-efficient tuning of pre-trained models using dynamic search-free low-rank adaptation
Mojtaba Valipour, Mehdi Rezagholizadeh, Ivan Kobyzev, and Ali Ghodsi · 2023
Later among the works it cites.
Sparse low-rank adaptation of pre-trained language models
Ning Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen, Bowen Zhou, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla, 2023
Tom Lieberum, Matthew Rahtz, János Kramár, Neel Nanda, Geoffrey Irving, Rohin Shah, and Vladimir Mikulik · 2023
Later among the works it cites.
TheoremQA: A theorem-driven question answering dataset
Wenhu Chen, Ming Yin, Max Ku, Pan Lu, Yixin Wan, Xueguang Ma, Jianyu Xu, Xinyi Wang, and Tony Xia · 2023
Later among the works it cites.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2024
Closest in time.
The geometry of categorical and hierarchical concepts in large language models, 2024
Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch · 2024
Closest in time.
Not all language model features are linear, 2024
Joshua Engels, Isaac Liao, Eric J. Michaud, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
The geometry of hidden representations of large transformer models
Lucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio, Alessio Ansuini, and Alberto Cazzaniga · 2024
Closest in time.
Emergence of a high-dimensional abstraction phase in language transformers
Emily Cheng, Diego Doimo, Corentin Kervadec, Iuri Macocco, Jade Yu, Alessandro Laio, and Marco Baroni · 2024
Closest in time.
Comparing specialised small and general large language models on text classification: 100 labelled samples to achieve break-even performance, 2024
Branislav Pecher, Ivan Srba, and Maria Bielikova · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Scibench: Evaluating college-level scientific problem-solving abilities of large language models, 2024
Xiaoxuan Wang, Ziniu Hu, Pan Lu, Yanqiao Zhu, Jieyu Zhang, Satyen Subramaniam, Arjun R. Loomba, Shichang Zhang, Yizhou Sun, and Wei Wang · 2024
Closest in time.