Fetching the paper…
Reading the bibliography…
This position paper argues that in many realistic (i.e., complex, contextualized, subjective) scenarios, one LLM is not enough to produce a reliable output.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. J · 1989
Earlier work this paper cites.
Distributed cognition
Hutchins, E · 2000
Earlier work this paper cites.
Collaboration, peer review and open source software
Johnson, J. P · 2006
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., teusz Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2020
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Pitfalls of static language modelling
Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., de Masson d’Autume, C., Ruder, S., Yogatama, D., et al · 2021
Earlier work this paper cites.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., and Choi, Y · 2021
Earlier work this paper cites.
Time-aware language models as temporal knowledge bases
Dhingra, B., Cole, J. R., Eisenschlos, J. M., Gillick, D., Eisenstein, J., and Cohen, W. W · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Earlier work this paper cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models
Press, O., Zhang, M., Min, S., Schmidt, L., Smith, N. A., and Lewis, M · 2022
Earlier work this paper cites.
Paradigm shift in natural language processing
Sun, T.-X., Liu, X.-Y., Qiu, X.-P., and Huang, X.-J · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Token-level adaptation of lora adapters for downstream task generalization
Belofsky, J · 2023
Earlier work this paper cites.
Frugalgpt: How to use large language models while reducing cost and improving performance
Chen, L., Zaharia, M., and Zou, J · 2023
Earlier work this paper cites.
Lm vs lm: Detecting factual errors via cross examination
Cohen, R., Hamri, M., Geva, M., and Globerson, A · 2023
Earlier work this paper cites.
Diao, S., Xu, T., Xu, R., Wang, J., and Zhang, T · 2023
Earlier work this paper cites.
The benefits of bad advice: Autocontrastive decoding across model layers
Gera, A., Friedman, R., Arviv, O., Gunasekara, C., Sznajder, B., Slonim, N., and Shnarch, E · 2023
Earlier work this paper cites.
Scaling expert language models with unsupervised domain discovery
Gururangan, S., Li, M., Lewis, M., Shi, W., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2023
Earlier work this paper cites.
Ssd-2: Scaling and inference-time fusion of diffusion language models
Han, X., Kumar, S., Tsvetkov, Y., and Ghazvininejad, M · 2023
Earlier work this paper cites.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Huang, C., Liu, Q., Lin, B. Y., Pang, T., Du, C., and Lin, M · 2023
Earlier work this paper cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2023
Earlier work this paper cites.
Personalized soups: Personalized large language model alignment via post-hoc parameter merging
Jang, J., Kim, S., Lin, B. Y., Wang, Y., Hessel, J., Zettlemoyer, L., Hajishirzi, H., Choi, Y., and Ammanabrolu, P · 2023
Earlier work this paper cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Earlier work this paper cites.
Active retrieval augmented generation
Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G · 2023
Earlier work this paper cites.
Copyright violations and large language models
Karamolegkou, A., Li, J., Zhou, L., and Søgaard, A · 2023
Earlier work this paper cites.
Language generation models can cause harm: So what can we do about it? an actionable survey
Kumar, S., Balachandran, V., Njoo, L., Anastasopoulos, A., and Tsvetkov, Y · 2023
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T., Zettlemoyer, L., and Lewis, M · 2023
Earlier work this paper cites.
Routing to the expert: Efficient reward-guided ensemble of large language models
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J · 2023
Cited alongside, same era.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2023
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Cited alongside, same era.
Having beer after prayer? measuring cultural bias in large language models
Naous, T., Ryan, M. J., Ritter, A., and Xu, W · 2023
Cited alongside, same era.
PREADD: Prefix-adaptive decoding for controlled text generation
Pei, J., Yang, K., and Klein, D · 2023
Realtime qa: what’s the answer right now?
Kasai, J., Sakaguchi, K., Le Bras, R., Asai, A., Yu, X., Radev, D., Smith, N. A., Choi, Y., Inui, K., et al · 2024
Later among the works it cites.
Kirk, H. R., Whitefield, A., Röttger, P., Bean, A., Margatina, K., Ciro, J., Mosquera, R., Bartolo, M., Williams, A., He, H., et al · 2024
Later among the works it cites.
Compo: Community preferences for language model personalization
Kumar, S., Park, C. Y., Tsvetkov, Y., Smith, N. A., and Hajishirzi, H · 2024
Later among the works it cites.
A theory of appropriateness with applications to generative artificial intelligence
Leibo, J. Z., Vezhnevets, A. S., Diaz, M., Agapiou, J. P., Cunningham, W. A., Sunehag, P., Haas, J., Koster, R., Duéñez-Guzmán, E. A., Isaac, W. S., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Whose opinions do language models reflect?
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P., and Hashimoto, T · 2023
Cited alongside, same era.
Replug: Retrieval-augmented black-box language models
Shi, W., Min, S., Yasunaga, M., Seo, M., James, R., Lewis, M., Zettlemoyer, L., and Yih, W.-t · 2023
Cited alongside, same era.
Large language model routing with benchmark datasets
Shnitzer, T., Ou, A., Silva, M., Soule, K., Sun, Y., Solomon, J., Thompson, N., and Yurochkin, M · 2023
Cited alongside, same era.
Globalbench: A benchmark for global progress in natural language processing
Song, Y., Khanuja, S., Liu, P., Faisal, F., Ostapenko, A., Winata, G., Aji, A., Cahyawijaya, S., Tsvetkov, Y., Anastasopoulos, A., et al · 2023
Cited alongside, same era.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
Stolfo, A., Belinkov, Y., and Sachan, M · 2023
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q., Chi, E., Zhou, D., et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., et al · 2023
Cited alongside, same era.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., and Bing, L · 2024
Later among the works it cites.
Mitigating the alignment tax of rlhf
Lin, Y., Lin, H., Xiong, W., Diao, S., Liu, J., Zhang, J., Pan, R., Wang, H., Hu, W., Zhang, H., et al · 2024
Later among the works it cites.
Tuning language models by proxy
Liu, A., Han, X., Wang, Y., Tsvetkov, Y., Choi, Y., and Smith, N. A · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Later among the works it cites.
Pack of LLMs: Model fusion at test-time via perplexity optimization
Mavromatis, C., Karypis, P., and Karypis, G · 2024
Later among the works it cites.
Silo language models: Isolating legal risk in a nonparametric datastore
Min, S., Gururangan, S., Wallace, E., Shi, W., Hajishirzi, H., Smith, N. A., and Zettlemoyer, L · 2024
Later among the works it cites.
Can llms keep a secret? testing privacy implications of language models via contextual integrity theory
Mireshghallah, N., Kim, H., Zhou, X., Tsvetkov, Y., Sap, M., Shokri, R., and Choi, Y · 2024
Later among the works it cites.
An emulator for fine-tuning large language models using small language models
Mitchell, E., Rafailov, R., Sharma, A., Finn, C., and Manning, C. D · 2024
Later among the works it cites.
Learning to route among specialized experts for zero-shot generalization
Muqeeth, M., Liu, H., Liu, Y., and Raffel, C · 2024
Later among the works it cites.
Routellm: Learning to route llms with preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Later among the works it cites.
Normad: A benchmark for measuring the cultural adaptability of large language models
Rao, A., Yerukola, A., Shah, V., Reinecke, K., and Sap, M · 2024
Later among the works it cites.
Mitigating hallucinations and off-target machine translation with source-contrastive and language-contrastive decoding
Sennrich, R., Vamvas, J., and Mohammadshahi, A · 2024
Later among the works it cites.
Learning to decode collaboratively with multiple language models
Shen, Z., Lang, H., Wang, B., Kim, Y., and Sontag, D · 2024
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, W.-t · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning
Sprague, Z., Yin, F., Rodriguez, J. D., Jiang, D., Wadhwa, M., Singhal, P., Zhao, X., Ye, X., Mahowald, K., and Durrett, G · 2024
Later among the works it cites.
Tensoropera router: A multi-model router for efficient llm inference
Stripelis, D., Hu, Z., Zhang, J., Xu, Z., Shah, A. D., Jin, H., Yao, Y., Avestimehr, S., and He, C · 2024
Later among the works it cites.
Debategpt: Fine-tuning large language models with multi-agent debate supervision
Subramaniam, V., Torralba, A., and Li, S · 2024
Later among the works it cites.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm
Sukhbaatar, S., Golovneva, O., Sharma, V., Xu, H., Lin, X. V., Rozière, B., Kahn, J., Li, D., Yih, W.-t., Weston, J., et al · 2024
Later among the works it cites.
Assessing programming task difficulty for efficient evaluation of large language models
Tambon, F., Nikanjam, A., Khomh, F., and Antoniol, G · 2024
Later among the works it cites.
Fusing models with complementary expertise
Wang, H., Polo, F. M., Sun, Y., Kundu, S., Xing, E., and Yurochkin, M · 2024
Later among the works it cites.
Evaluating copyright takedown methods for language models
Wei, B., Shi, W., Huang, Y., Smith, N. A., Zhang, C., Zettlemoyer, L., Li, K., and Henderson, P · 2024
Later among the works it cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., et al · 2024
Later among the works it cites.
Sayself: Teaching llms to express confidence with self-reflective rationales
Xu, T., Wu, S., Diao, S., Liu, X., Wang, X., Chen, Y., and Gao, J · 2024
Later among the works it cites.
Do large language models latently perform multi-hop reasoning?
Yang, S., Gribovskaya, E., Kassner, N., Geva, M., and Riedel, S · 2024
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, Z., and Zhang, Y · 2024
Later among the works it cites.
Lofit: Localized fine-tuning on llm representations
Yin, F., Ye, X., and Durrett, G · 2024
Later among the works it cites.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
Yu, L., Yu, B., Yu, H., Huang, F., and Li, Y · 2024
Later among the works it cites.
Zhang, K., Wang, J., Hua, E., Qi, B., Ding, N., and Zhou, B · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
A call to build models like we build open-source software, 2021
Raffel, C · 2025
Closest in time.
Multiagent finetuning: Self improvement with diverse reasoning chains
Subramaniam, V., Du, Y., Tenenbaum, J. B., Torralba, A., Li, S., and Mordatch, I · 2025
Closest in time.