Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have been widely used in various applications but are known to suffer from issues related to untruthfulness and toxicity.
Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification
Borkan, D.; Dixon, L.; Sorensen, J.; Thain, N.; and Vasserman, L. 2019 · 1903
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 2005
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T. B.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
ZeRO: memory optimizations toward training trillion parameter models
Rajbhandari, S.; Rasley, J.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters
Rasley, J.; Rajbhandari, S.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
Editable Neural Networks
Sinitsin, A.; Plokhotnyuk, V.; Pyrkin, D. V.; Popov, S.; and Babenko, A. 2020 · 2020
Earlier work this paper cites.
Neural Text Generation With Unlikelihood Training
Welleck, S.; Kulikov, I.; Roller, S.; Dinan, E.; Cho, K.; and Weston, J. 2020 · 2020
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
De Cao, N.; Aziz, W.; and Titov, I. 2021 · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
Liu, A.; Sap, M.; Lu, X.; Swayamdipta, S.; Bhagavatula, C.; Smith, N. A.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Challenges in Detoxifying Language Models
Welbl, J.; Glaese, A.; Uesato, J.; Dathathri, S.; Mellor, J.; Hendricks, L. A.; Anderson, K.; Kohli, P.; Coppin, B.; and Huang, P. 2021 · 2021
Earlier work this paper cites.
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
Geva, M.; Caciularu, A.; Wang, K. R.; and Goldberg, Y. 2022 · 2022
Earlier work this paper cites.
Towards a Unified View of Parameter-Efficient Transfer Learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2022 · 2022
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S.; Hilton, J.; and Evans, O. 2022 · 2022
Cited alongside, same era.
QUARK: Controllable Text Generation with Reinforced Unlearning
Lu, X.; Welleck, S.; Hessel, J.; Jiang, L.; Qin, L.; West, P.; Ammanabrolu, P.; and Choi, Y. 2022 · 2022
Cited alongside, same era.
Merging Models with Fisher-Weighted Averaging
Matena, M.; and Raffel, C. 2022 · 2022
Cited alongside, same era.
Locating and Editing Factual Associations in GPT
Meng, K.; Bau, D.; Andonian, A.; and Belinkov, Y. 2022 · 2022
Cited alongside, same era.
Fast Model Editing at Scale
Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; and Manning, C. D. 2022a · 2022
Cited alongside, same era.
LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition
Huang, C.; Liu, Q.; Lin, B. Y.; Pang, T.; Du, C.; and Lin, M. 2023 · 2023
Closest in time.
Editing models with task arithmetic
Ilharco, G.; Ribeiro, M. T.; Wortsman, M.; Schmidt, L.; Hajishirzi, H.; and Farhadi, A. 2023 · 2023
Closest in time.
Survey of Hallucination in Natural Language Generation
Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.; Madotto, A.; and Fung, P. 2023 · 2023
Closest in time.
Dataless Knowledge Fusion by Merging Weights of Language Models
Jin, X.; Ren, X.; Preotiuc-Pietro, D.; and Cheng, P. 2023 · 2023
Closest in time.
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Li, K.; Hopkins, A. K.; Bau, D.; Viégas, F. B.; Pfister, H.; and Wattenberg, M. 2023b · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Memory-Based Model Editing at Scale
Mitchell, E.; Lin, C.; Bosselut, A.; Manning, C. D.; and Finn, C. 2022b · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P. F.; Leike, J.; and Lowe, R. 2022 · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022 · 2022
Cited alongside, same era.
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Wortsman, M.; Ilharco, G.; Gadre, S. Y.; Roelofs, R.; Lopes, R. G.; Morcos, A. S.; Namkoong, H.; Farhadi, A.; Carmon, Y.; Kornblith, S.; and Schmidt, L. 2022 · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M. T.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Language and Task Arithmetic with Parameter-Efficient Layers for Zero-Shot Summarization
Chronopoulou, A.; Pfeiffer, J.; Maynez, J.; Wang, X.; Ruder, S.; and Agrawal, P. 2023 · 2023
Cited alongside, same era.
Li, X.; Zhang, T.; Dubois, Y.; Taori, R.; Gulrajani, I.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023d · 2023
Closest in time.
Evaluating Verifiability in Generative Search Engines
Liu, N. F.; Zhang, T.; and Liang, P. 2023 · 2023
Closest in time.
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Suzgun, M.; Scales, N.; Schärli, N.; Gehrmann, S.; Tay, Y.; Chung, H. W.; Chowdhery, A.; Le, Q. V.; Chi, E.; Zhou, D.; and Wei, J. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Closest in time.
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources
Wang, Y.; Ivison, H.; Dasigi, P.; Hessel, J.; Khot, T.; Chandu, K. R.; Wadden, D.; MacMillan, K.; Smith, N. A.; Beltagy, I.; and Hajishirzi, H. 2023 · 2023
Closest in time.
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N. A.; Ostendorf, M.; and Hajishirzi, H. 2023 · 2023
Closest in time.
WizardLM: Empowering Large Language Models to Follow Complex Instructions
Xu, C.; Sun, Q.; Zheng, K.; Geng, X.; Zhao, P.; Feng, J.; Tao, C.; and Jiang, D. 2023 · 2023
Closest in time.
AdaMerging: Adaptive Model Merging for Multi-Task Learning
Yang, E.; Wang, Z.; Shen, L.; Liu, S.; Guo, G.; Wang, X.; and Tao, D. 2023 · 2023
Closest in time.
Composing Parameter-Efficient Modules with Arithmetic Operations
Zhang, J.; Chen, S.; Liu, J.; and He, J. 2023 · 2023
Closest in time.
Zhao, R.; Li, X.; Chia, Y. K.; Ding, B.; and Bing, L. 2023 · 2023
Closest in time.