Fetching the paper…
Reading the bibliography…
Foundation models are first pre-trained on vast unsupervised datasets and then fine-tuned on labeled data.
On the mathematical foundations of theoretical statistics
Ronald A Fisher · 1922
Earlier work this paper cites.
Cours d’économie politique
Vilfredo Pareto · 1964
Earlier work this paper cites.
A value pluralism model of ideological reasoning
Philip E Tetlock · 1986
Earlier work this paper cites.
Choosing preferences by constructing institutions: A cultural theory of preference formation
Aaron Wildavsky · 1987
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker and Yann Le Cun · 1988
Earlier work this paper cites.
Neural network ensembles
Lars Kai Hansen and Peter Salamon · 1990
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, J. S. Denker, Sara A. Solla, R. E. Howard, and L.D. Jackel · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Robot shaping: Developing autonomous agents through learning
Marco Dorigo and Marco Colombetti · 1994
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
An overview of statistical learning theory
Vladimir N Vapnik · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Handling preferences in evolutionary multiobjective optimization: A survey
CA Coello · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, Kishore, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
Nicol N Schraudolph · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard Hovy · 2003
Earlier work this paper cites.
Multitask reinforcement learning on the distribution of mdps
Fumihide Tanaka and Masayuki Yamamura · 2003
Earlier work this paper cites.
METEOR: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Aaron Wilson, Alan Fern, Soumya Ray, and Prasad Tadepalli · 2007
Earlier work this paper cites.
Learning all optimal policies with multiple criteria
Leon Barrett and Srini Narayanan · 2008
Earlier work this paper cites.
On the limitations of scalarisation for multi-objective reinforcement learning of pareto fronts
Peter Vamplew, John Yearwood, Richard Dazeley, and Adam Berry · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Empirical evaluation methods for multiobjective reinforcement learning algorithms
Peter Vamplew, Richard Dazeley, Adam Berry, Rustam Issabekov, and Evan Dekker · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz et al · 2012
Earlier work this paper cites.
AVA: A large-scale database for aesthetic visual analysis
Naila Murray, Luca Marchesotti, and Florent Perronnin · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Efficient backprop
Yann LeCun, Léon Bottou, Genevieve B. Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Earlier work this paper cites.
Performance metric ensemble for multiobjective evolutionary algorithms
Gary G Yen and Zhenan He · 2013
Earlier work this paper cites.
A multiobjective reinforcement learning approach to water resources systems operation: Pareto frontier approximation in a single run
Andrea Castelletti, Francesca Pianosi, and Marcello Restelli · 2013
Earlier work this paper cites.
The Arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Learning and transferring mid-level image representations using convolutional neural networks
Maxime Oquab, Leon Bottou, Ivan Laptev, and Josef Sivic · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Kristof Van Moffaert and Ann Nowé · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Reinforcement learning and the reward engineering principle
Dan Dewey · 2014
Earlier work this paper cites.
A novel adaptive weight selection algorithm for multi-objective multi-agent reinforcement learning
Kristof Van Moffaert, Tim Brys, Arjun Chandra, Lukas Esterle, Peter R Lewis, and Ann Nowé · 2014
Earlier work this paper cites.
Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Alignment for advanced machine learning systems
Jessica Taylor, Eliezer Yudkowsky, Patrick LaVictoire, and Andrew Critch · 2016
Earlier work this paper cites.
Modeling context in referring expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg · 2016
Earlier work this paper cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and P. Abbeel · 2016
Earlier work this paper cites.
Multi-objective deep reinforcement learning
Hossam Mossalam, Yannis M Assael, Diederik M Roijers, and Shimon Whiteson · 2016
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Detecting opinion spam and fake news using n-gram analysis and semantic similarity
Hadeer Ahmed · 2017
Earlier work this paper cites.
Tl; dr: Mining reddit to learn automatic summarization
Michael Völske, Martin Potthast, Shahbaz Syed, and Benno Stein · 2017
Earlier work this paper cites.
Distral: Robust multitask reinforcement learning
Yee Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Human-aligned artificial intelligence is a multiobjective problem
Peter Vamplew, Richard Dazeley, Cameron Foale, Sally Firmin, and Jane Mummery · 2018
Earlier work this paper cites.
Deep learning using rectified linear units (relu)
Abien Fred Agarap · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
Detecting opinion spams and fake news using text classification
Hadeer Ahmed, Issa Traore, and Sherif Saad · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Cited alongside, same era.
Neuroaesthetics and art’s diversity and universality
Marcos Nadal and Anjan Chatterjee · 2019
Cited alongside, same era.
A generalized algorithm for multi-objective reinforcement learning and policy adaptation
Runzhe Yang, Xingyuan Sun, and Karthik Narasimhan · 2019
Cited alongside, same era.
Dynamic weights in multi-objective deep reinforcement learning
ColD fusion: Collaborative descent for distributed multitask finetuning
Shachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim, Yoav Katz, and Leshem Choshen · 2022
Later among the works it cites.
The role of permutation invariance in linear mode connectivity of neural networks
Rahim Entezari, Hanie Sedghi, Olga Saukh, and Behnam Neyshabur · 2022
Later among the works it cites.
Self-instruct: Aligning language model with self generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2022
Later among the works it cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Later among the works it cites.
ExpansionNet v2: Block static expansion in fast end to end training for image captioning
Jia Cheng Hu, Roberto Cavicchioli, and Alessandro Capotondi · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Axel Abels, Diederik Roijers, Tom Lenaerts, Ann Nowé, and Denis Steckelmacher · 2019
Cited alongside, same era.
Limitations of the empirical fisher approximation for natural gradient descent
Frederik Kunstner, Philipp Hennig, and Lukas Balles · 2019
Cited alongside, same era.
Masked language model scoring
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2020
Cited alongside, same era.
Swin Transformer V2: Scaling up capacity and resolution
Ze Liu, Han Hu, Yutong Lin, Zhuliang Yao, Zhenda Xie, Yixuan Wei, Jia Ning, Yue Cao, Zheng Zhang, Li Dong, Furu Wei, and Baining Guo · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trebacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Later among the works it cites.
Pareto set learning for expensive multi-objective optimization
Xi Lin, Zhiyuan Yang, Xiaoyuan Zhang, and Qingfu Zhang · 2022
Later among the works it cites.
Pareto manifold learning: Tackling multiple tasks via ensembles of single-task models
Nikolaos Dimitriadis, Pascal Frossard, and François Fleuret · 2022
Later among the works it cites.
Learning a subspace of policies for online adaptation in reinforcement learning
Jean-Baptiste Gaya, Laure Soulier, and Ludovic Denoyer · 2022
Later among the works it cites.
X-risk analysis for AI research
Dan Hendrycks and Mantas Mazeika · 2022
Later among the works it cites.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2022
Later among the works it cites.
Defining and characterizing reward gaming
Joar Max Viktor Skalse, Nikolaus H. R. Howe, Dmitrii Krasheninnikov, and David Krueger · 2022
Later among the works it cites.
Goal misgeneralization: Why correct specifications aren’t enough for correct goals
Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton · 2022
Later among the works it cites.
Goal misgeneralization in deep reinforcement learning
Lauro Langosco Di Langosco, Jack Koch, Lee D Sharkey, Jacob Pfau, and David Krueger · 2022
Later among the works it cites.
Momentum-based weight interpolation of strong zero-shot models for continual learning
Zafir Stojanovski, Karsten Roth, and Zeynep Akata · 2022
Later among the works it cites.
Weight averaging: A simple yet effective method to overcome catastrophic forgetting in automatic speech recognition
Steven Vander Eeckt et al · 2022
Later among the works it cites.
Branch-Train-Merge: Embarrassingly parallel training of expert language models
Margaret Li, Suchin Gururangan, Tim Dettmers, Mike Lewis, Tim Althoff, Noah A Smith, and Luke Zettlemoyer · 2022
Later among the works it cites.
Fishr: Invariant gradient variances for out-of-distribution generalization
Alexandre Rame, Corentin Dancette, and Matthieu Cord · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Mitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li, Hanna Hajishirzi, Ali Farhadi, Hongseok Namkoong, and Ludwig Schmidt · 2022
Later among the works it cites.
Tuning computer vision models with task rewards
André Susano Pinto, Alexander Kolesnikov, Yuge Shi, Lucas Beyer, and Xiaohua Zhai · 2023
Closest in time.
Gpt-4 technical report
OpenAI · 2023
Closest in time.
Stanford Alpaca: An instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Reinforcement learning for language models
Yoav Goldberg · 2023
Closest in time.
Aligning text-to-image models using human feedback
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Closest in time.
Better aligning text-to-image models with human preference
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li · 2023
Closest in time.
HIVE: Harnessing human feedback for instructional visual editing
Shu Zhang, Xinyi Yang, Yihao Feng, Can Qin, Chia-Chih Chen, Ning Yu, Zeyuan Chen, Huan Wang, Silvio Savarese, Stefano Ermon, et al · 2023
Closest in time.
Reward design with language models
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh · 2023
Closest in time.
ImageReward: Learning and evaluating human preferences for text-to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong · 2023
Closest in time.
Rewarding chatbots for real-world engagement with millions of users
Robert Irvine, Douglas Boubert, Vyas Raina, Adian Liusie, Vineet Mudupalli, Aliaksei Korshuk, Zongyi Liu, Fritz Cremer, Valentin Assassi, Christie-Carol Beauchamp, et al · 2023
Closest in time.
LLaMA: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
Generative CI through collective response systems
Aviv Ovadya · 2023
Closest in time.
Large language models as superpositions of cultural perspectives
Grgur Kovač, Masataka Sawayama, Rémy Portelas, Cédric Colas, Peter Ford Dominey, and Pierre-Yves Oudeyer · 2023
Closest in time.
Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale · 2023
Closest in time.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the MACHIAVELLI benchmark
Alexander Pan, Chan Jun Shern, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Closest in time.
Aligning human preferences with baseline objectives in reinforcement learning
Daniel Marta, Simon Holk, Christian Pek, Jana Tumova, and Iolanda Leite · 2023
Closest in time.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Closest in time.
Model Ratatouille: Recycling diverse models for out-of-distribution generalization
Alexandre Ramé, Kartik Ahuja, Jianyu Zhang, Matthieu Cord, Léon Bottou, and David Lopez-Paz · 2023
Closest in time.
Git Re-Basin: Merging models modulo permutation symmetries
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2023
Closest in time.
Fine-tuning 20B LLMs with RLHF on a 24GB consumer GPU
Edward Beeching, Younes Belkada, Leandro von Werra, Sourab Mangrulkar, Lewis Tunstall, and Kashif Rasul · 2023
Closest in time.
Huggingface h4 stack exchange preference dataset, 2023
Nathan Lambert, Lewis Tunstall, Nazneen Rajani, and Tristan Thrush · 2023
Closest in time.
Openassistant conversations–democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi-Rui Tam, Keith Stevens, Abdullah Barhoum, Nguyen Minh Duc, Oliver Stanley, Richárd Nagyfi, et al · 2023
Closest in time.
Enze Xie, Lewei Yao, Han Shi, Zhili Liu, Daquan Zhou, Zhaoqiang Liu, Jiawei Li, and Zhenguo Li · 2023
Closest in time.
Unified model for image, video, audio and language tasks
Mustafa Shukor, Corentin Dancette, Alexandre Rame, and Matthieu Cord · 2023
Closest in time.
Aligning language models with preferences through f-divergence minimization
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman · 2023
Closest in time.
RRHF: Rank responses to align language models with human feedback without tears
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Closest in time.
Simple emergent action representations from multi-task policy training
Pu Hua, Yubei Chen, and Huazhe Xu · 2023
Closest in time.
Seasoning model soups for robustness to adversarial and natural distribution shifts
Francesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, and Sven Gowal · 2023
Closest in time.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2023
Closest in time.
Merging decision transformers: Weight averaging for forming multi-task policies
Daniel Lawson and Ahmed H Qureshi · 2023
Closest in time.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2023
Closest in time.
Elastic weight removal for faithful and abstractive dialogue generation
Nico Daheim, Nouha Dziri, Mrinmaya Sachan, Iryna Gurevych, and Edoardo M Ponti · 2023
Closest in time.
Exploring the benefits of training expert language models over instruction tuning
Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee, and Minjoon Seo · 2023
Closest in time.
Natural selection favors AIs over humans
Dan Hendrycks · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
Lamp: When large language models meet personalization
Alireza Salemi, Sheshera Mysore, Michael Bendersky, and Hamed Zamani · 2023
Closest in time.
Alpaca-LoRA
Eric J. Wang · 2023
Closest in time.
StackLLaMA: An RL Fine-tuned LLaMA Model for Stack Exchange Question and Answering, 2023
Edward Beeching, Younes Belkada, Kashif Rasul, Lewis Tunstall, Leandro von Werra, Nazneen Rajani, and Nathan Lambert · 2023
Closest in time.
RAFT: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Closest in time.