Fetching the paper…
Reading the bibliography…
The widespread application of LLMs across various tasks and fields has necessitated the alignment of these models with human values and preferences.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 1908
Earlier work this paper cites.
A theory of organization and change within value-attitude systems
Milton Rokeach. 1968 · 1968
Earlier work this paper cites.
The nature of human values
Milton Rokeach. 1973 · 1973
Earlier work this paper cites.
Some unresolved issues in theories of beliefs, attitudes, and values
Milton Rokeach. 1979 · 1979
Earlier work this paper cites.
Distributed representations
Geoffrey E Hinton. 1984 · 1984
Earlier work this paper cites.
A Prolegomenon to Theory of Translation
Robert Darrell Firmage. 1986 · 1986
Earlier work this paper cites.
Translationese in swedish novels translated from english
Martin Gellerstam. 1986 · 1986
Earlier work this paper cites.
Distributed representations
GE Hinton, JL McClelland, and DE Rumelhart. 1986 · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986 · 1986
Earlier work this paper cites.
Finding structure in time
Jeffrey L Elman. 1990 · 1990
Earlier work this paper cites.
Beyond individualism/collectivism: New cultural dimensions of values
Shalom H Schwartz. 1994 · 1994
Earlier work this paper cites.
World values surveys and european values surveys, 1981-1984, 1990-1993, and 1995-1997
Ronald Inglehart, Miguel Basanez, Jaime Diez-Medrano, Loek Halman, and Ruud Luijkx. 2000 · 1997
Earlier work this paper cites.
A theory of cultural values and some implications for work
Shalom H Schwartz. 1999 · 1999
Earlier work this paper cites.
Culture’s consequences: Comparing values, behaviors, institutions and organizations across nations
Geert Hofstede. 2001 · 2001
Earlier work this paper cites.
Human beliefs and values: A cross-cultural sourcebook based on the 1999-2002 values surveys
Ronald Inglehart. 2004 · 2002
Earlier work this paper cites.
English as a global language
David Crystal. 2003 · 2003
Earlier work this paper cites.
Culture, leadership, and organizations: The GLOBE study of 62 societies
Robert J House, Paul J Hanges, Mansour Javidan, Peter W Dorfman, and Vipin Gupta. 2004 · 2004
Earlier work this paper cites.
Mapping and interpreting cultural differences around the world
Shalom H Schwartz. 2004 · 2004
Earlier work this paper cites.
The role of english in scientific communication: lingua franca or tyrannosaurus rex?
C Tardy. 2004 · 2004
Earlier work this paper cites.
Cultures and organizations: Software of the mind , volume 2
Geert Hofstede, Gert Jan Hofstede, and Michael Minkov. 2005 · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Mapping global values
Ronald Inglehart. 2006 · 2006
Earlier work this paper cites.
Die messung von werten mit dem “portraits value questionnaire”
Peter Schmidt, Sebastian Bamberg, Eldad Davidov, Johannes Herrmann, and Shalom H Schwartz. 2007 · 2007
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2020 · 2008
Earlier work this paper cites.
Understanding human values
Milton Rokeach. 2008 · 2008
Earlier work this paper cites.
Cultural value orientations: Nature & implications of national differences
Shalom Schwartz. 2008 · 2008
Earlier work this paper cites.
Testing the discriminant validity of schwartz’portrait value questionnaire items–a replication and extension of knoppen and saris (2009)
Constanze Beierlein, Eldad Davidov, Peter Schmidt, Shalom H Schwartz, and Beatrice Rammstedt. 2012 · 2009
Earlier work this paper cites.
Managerial implications of the globe project: A study of 62 societies
Mansour Javidan and Ali Dastmalchian. 2009 · 2009
Earlier work this paper cites.
Distance metric learning for large margin nearest neighbor classification
Kilian Q Weinberger and Lawrence K Saul. 2009 · 2009
Earlier work this paper cites.
Identification of Translationese: A Machine Learning Approach , page 503–511
Iustina Ilisei, Diana Inkpen, Gloria Corpas Pastor, and Ruslan Mitkov. 2010 · 2010
Earlier work this paper cites.
Dimensionalizing cultures: The hofstede model in context
Geert Hofstede. 2011 · 2011
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz. 2012 · 2012
Earlier work this paper cites.
Vsm 2013
Geert Hofstede and Michael Minkov. 2013 · 2013
Earlier work this paper cites.
Automatic detection of machine translated text and translation quality estimation
Roee Aharoni, Moshe Koppel, and Yoav Goldberg. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014a · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014b · 2014
Earlier work this paper cites.
The phenomenon of linguistic globalization: English as the global lingua franca (eglf)
Vladimir M. Smokotin, Anna S. Alekseyenko, and Galina I. Petrova. 2014 · 2014
Earlier work this paper cites.
Deepface: Closing the gap to human-level performance in face verification
Yaniv Taigman, Ming Yang, Marc’Aurelio Ranzato, and Lior Wolf. 2014 · 2014
Earlier work this paper cites.
Siamese neural networks for one-shot image recognition
Gregory Koch, Richard Zemel, Ruslan Salakhutdinov, et al. 2015 · 2015
Earlier work this paper cites.
Deep learning and the information bottleneck principle
Naftali Tishby and Noga Zaslavsky. 2015 · 2015
Earlier work this paper cites.
Fully-convolutional siamese networks for object tracking
Luca Bertinetto, Jack Valmadre, Joao F Henriques, Andrea Vedaldi, and Philip HS Torr. 2016 · 2016
Earlier work this paper cites.
Multi-layer representation learning for medical concepts
Edward Choi, Mohammad Taha Bahadori, Elizabeth Searles, Catherine Coffey, Michael Thompson, James Bost, Javier Tejedor-Sojo, and Jimeng Sun. 2016 · 2016
Earlier work this paper cites.
Deep neural networks for youtube recommendations
Paul Covington, Jay Adams, and Emre Sargin. 2016 · 2016
Earlier work this paper cites.
Large-scale embedding learning in heterogeneous event data
Huan Gui, Jialu Liu, Fangbo Tao, Meng Jiang, Brandon Norick, and Jiawei Han. 2016 · 2016
Earlier work this paper cites.
Massive exploration of neural machine translation architectures
Denny Britz, Anna Goldie, Minh-Thang Luong, and Quoc Le. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Cited alongside, same era.
Neural collaborative filtering
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017 · 2017
Cited alongside, same era.
Spatial-aware object embeddings for zero-shot localization and classification of actions
Pascal Mettes and Cees GM Snoek. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
The Refined Theory of Basic Values , page 51–72
The CRINGE loss: Learning what language not to model
Leonard Adolphs, Tianyu Gao, Jing Xu, Kurt Shuster, Sainbayar Sukhbaatar, and Jason Weston. 2023 · 2023
Later among the works it cites.
Probing pre-trained language models for cross-cultural differences in values
Arnav Arora, Lucie-aimée Kaffee, and Isabelle Augenstein. 2023 · 2023
Later among the works it cites.
Maybe only 0.5% data is needed: A preliminary exploration of low training data instruction tuning
Hao Chen, Yiming Zhang, Qi Zhang, Hantao Yang, Xiaomeng Hu, Xuetao Ma, Yifan Yanggong, and Junbo Zhao. 2023 · 2023
Later among the works it cites.
Towards measuring the representation of subjective global opinions in language models
Esin Durmus, Karina Nyugen, Thomas I Liao, Nicholas Schiefer, Amanda Askell, Anton Bakhtin, Carol Chen, Zac Hatfield-Dodds, Danny Hernandez, Nicholas Joseph, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shalom H. Schwartz. 2017 · 2017
Cited alongside, same era.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
UMAP: uniform manifold approximation and projection for dimension reduction
Leland McInnes and John Healy. 2018 · 2018
Cited alongside, same era.
On the information bottleneck theory of deep learning
Andrew Michael Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan Daniel Tracey, and David Daniel Cox. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019 · 2019
Cited alongside, same era.
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte. 2023 · 2023
Later among the works it cites.
Multilingual language models are not multicultural: A case study in emotion
Shreya Havaldar, Bhumika Singhal, Sunny Rai, Langchen Liu, Sharath Chandra Guntuku, and Lyle Ungar. 2023 · 2023
Later among the works it cites.
Cyclealign: Iterative distillation from black-box llm to white-box models for better human alignment
Jixiang Hong, Quan Tu, Changyu Chen, Xing Gao, Ji Zhang, and Rui Yan. 2023 · 2023
Later among the works it cites.
Unnatural instructions: Tuning language models with (almost) no human labor
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023 · 2023
Later among the works it cites.
Learn what not to learn: Towards generative safety in chatbots
Leila Khalatbari, Yejin Bang, Dan Su, Willy Chung, Saeed Ghadimi, Hossein Sameti, and Pascale Fung. 2023 · 2023
Later among the works it cites.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi. 2023 · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts. 2023 · 2023
Later among the works it cites.
Using in-context learning to improve dialogue safety
Nicholas Meade, Spandana Gella, Devamanyu Hazarika, Prakhar Gupta, Di Jin, Siva Reddy, Yang Liu, and Dilek Hakkani-Tur. 2023 · 2023
Later among the works it cites.
Having beer after prayer? measuring cultural bias in large language models
Tarek Naous, Michael J Ryan, and Wei Xu. 2023 · 2023
Later among the works it cites.
Seallms–large language models for southeast asia
Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, et al. 2023 · 2023
Later among the works it cites.
Is ChatGPT a general-purpose natural language processing task solver?
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang. 2023 · 2023
Later among the works it cites.
Knowledge of cultural moral norms in large language models
Aida Ramezani and Yang Xu. 2023 · 2023
Later among the works it cites.
Training language models with language feedback at scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez. 2023 · 2023
Later among the works it cites.
Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, Lalit Pradhan, Zain Muhammad Mujahid, Massa Baali, Alham Fikri Aji, Zhengzhong Liu, Andy Hock, Andrew Feldman, Jonathan Lee, Andrew Jackson, Preslav Nakov, Timothy Baldwin, and Eric Xing. 2023 · 2023
Later among the works it cites.
Large language model alignment: A survey
Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023 · 2023
Later among the works it cites.
An empirical study of instruction-tuning large language models in chinese
Qingyi Si, Tong Wang, Zheng Lin, Xu Zhang, Yanan Cao, and Weiping Wang. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. 2023 · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al. 2023 · 2023
Later among the works it cites.
From instructions to intrinsic human values–a survey of alignment goals for big models
Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang, and Xing Xie. 2023 · 2023
Later among the works it cites.
InstructSafety: A unified framework for building multidimensional and explainable safety detector through instruction tuning
Zhexin Zhang, Jiale Cheng, Hao Sun, Jiawen Deng, and Minlie Huang. 2023b · 2023
Later among the works it cites.
LIMA: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, LILI YU, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023 · 2023
Later among the works it cites.
The multilingual alignment prism: Aligning global and local preferences to reduce harm
Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01.ai
01. AI, :, Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Tao Yu, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zhiyuan Liu, and Zonghong Dai. 2024 · 2024
Closest in time.
Llama 3 model card
AI@Meta. 2024 · 2024
Closest in time.
Investigating cultural alignment of large language models
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, and Mona Diab. 2024 · 2024
Closest in time.
Safety-tuned LLaMAs: Lessons from improving the safety of large language models that follow instructions
Federico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Rottger, Dan Jurafsky, Tatsunori Hashimoto, and James Zou. 2024 · 2024
Closest in time.
Subobject-level image tokenization
Delong Chen, Samuel Cahyawijaya, Jianfeng Liu, Baoyuan Wang, and Pascale Fung. 2024 · 2024
Closest in time.
Self-improving robust preference optimization
Eugene Choi, Arash Ahmadian, Matthieu Geist, Oilvier Pietquin, and Mohammad Gheshlaghi Azar. 2024 · 2024
Closest in time.
Human feedback is not gold standard
Tom Hosking, Phil Blunsom, and Max Bartolo. 2024 · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Closest in time.
Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling
Dahyun Kim, Chanjun Park, Sanghoon Kim, Wonsung Lee, Wonho Song, Yunsu Kim, Hyeonwoo Kim, Yungi Kim, Hyeonju Lee, Jihoo Kim, Changbae Ahn, Seonghoon Yang, Sukyung Lee, Hyunbyung Park, Gyoungjin Gim, Mikyoung Cha, Hwalsuk Lee, and Sunghun Kim. 2024 · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al. 2024 · 2024
Closest in time.
Nomic embed: Training a reproducible long context text embedder
Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. 2024 · 2024
Closest in time.
From one to many: Expanding the scope of toxicity mitigation in language models
Luiza Pozzobon, Patrick Lewis, Sara Hooker, and Beyza Ermis. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024 · 2024
Closest in time.
To compress or not to compress—self-supervised learning and information theory: A review
Ravid Shwartz Ziv and Yann LeCun. 2024 · 2024
Closest in time.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. 2024 · 2024
Closest in time.
Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties
Taylor Sorensen, Liwei Jiang, Jena D Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, et al. 2024 · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Ahmet Ustun, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024 · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024 · 2024
Closest in time.
Heterogeneous value alignment evaluation for large language models
Zhaowei Zhang, Ceyao Zhang, Nian Liu, Siyuan Qi, Ziqi Rong, Song-Chun Zhu, Shuguang Cui, and Yaodong Yang. 2024 · 2024
Closest in time.