Fetching the paper…
Reading the bibliography…
The proliferation of large language models (LLMs) requires robust evaluation of their alignment with local values and ethical standards, especially as existing benchmarks often reflect the cultural, legal, and ideological values of their creators.
Z. Jiang, A. Anastasopoulos, J. Araki, H. Ding, and G. Neubig, “X-factr: Multilingual factual knowledge retrieval from pretrained language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 5943–5959
2020
Earlier work this paper cites.
S. Fazelpour and Z. C. Lipton, “Algorithmic fairness from a non-ideal perspective,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , 2020, pp. 57–63
2020
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
J. Fries, L. Weber, N. Seelam, G. Altay, D. Datta, S. Garda, S. Kang, R. Su, W. Kusa, S. Cahyawijaya et al. , “Bigbio: A framework for data-centric biomedical natural language processing,” Advances in Neural Information Processing Systems , vol. 35, pp. 25 792–25 806, 2022
2022
Earlier work this paper cites.
J. Zhu, Q. Dai, L. Su, R. Ma, J. Liu, G. Cai, X. Xiao, and R. Zhang, “Bars: towards open benchmarking for recommender systems,” in Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2022, pp. 2912–2923
2022
Earlier work this paper cites.
Z. Talat, A. Névéol, S. Biderman, M. Clinciu, M. Dey, S. Longpre, S. Luccioni, M. Masoud, M. Mitchell, D. Radev et al. , “You reap what you sow: On the challenges of bias evaluation under multilingual settings,” in Proceedings of BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models , 2022, pp. 26–41
2022
Earlier work this paper cites.
P. Schramowski, C. Turan, N. Andersen, C. A. Rothkopf, and K. Kersting, “Large pre-trained language models contain human-like biases of what is right and wrong to do,” Nature Machine Intelligence , vol. 4, no. 3, pp. 258–268, 2022
2022
Earlier work this paper cites.
T. R. McIntosh, T. Liu, T. Susnjak, P. Watters, A. Ng, and M. N. Halgamuge, “A culturally sensitive test to evaluate nuanced gpt hallucination,” IEEE Transactions on Artificial Intelligence , 2023
2023
Earlier work this paper cites.
T. Guo, B. Nan, Z. Liang, Z. Guo, N. Chawla, O. Wiest, X. Zhang et al. , “What can large language models do in chemistry? a comprehensive benchmark on eight tasks,” Advances in Neural Information Processing Systems , vol. 36, pp. 59 662–59 688, 2023
2023
Earlier work this paper cites.
J. Yan, V. Yadav, S. Li, L. Chen, Z. Tang, H. Wang, V. Srinivasan, X. Ren, and H. Jin, “Backdooring instruction-tuned large language models with virtual prompt injection,” in NeurIPS 2023 Workshop on Backdoors in Deep Learning-The Good, the Bad, and the Ugly , 2023
2023
Earlier work this paper cites.
T. R. McIntosh, T. Susnjak, T. Liu, P. Watters, A. Ng, and M. N. Halgamuge, “A game-theoretic approach to containing artificial general intelligence: Insights from highly autonomous aggressive malware,” IEEE Transactions on Artificial Intelligence , 2024
2024
Earlier work this paper cites.
2024
Cited alongside, same era.
L. Yuan, Y. Chen, G. Cui, H. Gao, F. Zou, X. Cheng, H. Ji, Z. Liu, and M. Sun, “Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, S. Zhong, B. Yin, and X. Hu, “Harnessing the power of llms in practice: A survey on chatgpt and beyond,” ACM Transactions on Knowledge Discovery from Data , vol. 18, no. 6, pp. 1–32, 2024
2024
Cited alongside, same era.
W. Zhang, M. Aljunied, C. Gao, Y. K. Chia, and L. Bing, “M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, B. Chen, R. Sun, Y. Wang, and Y. Yang, “Beavertails: Towards improved safety alignment of llm via a human-preference dataset,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
T. R. McIntosh, T. Susnjak, T. Liu, P. Watters, and M. N. Halgamuge, “The inadequacy of reinforcement learning from human feedback-radicalizing large language models via semantic vulnerabilities,” IEEE Transactions on Cognitive and Developmental Systems , 2024
2024
Closest in time.
Z. Sun, Y. Shen, Q. Zhou, H. Zhang, Z. Chen, D. Cox, Y. Yang, and C. Gan, “Principle-driven self-alignment of language models from scratch with minimal human supervision,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Zhang, G. Shen, G. Tao, S. Cheng, and X. Zhang, “On large language models’ resilience to coercive interrogation,” in 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 2024, pp. 252–252
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
L. Ruis, A. Khan, S. Biderman, S. Hooker, T. Rocktäschel, and E. Grefenstette, “The goldilocks of pragmatic understanding: Fine-tuning strategy matters for implicature resolution by llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
Z. Chen, H. Mao, H. Li, W. Jin, H. Wen, X. Wei, S. Wang, D. Yin, W. Fan, H. Liu et al. , “Exploring the potential of large language models (llms) in learning on graphs,” ACM SIGKDD Explorations Newsletter , vol. 25, no. 2, pp. 42–61, 2024
2024
Cited alongside, same era.
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel, “Scaling data-constrained language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y. Belkada, P. Samygin, and C. A. Raffel, “Distributed inference and fine-tuning of large language models over the internet,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
L. Fan, D. Krishnan, P. Isola, D. Katabi, and Y. Tian, “Improving clip training with language rewrites,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang et al. , “Self-refine: Iterative refinement with self-feedback,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
Y. Niu, Y. Pu, Z. Yang, X. Li, T. Zhou, J. Ren, S. Hu, H. Li, and Y. Liu, “Lightzero: A unified benchmark for monte carlo tree search in general sequential decision scenarios,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
Z. Jiang, H. Lin, Y. Zhong, Q. Huang, Y. Chen, Z. Zhang, Y. Peng, X. Li, C. Xie, S. Nong et al. , “ { \{ MegaScale } \} : Scaling large language model training to more than 10,000 { \{ GPUs } \} ,” in 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) , 2024, pp. 745–760
Cited in the paper.
2024
Closest in time.
N. Kiehne, A. Ljapunov, M. Bätje, and W.-T. Balke, “Analyzing effects of learning downstream tasks on moral bias in large language models,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , 2024, pp. 904–923
2024
Closest in time.
N. Scherrer, C. Shi, A. Feder, and D. Blei, “Evaluating the moral beliefs encoded in llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
J. Lee, S. Kim, S. Won, J. Lee, M. Ghassemi, J. Thorne, J. Choi, O.-K. Kwon, and E. Choi, “Visalign: Dataset for measuring the alignment between ai and humans in visual perception,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Zhao, W. Fan, J. Li, Y. Liu, X. Mei, Y. Wang, Z. Wen, F. Wang, X. Zhao, J. Tang et al. , “Recommender systems in the era of large language models (llms),” IEEE Transactions on Knowledge and Data Engineering , 2024
2024
Closest in time.
Y. Zhang, J. Sun, L. Feng, C. Yao, M. Fan, L. Zhang, Q. Wang, X. Geng, and Y. Rui, “See widely, think wisely: Toward designing a generative multi-agent system to burst filter bubbles,” in Proceedings of the CHI Conference on Human Factors in Computing Systems , 2024, pp. 1–24
2024
Closest in time.
C. Ziems, W. Held, O. Shaikh, J. Chen, Z. Zhang, and D. Yang, “Can large language models transform computational social science?” Computational Linguistics , vol. 50, no. 1, pp. 237–291, 2024
2024
Closest in time.