Fetching the paper…
Reading the bibliography…
Guardrail, an emerging mechanism designed to ensure that large language models (LLMs) align with human values by moderating harmful or toxic responses, requires a sociotechnical approach in their design.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
A morality based on trust: Some reflections on japanese morality
Y. Yamamoto · 1990
Earlier work this paper cites.
Perceived risk, trust, and democracy
P. Slovic · 1993
Earlier work this paper cites.
Familiarity, confidence, trust: Problems and alternatives
N. Luhmann et al · 2000
Earlier work this paper cites.
Social trust: A cognitive approach
R. Falcone and C. Castelfranchi · 2001
Earlier work this paper cites.
The beta reputation system
A. Josang and R. Ismail · 2002
Earlier work this paper cites.
The role of trust in information science and technology
S. Marsh and M. R. Dibben · 2003
Earlier work this paper cites.
Task delegation using experience-based multi-dimensional trust
N. Griffiths · 2005
Earlier work this paper cites.
A survey of trust and reputation systems for online service provision
A. Jøsang, R. Ismail, and C. Boyd · 2007
Earlier work this paper cites.
Experience-based service provider selection in agent-mediated e-commerce
M. Şensoy, F. C. Pembe, H. Zırtıloğlu, P. Yolum, and A. Bener · 2007
Earlier work this paper cites.
Computing confidence values: Does trust dynamics matter?
J. Urbano, A. P. Rocha, and E. Oliveira · 2009
Earlier work this paper cites.
Continuous and transparent user identity verification for secure internet services
A. Ceccarelli, L. Montecchi, F. Brancati, P. Lollini, A. Marguglio, and A. Bondavalli · 2013
Earlier work this paper cites.
How trust is formed in online health communities: a process perspective
H. Fan, R. Lederman, S. P. Smith, and S. Chang · 2014
Earlier work this paper cites.
A survey on trust modeling
J.-H. Cho, K. Chan, and S. Adali · 2015
Earlier work this paper cites.
Social trust: a major challenge for the future of autonomous systems
M. Lahijanian and M. Kwiatkowska · 2016
Earlier work this paper cites.
Trust-based multi-robot symbolic motion planning with a human-in-the-loop
Y. Wang, L. R. Humphrey, Z. Liao, and H. Zheng · 2018
Cited alongside, same era.
Reasoning about cognitive trust in stochastic multiagent systems
X. Huang, M. Kwiatkowska, and M. Olejnik · 2019
Cited alongside, same era.
Adaptive trust calibration for human-ai collaboration
K. Okamura and S. Yamada · 2020
Cited alongside, same era.
An adaptive trust model based on recommendation filtering algorithm for the internet of things systems
G. Chen, F. Zeng, J. Zhang, T. Lu, J. Shen, and W. Shu · 2021
Cited alongside, same era.
What makes good in-context examples for gpt- 3 3 ?, 2021
J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen · 2021
Cited alongside, same era.
Combining direct trust and indirect trust in multi-agent systems
E. Parhizkar, M. H. Nikravan, R. C. Holte, and S. Zilles · 2021
Universal and transferable adversarial attacks on aligned language models, 2023
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Later among the works it cites.
Position: Social choice should guide AI alignment in dealing with diverse human feedback
V. Conitzer, R. Freedman, J. Heitzig, W. H. Holliday, B. M. Jacobs, N. Lambert, M. Mossé, E. Pacuit, S. Russell, H. Schoelkopf, E. Tewolde, and W. S. Zwicker · 2024
Closest in time.
Merriam-webster dictionary
M.-W. Dictionary · 2024
Closest in time.
A survey on rag meeting llms: Towards retrieval-augmented large language models
W. Fan, Y. Ding, L. Ning, S. Wang, H. Li, D. Yin, T.-S. Chua, and Q. Li · 2024
Closest in time.
I. Frisch and M. Giulianelli · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust and risk perception: A critical review of the literature
M. Siegrist · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh · 2021
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa · 2023
Cited alongside, same era.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation, 2023
Z. Lin, Z. Wang, Y. Tong, Y. Wang, Y. Guo, Y. Wang, and J. Shang · 2023
Cited alongside, same era.
Guardrails ai
S. Rajpal · 2023
Cited alongside, same era.
A survey of safety and trustworthiness of large language models through the lens of verification and validation
X. Huang, W. Ruan, W. Huang, G. Jin, Y. Dong, C. Wu, S. Bensalem, R. Mu, Y. Qi, X. Zhao, et al · 2024
Closest in time.
In-context learning learns label relationships but is not conventional learning
J. Kossen, Y. Gal, and T. Rainforth · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks · 2024
Closest in time.
Evidence-backed fact checking using rag and few-shot in-context learning with llms
R. Singal, P. Patwa, P. Patwa, A. Chadha, and A. Das · 2024
Closest in time.
The role of llms in sustainable smart cities: Applications, challenges, and future directions
A. Ullah, G. Qi, S. Hussain, I. Ullah, and Z. Ali · 2024
Closest in time.
Plug in the safety chip: Enforcing constraints for llm-driven robot agents
Z. Yang, S. S. Raman, A. Shah, and S. Tellex · 2024
Closest in time.
X. Zeng, F. Song, and A. Liu · 2024
Closest in time.
Position: Towards a responsible llm-empowered multi-agent systems, 2025
J. Hu, Y. Dong, S. Ao, Z. Li, B. Wang, L. Singh, G. Cheng, S. D. Ramchurn, and X. Huang · 2025
Closest in time.
Why ai is weird and shouldn’t be this way: Towards ai for everyone, with everyone, by everyone
R. Mihalcea, O. Ignat, L. Bai, A. Borah, L. Chiruzzo, Z. Jin, C. Kwizera, J. Nwatu, S. Poria, and T. Solorio · 2025
Closest in time.