Fetching the paper…
Reading the bibliography…
When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framework for systematically evaluating distribution shifts in LLMs decision-making processes, where they are given control of mechanisms governed by pre-defined rules.
On information and sufficiency
Kullback, S. and Leibler, R. A · 1951
Earlier work this paper cites.
Comparing Distributions: The Two-Sample Anderson-Darling Test as an Alternative to the Kolmogorov-Smirnoff Test
Engmann, S. and Cousineau, D · 2011
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Caliskan, A., Bryson, J. J., and Narayanan, A · 2017
Earlier work this paper cites.
Persistent anti-muslim bias in large language models, 2021
Abid, A., Farooqi, M., and Zou, J · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
Stereotyping Norwegian salmon: An inventory of pitfalls in fairness benchmark datasets
Blodgett, S. L., Lopez, G., Olteanu, A., Sim, R., and Wallach, H · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models, 2021
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., and Gabriel, I · 2021
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods, 2022
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Instructed to bias: instruction-tuned language models exhibit emergent cognitive bias
Itzhak, I., Stanovsky, G., Rosenfeld, N., and Belinkov, Y · 2023
Cited alongside, same era.
Towards a unified agent with foundation models, 2023
Palo, N. D., Byravan, A., Hasenclever, L., Wulfmeier, M., Heess, N., and Riedmiller, M · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency, 2023
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., Goel, S., Li, N., Byun, M. J., Wang, Z., Mallen, A., Basart, S., Koyejo, S., Song, D., Fredrikson, M., Kolter, J. Z., and Hendrycks, D · 2023
Later among the works it cites.
Agent ai: Surveying the horizons of multimodal interaction, 2024
Durante, Z., Huang, Q., Wake, N., Gong, R., Park, J. S., Sarkar, B., Taori, R., Noda, Y., Terzopoulos, D., Choi, Y., Ikeuchi, K., Vo, H., Fei-Fei, L., and Gao, J · 2024
Closest in time.
Hagendorff, T., Dasgupta, I., Binz, M., Chan, S. C. Y., Lampinen, A., Wang, J. X., Akata, Z., and Schulz, E · 2024
Closest in time.
Lamparth, M., Corso, A., Ganz, J., Mastro, O. S., Schneider, J., and Trinkunas, H · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vezhnevets, A. S., Agapiou, J. P., Aharon, A., Ziv, R., Matyas, J., Duéñez-Guzmán, E. A., Cunningham, W. A., Osindero, S., Karmon, D., and Leibo, J. Z · 2023
Cited alongside, same era.
Webshop: Towards scalable real-world web interaction with grounded language agents, 2023
Yao, S., Chen, H., Yang, J., and Narasimhan, K · 2023
Cited alongside, same era.
X. on the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling
Pearson, K
Cited in the paper.
Do llms exhibit human-like response biases? a case study in survey design
Tjuatja, L., Chen, V., Wu, T., Talwalkwar, A., and Neubig, G · 2024
Closest in time.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., and Wen, J · 2095
Closest in time.