Fetching the paper…
Reading the bibliography…
While Large Language Models (LLMs) have seen widespread applications across numerous fields, their limited interpretability poses concerns regarding their safe operations from multiple aspects, e.g., truthfulness, robustness, and fairness.
Genetic k-means algorithm
K Krishna and M Narasimha Murty · 1999
Earlier work this paper cites.
Ensemble learning
Thomas G Dietterich et al · 2002
Earlier work this paper cites.
Greedy decoding for statistical machine translation in almost linear time
Ulrich Germann · 2003
Earlier work this paper cites.
An introduction to roc analysis
Tom Fawcett · 2006
Earlier work this paper cites.
Artificial general intelligence: concept, state of the art, and future prospects
Ben Goertzel · 2014
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al · 2014
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Learning with rejection
Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri · 2016
Earlier work this paper cites.
Deepxplore: Automated whitebox testing of deep learning systems
Kexin Pei, Yinzhi Cao, Junfeng Yang, and Suman Jana · 2017
Earlier work this paper cites.
Reluplex: An efficient smt solver for verifying deep neural networks
Guy Katz, Clark Barrett, David L Dill, Kyle Julian, and Mykel J Kochenderfer · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Safety verification of deep neural networks
Xiaowei Huang, Marta Kwiatkowska, Sen Wang, and Min Wu · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Robust online monitoring of signal temporal logic
Jyotirmoy V Deshmukh, Alexandre Donzé, Shromona Ghosh, Xiaoqing Jin, Garvit Juniwal, and Sanjit A Seshia · 2017
Earlier work this paper cites.
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv · 2017
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Earlier work this paper cites.
Deepgauge: Multi-granularity testing criteria for deep learning systems
Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chunyang Chen, Ting Su, Li Li, Yang Liu, et al · 2018
Earlier work this paper cites.
Ai2: Safety and robustness certification of neural networks with abstract interpretation
Timon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov, Swarat Chaudhuri, and Martin Vechev · 2018
Earlier work this paper cites.
Deeptest: Automated testing of deep-neural-network-driven autonomous cars
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray · 2018
Earlier work this paper cites.
Youcheng Sun, Xiaowei Huang, Daniel Kroening, James Sharp, Matthew Hill, and Rob Ashmore · 2018
Earlier work this paper cites.
Efficient formal safety analysis of neural networks
Shiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang, and Suman Jana · 2018
Earlier work this paper cites.
Planning and decision-making for autonomous vehicles
Wilko Schwarting, Javier Alonso-Mora, and Daniela Rus · 2018
Earlier work this paper cites.
Specification-based monitoring of cyber-physical systems: a survey on theory, tools and applications
Ezio Bartocci, Jyotirmoy Deshmukh, Alexandre Donzé, Georgios Fainekos, Oded Maler, Dejan Ničković, and Sriram Sankaranarayanan · 2018
Earlier work this paper cites.
Ensemble learning: A survey
Omer Sagi and Lior Rokach · 2018
Earlier work this paper cites.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin · 2018
Earlier work this paper cites.
Enhancing the reliability of out-of-distribution image detection in neural networks
Shiyu Liang, Yixuan Li, and R. Srikant · 2018
Earlier work this paper cites.
Adversarially learned one-class classifier for novelty detection
Mohammad Sabokrou, Mohammad Khalooei, Mahmood Fathy, and Ehsan Adeli · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
A call for clarity in reporting bleu scores
Matt Post · 2018
Earlier work this paper cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave · 2019
Earlier work this paper cites.
Guiding deep learning system testing using surprise adequacy
Jinhan Kim, Robert Feldt, and Shin Yoo · 2019
Earlier work this paper cites.
Outside the box: Abstraction-based monitoring of neural networks
Thomas A Henzinger, Anna Lukina, and Christian Schilling · 2019
Earlier work this paper cites.
Runtime monitoring neuron activation patterns
Chih-Hong Cheng, Georg Nührenberg, and Hirotoshi Yasuoka · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Adversarial sample detection for deep neural network through model mutation testing
Jingyi Wang, Guoliang Dong, Jun Sun, Xinyu Wang, and Peixin Zhang · 2019
Earlier work this paper cites.
A survey of challenges for runtime verification from advanced application domains (beyond software)
César Sánchez, Gerardo Schneider, Wolfgang Ahrendt, Ezio Bartocci, Domenico Bianculli, Christian Colombo, Yliès Falcone, Adrian Francalanza, Sran Krstić, Joao M Lourenço, et al · 2019
Earlier work this paper cites.
A roadmap toward the resilient internet of things for cyber-physical systems
Denise Ratasich, Faiq Khalid, Florian Geissler, Radu Grosu, Muhammad Shafique, and Ezio Bartocci · 2019
Earlier work this paper cites.
A safety analysis method for perceptual components in automated driving
Rick Salay, Matt Angus, and Krzysztof Czarnecki · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Software engineering for machine learning: A case study
Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann · 2019
Earlier work this paper cites.
Secure deep learning engineering: A road towards quality assurance of intelligent systems
Yang Liu, Lei Ma, and Jianjun Zhao · 2019
Earlier work this paper cites.
Selectivenet: A deep neural network with an integrated reject option
Yonatan Geifman and Ran El-Yaniv · 2019
Earlier work this paper cites.
Deep anomaly detection with outlier exposure
Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Testing machine learning based systems: a systematic mapping
Vincenzo Riccio, Gunel Jahangirova, Andrea Stocco, Nargiz Humbatova, Michael Weiss, and Paolo Tonella · 2020
Earlier work this paper cites.
Large-scale machine learning systems in real-world industrial settings: A review of challenges and solutions
Lucy Ellen Lwakatare, Aiswarya Raj, Ivica Crnkovic, Jan Bosch, and Helena Holmström Olsson · 2020
Earlier work this paper cites.
Deepgini: prioritizing massive tests to enhance the robustness of deep neural networks
Yang Feng, Qingkai Shi, Xinyu Gao, Jun Wan, Chunrong Fang, and Zhenyu Chen · 2020
Earlier work this paper cites.
Misbehaviour prediction for autonomous driving systems
Andrea Stocco, Michael Weiss, Marco Calzana, and Paolo Tonella · 2020
Earlier work this paper cites.
A survey of safety and trustworthiness of deep neural networks: Verification, testing, adversarial attack and defence, and interpretability
Xiaowei Huang, Daniel Kroening, Wenjie Ruan, James Sharp, Youcheng Sun, Emese Thamo, Min Wu, and Xinping Yi · 2020
Earlier work this paper cites.
Safeml: safety monitoring of machine learning classifiers through statistical difference measures
Koorosh Aslansefat, Ioannis Sorokos, Declan Whiting, Ramin Tavakoli Kolagari, and Yiannis Papadopoulos · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Aligning ai with shared human values
Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
On testing machine learning programs
Houssem Ben Braiek and Foutse Khomh · 2020
Cited alongside, same era.
Nnv: the neural network verification tool for deep neural networks and learning-enabled cyber-physical systems
Hoang-Dung Tran, Xiaodong Yang, Diego Manzanas Lopez, Patrick Musau, Luan Viet Nguyen, Weiming Xiang, Stanley Bak, and Taylor T Johnson · 2020
Cited alongside, same era.
Adoption and effects of software engineering best practices in machine learning
Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser · 2020
Cited alongside, same era.
Generalized odin: Detecting out-of-distribution image without learning from out-of-distribution data
Yen-Chang Hsu, Yilin Shen, Hongxia Jin, and Zsolt Kira · 2020
Cited alongside, same era.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Later among the works it cites.
Weakly supervised detection of hallucinations in llm activations
Miriam Rateike, Celia Cintas, John Wamburu, Tanya Akumu, and Skyler Speakman · 2023
Later among the works it cites.
When and why test generators for deep learning produce invalid inputs: an empirical study
Vincenzo Riccio and Paolo Tonella · 2023
Later among the works it cites.
Repairing dnn architecture: Are we there yet?
Jinhan Kim, Nargiz Humbatova, Gunel Jahangirova, Paolo Tonella, and Shin Yoo · 2023
Later among the works it cites.
Smarla: A safety monitoring approach for deep reinforcement learning agents
Amirhossein Zolfagharian, Manel Abdellatif, Lionel C Briand, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Energy-based out-of-distribution detection
Weitang Liu, Xiaoyun Wang, John Owens, and Yixuan Li · 2020
Cited alongside, same era.
Cats are not fish: Deep learning testing calls for out-of-distribution awareness
David Berend, Xiaofei Xie, Lei Ma, Lingjun Zhou, Yang Liu, Chi Xu, and Jianjun Zhao · 2020
Cited alongside, same era.
The k-means algorithm: A comprehensive survey and performance evaluation
Mohiuddin Ahmed, Raihan Seraj, and Syed Mohammed Shamsul Islam · 2020
Cited alongside, same era.
A token-level reference-free hallucination detection benchmark for free-form text generation
Tianyu Liu, Yizhe Zhang, Chris Brockett, Yi Mao, Zhifang Sui, Weizhu Chen, and Bill Dolan · 2021
Cited alongside, same era.
Probing toxic content in large pre-trained language models
Nedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, and Dit-Yan Yeung · 2021
Cited alongside, same era.
Robot: Robustness-oriented testing for deep learning systems
Jingyi Wang, Jialuo Chen, Youcheng Sun, Xingjun Ma, Dongxia Wang, Jun Sun, and Peng Cheng · 2021
Cited alongside, same era.
Prioritizing test inputs for deep neural networks via mutation analysis
Zan Wang, Hanmo You, Junjie Chen, Yingyi Zhang, Xuyuan Dong, and Wenbin Zhang · 2021
Cited alongside, same era.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Later among the works it cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, et al · 2023
Later among the works it cites.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Zhen Lin, Shubhendu Trivedi, and Jimeng Sun · 2023
Later among the works it cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar · 2023
Later among the works it cites.
The internal state of an llm knows when it’s lying, 2023
Amos Azaria and Tom Mitchell · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
Gpt 3.5, 2023
OpenAI · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Out-of-distribution detection is not all you need
Joris Guérin, Kevin Delmas, Raul Ferreira, and Jérémie Guiochet · 2023
Later among the works it cites.
Mosaic: Model-based safety analysis framework for ai-enabled cyber-physical systems
Xuan Xie, Jiayang Song, Zhehua Zhou, Fuyuan Zhang, and Lei Ma · 2023
Later among the works it cites.
Jiayang Song, Zhehua Zhou, Jiawei Liu, Chunrong Fang, Zhan Shu, and Lei Ma · 2023
Later among the works it cites.
Line: Out-of-distribution detection by leveraging important neurons
Yong Hyun Ahn, Gyeong-Moon Park, and Seong Tae Kim · 2023
Later among the works it cites.
Metamorphic runtime monitoring of autonomous driving systems
Jon Ayerdi, Asier Iriarte, Pablo Valle, Ibai Roman, Miren Illarramendi, and Aitor Arrieta · 2023
Later among the works it cites.
Online shielding for reinforcement learning
Bettina Könighofer, Julian Rudolf, Alexander Palmisano, Martin Tappler, and Roderick Bloem · 2023
Later among the works it cites.
Systematic rectification of language models via dead-end analysis
Meng Cao, Mehdi Fatemi, Jackie CK Cheung, and Samira Shabanian · 2023
Later among the works it cites.
Controlled decoding from language models
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanping Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, et al · 2023
Later among the works it cites.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2023
Later among the works it cites.
Summarization is (almost) dead
Xiao Pu, Mingqi Gao, and Xiaojun Wan · 2023
Later among the works it cites.
Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models
Caroline Lemieux, Jeevana Priya Inala, Shuvendu K Lahiri, and Siddhartha Sen · 2023
Later among the works it cites.
Fill in the blank: Context-aware automated text input generation for mobile gui testing
Zhe Liu, Chunyang Chen, Junjie Wang, Xing Che, Yuekai Huang, Jun Hu, and Qing Wang · 2023
Later among the works it cites.
An analysis of the automatic bug fixing performance of chatgpt
Dominik Sobania, Martin Briesch, Carol Hanna, and Justyna Petke · 2023
Later among the works it cites.
Codegen: An open large language model for code with multi-turn program synthesis, 2023
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2023
Later among the works it cites.
Large language models for software engineering: Survey and open problems
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Later among the works it cites.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou · 2023
Later among the works it cites.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Later among the works it cites.
A multitask, multilingual, multimodal evaluation of chatgpt on reasoning, hallucination, and interactivity, 2023
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, and Pascale Fung · 2023
Later among the works it cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda · 2023
Later among the works it cites.
Is chatgpt a general-purpose natural language processing task solver?
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang · 2023
Later among the works it cites.
Chain-of-thought hub: A continuous effort to measure large language models’ reasoning performance
Yao Fu, Litu Ou, Mingyu Chen, Yuhao Wan, Hao Peng, and Tushar Khot · 2023
Later among the works it cites.
Classeval: A manually-crafted benchmark for evaluating llms on class-level code generation
Xueying Du, Mingwei Liu, Kaixin Wang, Hanlin Wang, Junwei Liu, Yixuan Chen, Jiayi Feng, Chaofeng Sha, Xin Peng, and Yiling Lou · 2023
Later among the works it cites.
Genaipabench: A benchmark for generative ai-based privacy assistants
Aamir Hamid, Hemanth Reddy Samidi, Tim Finin, Primal Pappachan, and Roberto Yus · 2023
Later among the works it cites.
B.c. lawyer who used fake, ai-generated cases faces law society probe, possible costs, 2024
Global News · 2024
Closest in time.
Detectors for safe and reliable llms: Implementations, uses, and limitations
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor, Ioana Baldini, Sara E Berger, Bishwaranjan Bhattacharjee, Djallel Bouneffouf, Subhajit Chaudhury, Pin-Yu Chen, Lamogha Chiazor, et al · 2024
Closest in time.
Shieldlm: Empowering llms as aligned, customizable and explainable safety detectors
Zhexin Zhang, Yida Lu, Jingyuan Ma, Di Zhang, Rui Li, Pei Ke, Hao Sun, Lei Sha, Zhifang Sui, Hongning Wang, et al · 2024
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Miao Xiong, Zhiyuan Hu, Xinyang Lu, YIFEI LI, Jie Fu, Junxian He, and Bryan Hooi · 2024
Closest in time.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang · 2024
Closest in time.
Perspective api, 2024
Google · 2024
Closest in time.
Sharegpt, 2024
ShareGPT · 2024
Closest in time.
Reality bites: Assessing the realism of driving scenarios with large language models
Jiahui Wu, Chengjie Lu, Aitor Arrieta, Tao Yue, and Shaukat Ali · 2024
Closest in time.
A data-driven measure of relative uncertainty for misclassification detection
Eduardo Dadalto Câmara Gomes, Marco Romanelli, Georg Pichler, and Pablo Piantanida · 2024
Closest in time.
Reinforcement learning with ensemble model predictive safety certification
Sven Gronauer, Tom Haider, Felippe Schmoeller da Roza, and Klaus Diepold · 2024
Closest in time.
Large language models for software engineering: A systematic literature review, 2024
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2024
Closest in time.
Software testing with large language models: Survey, landscape, and vision
Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang · 2024
Closest in time.
Hallucination is inevitable: An innate limitation of large language models
Ziwei Xu, Sanjay Jain, and Mohan Kankanhalli · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2024
Closest in time.
CRITIC: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Nan Duan, and Weizhu Chen · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Closest in time.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner · 2024
Closest in time.