Fetching the paper…
Reading the bibliography…
In this paper, we present a very first study to investigate trust and ethical implications of on-device artificial intelligence (AI), focusing on small language models (SLMs) amenable for personal devices like smartphones.
Learning fair representations
Rich Zemel, Yu Wu, Kevin Swersky, Toni Pitassi, and Cynthia Dwork · 2013
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher
Guangda Ji and Zhanxing Zhu · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Earlier work this paper cites.
Adversarial glue: A multi-task benchmark for robustness evaluation of language models
Boxin Wang, Chejian Xu, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Hassan Awadallah, and Bo Li · 2021
Earlier work this paper cites.
Textflint: Unified multilingual robustness evaluation toolkit for natural language processing
Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Studying the usage of text-to-text transfer transformer to support code-related tasks
Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota · 2021
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Mirac Suzgun, Nathan Scales, Nathanael Schärli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, et al · 2022
Earlier work this paper cites.
Inherent tradeoffs in learning fair representations
Han Zhao and Geoffrey J Gordon · 2022
Earlier work this paper cites.
Certifying some distributional fairness with subpopulation decomposition
Mintong Kang, Linyi Li, Maurice Weber, Yang Liu, Ce Zhang, and Bo Li · 2022
Earlier work this paper cites.
Are large pre-trained language models leaking your personal information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang · 2022
Earlier work this paper cites.
Xraygpt: Chest radiographs summarization using medical vision-language models
Omkar Thawkar, Abdelrahman Shaker, Sahal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, and Fahad Shahbaz Khan · 2023
Earlier work this paper cites.
Unlocking the power of chatgpt: A framework for applying generative ai in education
Jiahong Su and Weipeng Yang · 2023
Earlier work this paper cites.
Artificial intelligence in higher education: the state of the field
Helen Crompton and Diane Burke · 2023
Earlier work this paper cites.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Earlier work this paper cites.
Fingpt: Open-source financial large language models
Hongyang Yang, Xiao-Yang Liu, and Christina Dan Wang · 2023
Earlier work this paper cites.
The possibility of applying chatgpt (ai) for calculations in mechanical engineering
Dragi Tiro · 2023
Earlier work this paper cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al · 2023
Earlier work this paper cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al · 2023
Earlier work this paper cites.
The case for 4-bit precision: k-bit inference scaling laws
Tim Dettmers and Luke Zettlemoyer · 2023
Cited alongside, same era.
Do-not-answer: A dataset for evaluating safeguards in llms
Yuxia Wang, Haonan Li, Xudong Han, Preslav Nakov, and Timothy Baldwin · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Cited alongside, same era.
Beyond the safeguards: Exploring the security risks of chatgpt
Erik Derner and Kristina Batistič · 2023
Cited alongside, same era.
Fast as chita: Neural network pruning with combinatorial optimization
Riade Benbaki, Wenyu Chen, Xiang Meng, Hussein Hazimeh, Natalia Ponomareva, Zhe Zhao, and Rahul Mazumder · 2023
https://store.google.com/intl/en/ideas/articles/gemini-nano-google-pixel/
Gemini Nano Multimodal Capabilities on Pixel Phones · 2024
Closest in time.
https://www.samsung.com/us/smartphones/galaxy-s24/galaxy-ai/
Galaxy AI · 2024
Closest in time.
https://www.reddit.com/r/Instagram/comments/1d4vm1u/instagram_prompts/
Instagram Prompts · 2024
Closest in time.
https://blog.google/technology/developers/gemma-open-models/
Gemma: Introducing new state-of-the-art open models · 2024
Closest in time.
Phi-2: The surprising power of small language models
Alyssa Hughes · 2024
Closest in time.
https://www.together.ai/blog/redpajama-models-v1
Releasing 3B and 7B RedPajama-INCITE family of models · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Cited alongside, same era.
A comprehensive capability analysis of gpt-3 and gpt-3.5 series models
Junjie Ye, Xuanting Chen, Nuo Xu, Can Zu, Zekai Shao, Shichun Liu, Yuhan Cui, Zeyang Zhou, Chao Gong, Yang Shen, et al · 2023
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
https://www.embedl.com/knowledge/the-big-story-of-ai-in-2023-llms
The Big Story of AI in 2023: LLMs · 2024
Cited alongside, same era.
https://www.khanmigo.ai/teachers
Khanmigo: Free, AI-powered teacher assistant · 2024
Cited alongside, same era.
https://ai.google/responsibility/principles/
Google AI Principles · 2024
Closest in time.
https://www.bing.com/new/termsofuse
Copilot code of conduct · 2024
Closest in time.
From chatbots to phishbots?: Phishing scam generation in commercial large language models
Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, and Shirin Nilizadeh · 2024
Closest in time.
https://www.qualcomm.com/news/onq/2023/12/optimizing-generative-ai-for-edge-devices
Optimizing generative AI for edge devices · 2024
Closest in time.
https://pytorch.org/docs/stable/quantization.html
Quantization in PyTorch · 2024
Closest in time.
https://www.tensorflow.org/lite/performance/post_training_quantization
Post-training quantization | TensorFlow Lite · 2024
Closest in time.
A comprehensive evaluation of quantization strategies for large language models
Renren Jin, Jiangcun Du, Wuwei Huang, Wei Liu, Jian Luan, Bin Wang, and Deyi Xiong · 2024
Closest in time.
https://github.com/AI-secure/DecodingTrust/tree/main/data/fairness/fairness_data
Fairness Data - DecodingTrust · 2024
Closest in time.
https://developers.google.com/machine-learning/glossary#accuracy
Machine Learning Glossary · 2024
Closest in time.
https://github.com/AI-secure/DecodingTrust
GitHub - AI-secure/DecodingTrust · 2024
Closest in time.
https://github.com/mlc-ai/mlc-llm/tree/main/android
Github: Mlc chat android · 2024
Closest in time.
https://www.qualcomm.com/products/mobile/snapdragon/smartphones/snapdragon-8-series-mobile-platforms/snapdragon-8-gen-3-mobile-platform
Snapdragon 8 Gen-3 Mobile Platform · 2024
Closest in time.
https://github.com/mlc-ai/binary-mlc-llm-libs/releases/download/Android-06042024/mlc-chat.apk
MLC Chat apk file · 2024
Closest in time.
https://www.reddit.com/r/ChatGPTJailbreak/comments/12oxynu/how_to_make_malicious_codes_with_chatgpt/
Malicious Code Prompt · 2024
Closest in time.
Scieval: A multi-level large language model evaluation benchmark for scientific research
Liangtai Sun, Yang Han, Zihan Zhao, Da Ma, Zhennan Shen, Baocai Chen, Lu Chen, and Kai Yu · 2024
Closest in time.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al · 2024
Closest in time.
Sciassess: Benchmarking llm proficiency in scientific literature analysis
Hengxing Cai, Xiaochen Cai, Junhan Chang, Sihang Li, Lin Yao, Changxin Wang, Zhifeng Gao, Yongge Li, Mujie Lin, Shuwen Yang, et al · 2024
Closest in time.
Mercury: An efficiency benchmark for llm code synthesis
Mingzhe Du, Anh Tuan Luu, Bin Ji, and See-Kiong Ng · 2024
Closest in time.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al · 2024
Closest in time.
Pub: A pragmatics understanding benchmark for assessing llms’ pragmatics capabilities
Settaluri Lakshmi Sravanthi, Meet Doshi, Tankala Pavan Kalyan, Rudra Murthy, Pushpak Bhattacharyya, and Raj Dabre · 2024
Closest in time.