Fetching the paper…
Reading the bibliography…
We study privacy leakage in the reasoning traces of large reasoning models used as personal agents.
Controlling the false discovery rate: A practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Privacy as contextual integrity
Helen Nissenbaum. 2004 · 2004
Earlier work this paper cites.
Will we run out of data? limits of llm scaling based on human-generated data
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. 2022 · 2022
Earlier work this paper cites.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2023 · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Decodingtrust: A comprehensive assessment of trustworthiness in GPT models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2023 · 2023
Earlier work this paper cites.
System 2 attention (is something you might need too)
Jason Weston and Sainbayar Sukhbaatar. 2023 · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R. Narasimhan, and Yuan Cao. 2023 · 2023
Earlier work this paper cites.
TextObfuscator: Making pre-trained language model a privacy protector via obfuscating word representations
Xin Zhou, Yi Lu, Ruotian Ma, Tao Gui, Yuran Wang, Yong Ding, Yibo Zhang, Qi Zhang, and Xuanjing Huang. 2023 · 2023
Earlier work this paper cites.
Airgapagent: Protecting privacy-conscious conversational agents
Eugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage. 2024 · 2024
Earlier work this paper cites.
Improved neural word segmentation for standard Tibetan
Collin J. Brown. 2024 · 2024
Earlier work this paper cites.
Ci-bench: Benchmarking contextual integrity of ai assistants on synthetic data
Zhao Cheng, Diane Wan, Matthew Abueg, Sahra Ghalebikesabi, Ren Yi, Eugene Bagdasarian, Borja Balle, Stefan Mellem, and Shawn O’Banion. 2024 · 2024
Earlier work this paper cites.
Whispers in the machine: Confidentiality in llm-integrated systems
Jonathan Evertz, Merlin Chlosta, Lea Schönherr, and Thorsten Eisenhofer. 2024 · 2024
Earlier work this paper cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024 · 2024
Earlier work this paper cites.
Yukun Huang, Sanxing Chen, Hongyi Cai, and Bhuwan Dhingra. 2024 · 2024
Cited alongside, same era.
VisualWebArena: Evaluating multimodal agents on realistic visual web tasks
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024 · 2024
Cited alongside, same era.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, Rui Kong, Yile Wang, Hanfei Geng, Jian Luan, Xuefeng Jin, Zilong Ye, Guanjing Xiong, Fan Zhang, Xiang Li, and 6 others. 2024 · 2024
Cited alongside, same era.
Can llms keep a secret? testing privacy implications of language models via contextual integrity theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2024 · 2024
Cited alongside, same era.
Safety tax: Safety alignment makes your large reasoning models less reasonable
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, Zachary Yahn, Yichang Xu, and Ling Liu. 2025 · 2025
Closest in time.
Test-time computing: from system-1 thinking to system-2 thinking
Yixin Ji, Juntao Li, Hai Ye, Kaixin Wu, Jia Xu, Linjian Mo, and Min Zhang. 2025 · 2025
Closest in time.
Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. 2025 · 2025
Closest in time.
LLM Reasoning: from OpenAI O1 to DeepSeek R1
Wang Jiaqi, Li Xinliang, Liu Zhengliang, Wu Zihao, Zhong Tianyang, Shu Peng, Li Yiwei, Jiang Hanqi, Zhou Yifan, Chen Junhao, Ruan Wei, Pan Yi, Zhao Huaqin, Ma Chong, Yang Zhenyuan, Xu Shaochen, Zhang Ruidong, Dai Haixing, Zhao Lin, and 12 others. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
David Pape, Sina Mavali, Thorsten Eisenhofer, and Lea Schönherr. 2024 · 2024
Cited alongside, same era.
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2024 · 2024
Cited alongside, same era.
Privacylens: Evaluating privacy norm awareness of language models in action
Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2024a · 2024
Cited alongside, same era.
Papillon: Privacy preservation from internet-based and local language model ensembles
Li Siyan, Vethavikashini Chithrra Raghuram, Omar Khattab, Julia Hirschberg, and Zhou Yu. 2024 · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Cited alongside, same era.
How interpretable are reasoning explanations from prompting large language models?
Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024 · 2024
Cited alongside, same era.
Deliberative alignment: Reasoning enables safer language models
Xuechen Zhou, Shibani Santurkar, Deep Ganguli, Amanda Askell, David Krueger, Adam Lerer, Alex Kim, Aditya Malik, Miljan Martic, Cameron McKinnon, and Others. 2024 · 2024
Cited alongside, same era.
Phi-4-reasoning technical report
Marah Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, and 4 others. 2025 · 2025
Cited alongside, same era.
Wenjie Ma, Jingxuan He, Charlie Snell, Tyler Griggs, Sewon Min, and Matei Zaharia. 2025 · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025 · 2025
Closest in time.
Openai o4-mini: A compact reasoning language model
OpenAI. 2025 · 2025
Closest in time.
Scaling up membership inference: When and how attacks succeed on large language models
Haritz Puerto, Martin Gubri, Sangdoo Yun, and Seong Joon Oh. 2025 · 2025
Closest in time.
Position: Contextual integrity washing for language models
Yan Shvartzshnaider and Vasisht Duddu. 2025 · 2025
Closest in time.
Stop overthinking: A survey on efficient reasoning for large language models
Yang Sui, Yu-Neng Chuang, Guanchu Wang, Jiamu Zhang, Tianyi Zhang, Jiayi Yuan, Hongyi Liu, Andrew Wen, Shaochen Zhong, Hanjie Chen, and Xia Hu. 2025 · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning
Qwen Team. 2025 · 2025
Closest in time.
Safety in large reasoning models: A survey
Cheng Wang, Yue Liu, Baolong Li, Duzhen Zhang, Zhongzhi Li, and Junfeng Fang. 2025 · 2025
Closest in time.
On protecting the data privacy of large language models (llms) and llm agents: A literature review
Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzhen Cheng. 2025 · 2025
Closest in time.
Agentdam: Privacy leakage evaluation for autonomous web agents
Arman Zharmagambetov, Chuan Guo, Ivan Evtimov, Maya Pavlova, Ruslan Salakhutdinov, and Kamalika Chaudhuri. 2025 · 2025
Closest in time.