Fetching the paper…
Reading the bibliography…
Large language models can exhibit unexpected behavior in the blink of an eye.
Binomial approximation to the poisson binomial distribution: The krawtchouk expansion
B. Roos · 2001
Earlier work this paper cites.
Thin shell implies spectral gap up to polylog via a stochastic localization scheme
Ronen Eldan · 2013
Earlier work this paper cites.
Generative adversarial networks, 2014
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Limits on support recovery with probabilistic models: An information-theoretic framework, 2016
Jonathan Scarlett and Volkan Cevher · 2016
Earlier work this paper cites.
Probability in high dimension
Ramon van Handel · 2016
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings, 2019
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2019
Earlier work this paper cites.
High-dimensional regression with binary coefficients. estimating squared error and a phase transition, 2019
David Gamarnik and Ilias Zadik · 2019
Earlier work this paper cites.
The all-or-nothing phenomenon in sparse linear regression, 2019
Galen Reeves, Jiaming Xu, and Ilias Zadik · 2019
Earlier work this paper cites.
All-or-nothing statistical and computational phase transitions in sparse spiked matrix estimation, 2020
Jean Barbier, Nicolas Macris, and Cynthia Rush · 2020
Earlier work this paper cites.
Taming correlations through entropy-efficient measure decompositions with applications to mean-field approximation
Ronen Eldan · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang · 2020
Earlier work this paper cites.
The all-or-nothing phenomenon in sparse tensor pca
Jonathan Niles-Weed and Ilias Zadik · 2020
Earlier work this paper cites.
Support recovery in the phase retrieval model: Information-theoretic fundamental limits, 2020
Lan V. Truong and Jonathan Scarlett · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
On the all-or-nothing behavior of bernoulli group testing, 2021
Lan V. Truong, Matthew Aldridge, and Jonathan Scarlett · 2021
Earlier work this paper cites.
Localization schemes: A framework for proving mixing bounds for markov chains
Yuansi Chen and Ronen Eldan · 2022
Earlier work this paper cites.
Perception prioritized training of diffusion models
Jooyoung Choi, Jungbeom Lee, Chaehun Shin, Sungwon Kim, Hyunwoo Kim, and Sungroh Yoon · 2022
Earlier work this paper cites.
Statistical and computational phase transitions in group testing
Amin Coja-Oghlan, Oliver Gebhard, Max Hahn-Klimroth, Alexander S Wein, and Ilias Zadik · 2022
Earlier work this paper cites.
Sampling from the sherrington-kirkpatrick gibbs measure via algorithmic stochastic localization
Ahmed El Alaoui, Andrea Montanari, and Mark Sellke · 2022
Earlier work this paper cites.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Earlier work this paper cites.
SDEdit: Guided image synthesis and editing with stochastic differential equations
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon · 2022
Earlier work this paper cites.
An explanation of in-context learning as implicit bayesian inference, 2022
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Earlier work this paper cites.
Detecting language model attacks with perplexity, 2023
Gabriel Alon and Michael Kamfonas · 2023
Cited alongside, same era.
Sampling from mean-field gibbs measures via diffusion processes
Ahmed El Alaoui, Andrea Montanari, and Mark Sellke · 2023
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models, 2023
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou · 2023
Cited alongside, same era.
Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions
Sitan Chen, Sinho Chewi, Jerry Li, Yuanzhi Li, Adil Salim, and Anru Zhang · 2023
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries, 2023
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong · 2023
Cited alongside, same era.
Bayesian scaling laws for in-context learning, 2024
Aryaman Arora, Dan Jurafsky, Christopher Potts, and Noah D. Goodman · 2024
Later among the works it cites.
Developing a computer use model
Anthropic · 2024
Later among the works it cites.
Dynamical regimes of diffusion models, 2024
Giulio Biroli, Tony Bonnaire, Valentin de Bortoli, and Marc Mézard · 2024
Later among the works it cites.
Obfuscated activations bypass llm latent-space defenses, 2024
Luke Bailey, Alex Serrano, Abhay Sheshadri, Mikhail Seleznyov, Jordan Taylor, Erik Jenner, Jacob Hilton, Stephen Casper, Carlos Guestrin, and Scott Emmons · 2024
Later among the works it cites.
A survey on in-context learning, 2024
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, Baobao Chang, Xu Sun, Lei Li, and Zhifang Sui · 2024
Later among the works it cites.
Challenges and solutions for aging adults
Gemini · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Enhancing chat language models by scaling high-quality instructional conversations, 2023
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Zhi Zheng, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou · 2023
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes, 2023
Shivam Garg, Dimitris Tsipras, Percy Liang, and Gregory Valiant · 2023
Cited alongside, same era.
The journey, not the destination: How data guides diffusion models
Kristian Georgiev, Joshua Vendrow, Hadi Salman, Sung Min Park, and Aleksander Madry · 2023
Cited alongside, same era.
Measuring faithfulness in chain-of-thought reasoning, 2023
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, Saurav Kadavath, Shannon Yang, Thomas Henighan, Timothy Maxwell, Timothy Telleen-Lawton, Tristan Hume, Zac Hatfield-Dodds, Jared Kaplan, Jan Brauner, Samuel R. Bowman, and Ethan Perez · 2023
Cited alongside, same era.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
The unlocking spell on base llms: Rethinking alignment via in-context learning, 2023
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi · 2023
Cited alongside, same era.
Sharp thresholds in inference of planted subgraphs
Elchanan Mossel, Jonathan Niles-Weed, Youngtak Sohn, Nike Sun, and Ilias Zadik · 2023
Cited alongside, same era.
Jailbroken llama-3.1-8b-instruct via lora
grimjim · 2024
Later among the works it cites.
Automated multi-turn red-teaming with cascade
Haize Labs · 2024
Later among the works it cites.
A trivial jailbreak against llama 3
Haize Labs · 2024
Later among the works it cites.
Sampling from spherical spin glasses in total variation via algorithmic stochastic localization
Brice Huang, Andrea Montanari, and Huy Tuan Pham · 2024
Later among the works it cites.
What is in your safe data? identifying benign data that breaks safety, 2024
Luxi He, Mengzhou Xia, and Peter Henderson · 2024
Later among the works it cites.
Critical windows: non-asymptotic theory for feature emergence in diffusion models, 2024
Marvin Li and Sitan Chen · 2024
Later among the works it cites.
Llm defenses are not robust to multi-turn human jailbreaks yet, 2024
Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue · 2024
Later among the works it cites.
Critical tokens matter: Token-level contrastive estimation enhances llm’s reasoning capability, 2024
Zicheng Lin, Tian Liang, Jiahao Xu, Xing Wang, Ruilin Luo, Chufan Shi, Siheng Li, Yujiu Yang, and Zhaopeng Tu · 2024
Later among the works it cites.
Discrete diffusion modeling by estimating the ratios of the data distribution, 2024
Aaron Lou, Chenlin Meng, and Stefano Ermon · 2024
Later among the works it cites.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao · 2024
Later among the works it cites.
Safety alignment should be made more than just a few tokens deep, 2024
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma, Subhrajit Roy, Ahmad Beirami, Prateek Mittal, and Peter Henderson · 2024
Later among the works it cites.
Xstest: A test suite for identifying exaggerated safety behaviours in large language models, 2024
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy · 2024
Later among the works it cites.
Probing the latent hierarchical structure of data via diffusion models, 2024
Antonio Sclocchi, Alessandro Favero, Noam Itzhak Levi, and Matthieu Wyart · 2024
Later among the works it cites.
A strongreject for empty jailbreaks, 2024
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers, 2024
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Dissecting learning and forgetting in language model finetuning
Xiao Zhang and Ji Wu · 2024
Later among the works it cites.
Adversarial contrastive decoding: Boosting safety alignment of large language models via opposite prompt optimization, 2024
Zhengyue Zhao, Xiaoyun Zhang, Kaidi Xu, Xing Hu, Rui Zhang, Zidong Du, Qi Guo, and Yunji Chen · 2024
Later among the works it cites.
Detecting misbehavior in frontier reasoning models
OpenAI · 2025
Closest in time.
A phase transition in diffusion models reveals the hierarchical nature of data
Antonio Sclocchi, Alessandro Favero, and Matthieu Wyart · 2025
Closest in time.