Fetching the paper…
Reading the bibliography…
Techniques that enhance inference through increased computation at test-time have recently gained attention.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Minimum Bayes-risk decoding for statistical machine translation
Shankar Kumar and William Byrne. 2004 · 2004
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen. 2017 · 2017
Earlier work this paper cites.
Comparison of diverse decoding methods from conditional language models
Daphne Ippolito, Reno Kriz, João Sedoc, Maria Kustikova, and Chris Callison-Burch. 2019 · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Mitigating gender bias in natural language processing: Literature review
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019 · 2019
Earlier work this paper cites.
Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology
Ran Zmigrod, Sabrina J. Mielke, Hanna Wallach, and Ryan Cotterell. 2019 · 2019
Earlier work this paper cites.
Language (technology) is power: A critical survey of “bias” in NLP
Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020 · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
DExperts: Decoding-time controlled text generation with experts and anti-experts
Alisa Liu, Maarten Sap, Ximing Lu, Swabha Swayamdipta, Chandra Bhagavatula, Noah A. Smith, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
NeuroLogic decoding: (un)supervised neural text generation with predicate logic constraints
Ximing Lu, Peter West, Rowan Zellers, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
NeuroLogic a*esque decoding: Constrained text generation with lookahead heuristics
Ximing Lu, Sean Welleck, Peter West, Liwei Jiang, Jungo Kasai, Daniel Khashabi, Ronan Le Bras, Lianhui Qin, Youngjae Yu, Rowan Zellers, Noah A. Smith, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
An empirical survey of the effectiveness of debiasing techniques for pre-trained language models
Nicholas Meade, Elinor Poole-Dayan, and Siva Reddy. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Monitor-guided decoding of code LMs with static analysis of repository context
Lakshya Agrawal, Aditya Kanade, Navin Goyal, Shuvendu K Lahiri, and Sriram Rajamani. 2023 · 2023
Earlier work this paper cites.
Kushal Arora, Timothy J O’Donnell, Doina Precup, Jason Weston, and Jackie CK Cheung. 2023 · 2023
Earlier work this paper cites.
NEUROSTRUCTURAL DECODING: Neural text generation with structural constraints
Mohaddeseh Bastan, Mihai Surdeanu, and Niranjan Balasubramanian. 2023 · 2023
Earlier work this paper cites.
Faster minimum Bayes risk decoding with confidence-based pruning
Julius Cheng and Andreas Vlachos. 2023 · 2023
Earlier work this paper cites.
KCTS: Knowledge-constrained tree search decoding with token-level hallucination detection
Sehyun Choi, Tianqing Fang, Zhaowei Wang, and Yangqiu Song. 2023 · 2023
Earlier work this paper cites.
Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model
Haikang Deng and Colin Raffel. 2023 · 2023
Earlier work this paper cites.
Co 2 PT: Mitigating bias in pre-trained language models through counterfactual contrastive prompt tuning
Xiangjue Dong, Ziwei Zhu, Zhuoer Wang, Maria Teleki, and James Caverlee. 2023 · 2023
Earlier work this paper cites.
Grammar-constrained decoding for structured NLP tasks without finetuning
Saibo Geng, Martin Josifoski, Maxime Peyrard, and Robert West. 2023 · 2023
Earlier work this paper cites.
The benefits of bad advice: Autocontrastive decoding across model layers
Ariel Gera, Roni Friedman, Ofir Arviv, Chulaka Gunasekara, Benjamin Sznajder, Noam Slonim, and Eyal Shnarch. 2023 · 2023
Earlier work this paper cites.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023 · 2023
Earlier work this paper cites.
Grounded decoding: Guiding text generation with grounded models for embodied agents
Wenlong Huang, Fei Xia, Dhruv Shah, Danny Driess, Andy Zeng, Yao Lu, Pete Florence, Igor Mordatch, Sergey Levine, Karol Hausman, and brian ichter. 2023 · 2023
Earlier work this paper cites.
Speculative decoding with big little decoder
Sehoon Kim, Karttikeya Mangalam, Suhong Moon, Jitendra Malik, Michael W. Mahoney, Amir Gholami, and Kurt Keutzer. 2023 · 2023
Earlier work this paper cites.
Critic-driven decoding for mitigating hallucinations in data-to-text generation
Mateusz Lango and Ondrej Dusek. 2023 · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Earlier work this paper cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023 · 2023
Earlier work this paper cites.
Self-ensemble of n n -best generation hypotheses by lexically constrained decoding
Ryota Miyano, Tomoyuki Kajiwara, and Yuki Arase. 2023 · 2023
Earlier work this paper cites.
Pointwise mutual information based metric and decoding strategy for faithful generation in document grounded dialogs
Yatin Nandwani, Vineet Kumar, Dinesh Raghu, Sachindra Joshi, and Luis Lastras. 2023 · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Accelerating transformer inference for translation via parallel decoding
Andrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, and Emanuele Rodola. 2023 · 2023
Cited alongside, same era.
On second thought, let’s not think step by step! bias and toxicity in zero-shot reasoning
Omar Shaikh, Hongxin Zhang, William Held, Michael Bernstein, and Diyi Yang. 2023 · 2023
Cited alongside, same era.
Prompting GPT-3 to be reliable
Chenglei Si, Zhe Gan, Zhengyuan Yang, Shuohang Wang, Jianfeng Wang, Jordan Lee Boyd-Graber, and Lijuan Wang. 2023 · 2023
Cited alongside, same era.
Instructive decoding: Instruction-tuned large language models are self-refiner from noisy instructions
Taehyeon Kim, Joonkee Kim, Gihun Lee, and Se-Young Yun. 2024 · 2024
Closest in time.
On the multilingual ability of decoder-based pre-trained language models: Finding and controlling language-specific neurons
Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, and Yutaka Matsuo. 2024 · 2024
Closest in time.
Constrained decoding for cross-lingual label projection
Duong Minh Le, Yang Chen, Alan Ritter, and Wei Xu. 2024 · 2024
Closest in time.
The unlocking spell on base LLMs: Rethinking alignment via in-context learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2024 · 2024
Closest in time.
Reference trustable decoding: A training-free augmentation paradigm for large language models
Shi Luohe, Yao Yao, Zuchao Li, Lefei Zhang, and hai zhao. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spectr: Fast speculative decoding via optimal transport
Ziteng Sun, Ananda Theertha Suresh, Jae Hun Ro, Ahmad Beirami, Himanshu Jain, and Felix Yu. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023 · 2023
Cited alongside, same era.
Arithmetic sampling: Parallel diverse decoding for large language models
Luke Vilnis, Yury Zemlyanskiy, Patrick Murray, Alexandre Tachard Passos, and Sumit Sanghai. 2023 · 2023
Cited alongside, same era.
An empirical analysis of parameter-efficient methods for debiasing pre-trained language models
Zhongbin Xie and Thomas Lukasiewicz. 2023 · 2023
Cited alongside, same era.
Look-back decoding for open-ended text generation
Nan Xu, Chunting Zhou, Asli Celikyilmaz, and Xuezhe Ma. 2023 · 2023
Cited alongside, same era.
Fine-grained conversational decoding via isotropic and proximal search
Yuxuan Yao, Han Wu, Qiling Xu, and Linqi Song. 2023 · 2023
Cited alongside, same era.
Prompt-based Monte-Carlo tree search for goal-oriented dialogue policy planning
Xiao Yu, Maximillian Chen, and Zhou Yu. 2023 · 2023
Cited alongside, same era.
Mitigating hallucinations in large vision-language models (LVLMs) via language-contrastive decoding (LCD)
Avshalom Manevich and Reut Tsarfaty. 2024 · 2024
Closest in time.
EchoPrompt: Instructing the model to rephrase queries for improved in-context learning
Raja Sekhar Reddy Mekala, Yasaman Razeghi, and Sameer Singh. 2024 · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024 · 2024
Closest in time.
Controlled decoding from language models
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanping Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, Jilin Chen, Alex Beutel, and Ahmad Beirami. 2024 · 2024
Closest in time.
Skeleton-of-thought: Prompting LLMs for efficient parallel generation
Xuefei Ning, Zinan Lin, Zixuan Zhou, Zifu Wang, Huazhong Yang, and Yu Wang. 2024 · 2024
Closest in time.
Learning to reason with llms
OpenAI. 2024 · 2024
Closest in time.
Grammar-aligned decoding
Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova, and Loris D’Antoni. 2024 · 2024
Closest in time.
FLAP: Flow-adhering planning with constrained decoding in LLMs
Shamik Roy, Sailik Sengupta, Daniele Bonadiman, Saab Mansour, and Arshit Gupta. 2024 · 2024
Closest in time.
Superposed decoding: Multiple generations from a single autoregressive inference pass
Ethan Shen, Alan Fan, Sarah M Pratt, Jae Sung Park, Matthew Wallingford, Sham M. Kakade, Ari Holtzman, Ranjay Krishna, Ali Farhadi, and Aditya Kusupati. 2024 · 2024
Closest in time.
Trusting your evidence: Hallucinate less with context-aware decoding
Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024b · 2024
Closest in time.
Anti-LM decoding for zero-shot in-context machine translation
Suzanna Sia, Alexandra DeLucia, and Kevin Duh. 2024 · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao. 2024 · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2024 · 2024
Closest in time.
Mind the gap: Examining the self-improvement capabilities of large language models
Yuda Song, Hanlin Zhang, Carson Eisenach, Sham Kakade, Dean Foster, and Udaya Ghai. 2024 · 2024
Closest in time.
Specexec: Massively parallel speculative decoding for interactive LLM inference on consumer devices
Ruslan Svirschevski, Avner May, Zhuoming Chen, Beidi Chen, Zhihao Jia, and Max Ryabinin. 2024 · 2024
Closest in time.
Efficient minimum bayes risk decoding using low-rank matrix completion algorithms
Firas Trabelsi, David Vilar, Mara Finkelstein, and Markus Freitag. 2024 · 2024
Closest in time.
Alphazero-like tree-search can guide large language model decoding and training
Ziyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang. 2024 · 2024
Closest in time.
Mitigating hallucinations in large vision-language models with instruction contrastive decoding
Xintong Wang, Jingheng Pan, Liang Ding, and Chris Biemann. 2024b · 2024
Closest in time.
How interpretable are reasoning explanations from prompting large language models?
Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024 · 2024
Closest in time.
From decoding to meta-generation: Inference-time algorithms for large language models
Sean Welleck, Amanda Bertsch, Matthew Finlayson, Hailey Schoelkopf, Alex Xie, Graham Neubig, Ilia Kulikov, and Zaid Harchaoui. 2024 · 2024
Closest in time.
Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding
Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, and Zhifang Sui. 2024 · 2024
Closest in time.
SafeDecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. 2024 · 2024
Closest in time.
Language-informed beam search decoding for multilingual machine translation
Yilin Yang, Stefan Lee, and Prasad Tadepalli. 2024 · 2024
Closest in time.
A theoretical perspective for speculative decoding algorithm
Ming Yin, Minshuo Chen, Kaixuan Huang, and Mengdi Wang. 2024 · 2024
Closest in time.
Discerning and resolving knowledge conflicts through adaptive decoding with contextual information-entropy constraint
Xiaowei Yuan, Zhao Yang, Yequan Wang, Shengping Liu, Jun Zhao, and Kang Liu. 2024b · 2024
Closest in time.
Improving machine translation with large language models: A preliminary study with cooperative decoding
Jiali Zeng, Fandong Meng, Yongjing Yin, and Jie Zhou. 2024 · 2024
Closest in time.
Enhancing contextual understanding in large language models through contrastive decoding
Zheng Zhao, Emilio Monti, Jens Lehmann, and Haytham Assem. 2024 · 2024
Closest in time.
ROSE doesn’t do that: Boosting the safety of instruction-tuned large language models with reverse prompt contrastive decoding
Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2024 · 2024
Closest in time.
Improving open-ended text generation via adaptive decoding
Wenhong Zhu, Hongkun Hao, Zhiwei He, Yiming Ai, and Rui Wang. 2024 · 2024
Closest in time.
Masculine defaults via gendered discourse in podcasts and large language models
Maria Teleki, Xiangjue Dong, Haoran Liu, and James Caverlee. 2025 · 2025
Closest in time.