Fetching the paper…
Reading the bibliography…
To enhance reasoning capabilities, previous works have explored incorporating special-purpose tokens into the training process.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Augmenting self-attention with persistent memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019 · 1907
Earlier work this paper cites.
CTRL - A Conditional Transformer Language Model for Controllable Generation
Nitish Shirish Keskar, Bryan McCann, Lav Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Greedy decoding for statistical machine translation in almost linear time
Ulrich Germann. 2003 · 2003
Earlier work this paper cites.
Tldr: token loss dynamic reweighting for reducing repetitive utterance generation
Shaojie Jiang, Thomas Wolf, Christof Monz, and Maarten de Rijke. 2020 · 2003
Earlier work this paper cites.
Mikhail S Burtsev, Yuri Kuratov, Anton Peganov, and Grigory V Sapunov. 2020 · 2006
Earlier work this paper cites.
Posterior calibration and exploratory analysis for natural language processing models
Khanh Nguyen and Brendan O’Connor. 2015 · 2015
Earlier work this paper cites.
Beam search strategies for neural machine translation
Markus Freitag and Yaser Al-Onaizan. 2017 · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017 · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. 2018 · 2018
Earlier work this paper cites.
Confidence modeling for neural semantic parsing
Li Dong, Chris Quirk, and Mirella Lapata. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
End-to-end bias mitigation by modelling biases in corpora
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020 · 2020
Cited alongside, same era.
Calibrating deep neural networks using focal loss
Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz, Philip Torr, and Puneet Dokania. 2020 · 2020
Cited alongside, same era.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Cited alongside, same era.
Dynamically weighted balanced loss: class imbalanced learning and confidence calibration of deep neural networks
K Ruwani M Fernando and Chris P Tsokos. 2021 · 2021
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Later among the works it cites.
Phi-2: The surprising power of small language models
Microsoft. 2023 · 2023
Later among the works it cites.
From words to watts: Benchmarking the energy costs of large language model inference
Siddharth Samsi, Dan Zhao, Joseph McDonald, Baolin Li, Adam Michaleas, Michael Jones, William Bergeron, Jeremy Kepner, Devesh Tiwari, and Vijay Gadepally. 2023 · 2023
Later among the works it cites.
Dual focal loss for calibration
Linwei Tao, Minjing Dong, and Chang Xu. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dynamicvit: Efficient vision transformers with dynamic token sparsification
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. 2021 · 2021
Cited alongside, same era.
Recurrent memory transformer
Aydar Bulatov, Yury Kuratov, and Mikhail Burtsev. 2022 · 2022
Cited alongside, same era.
Large language models are reasoning teachers
Namgyu Ho, Laura Schmid, and Se-Young Yun. 2022 · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, et al. 2022 · 2022
Cited alongside, same era.
Distilling reasoning capabilities into smaller language models
Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2022 · 2022
Cited alongside, same era.
Can open-domain qa reader utilize external knowledge efficiently like humans?
Neeraj Varshney, Man Luo, and Chitta Baral. 2022 · 2022
Cited alongside, same era.
Language modeling via stochastic processes
Rose E Wang, Esin Durmus, Noah Goodman, and Tatsunori Hashimoto. 2022 · 2022
Cited alongside, same era.
Adaptive computation with elastic input sequence
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You. 2023 · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024 · 2024
Later among the works it cites.
Large language models for mathematical reasoning: Progresses and challenges
Janice Ahn, Rishu Verma, Renze Lou, Di Liu, Rui Zhang, and Wenpeng Yin. 2024 · 2024
Later among the works it cites.
Llama 3 model card
AI@Meta. 2024 · 2024
Later among the works it cites.
The pitfalls of next-token prediction
Gregor Bachmann and Vaishnavh Nagarajan. 2024 · 2024
Later among the works it cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2024 · 2024
Later among the works it cites.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al. 2024 · 2024
Later among the works it cites.
Think before you speak: Training language models with pause tokens
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan. 2024 · 2024
Later among the works it cites.
Thinking tokens for language modeling
David Herel and Tomas Mikolov. 2024 · 2024
Later among the works it cites.
Noiseboost: Alleviating hallucination with noise perturbation for multimodal large language models
Kai Wu, Boyuan Jiang, Zhengkai Jiang, Qingdong He, Donghao Luo, Shengzhi Wang, Qingwen Liu, and Chengjie Wang. 2024 · 2024
Later among the works it cites.