Fetching the paper…
Reading the bibliography…
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
The generalised product moment distribution in samples from a normal multivariate population
John Wishart. 1928 · 1928
Earlier work this paper cites.
Building a question answering test collection. In Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval
Ellen M Voorhees and Dawn M Tice. 2000 · 2000
Earlier work this paper cites.
Learning to Guide Random Search
Ozan Sener and Vladlen Koltun. 2020 · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge. In Machine learning challenges workshop
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge. In Proceedings of the second PASCAL challenges workshop on recognising textual entailment
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
A hardware Gaussian noise generator using the Box-Muller method and its error analysis
D-U Lee, John D Villasenor, Wayne Luk, and Philip Heng Wai Leong. 2006 · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge. In Proceedings of the ACL-PASCAL workshop on textual entailment and paraphrasing
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan. 2007 · 2007
Earlier work this paper cites.
The Fifth PASCAL Recognizing Textual Entailment Challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Earlier work this paper cites.
Efficient PGA LFSR implementation whitens pseudorandom numbers. In 2009 international conference on Reconfigurable Computing and FPGAs
Leonard Colavito and Dennis Silage. 2009 · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI spring symposium series
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
The winograd schema challenge. In Thirteenth international conference on the principles of knowledge representation and reasoning
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank. In EMNLP
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
FPGA gaussian random number generators with guaranteed statistical accuracy. In 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines
David B Thomas. 2014 · 2014
Earlier work this paper cites.
The table-Hadamard GRNG: An area-efficient FPGA Gaussian random number generator
David B Thomas. 2015 · 2015
Earlier work this paper cites.
FPGA-based CNN inference accelerator synthesized from multi-threaded C software. In 2017 30th IEEE International System-on-Chip Conference (SOCC)
Jin Hee Kim, Brett Grady, Ruolong Lian, John Brothers, and Jason H Anderson. 2017 · 2017
Earlier work this paper cites.
On-chip training of recurrent neural networks with limited numerical precision. In IJCNN
Taesik Na, Jong Hwan Ko, Jaeha Kung, and Saibal Mukhopadhyay. 2017 · 2017
Cited alongside, same era.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir. 2017 · 2017
Cited alongside, same era.
Zeroth-order stochastic variance reduction for nonconvex optimization
Sijia Liu, Bhavya Kailkhura, Pin-Yu Chen, Paishun Ting, Shiyu Chang, and Lisa Amini. 2018 · 2018
Cited alongside, same era.
WiC: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2018 · 2018
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Zarts: On zero-order optimization for neural architecture search
Xiaoxing Wang, Wenxuan Guo, Jianlin Su, Xiaokang Yang, and Junchi Yan. 2022 · 2022
Later among the works it cites.
You Already Have It: A Generator-Free Low-Precision DNN Training Framework Using Stochastic Rounding. In European Conference on Computer Vision
Geng Yuan, Sung-En Chang, Qing Jin, Alec Lu, Yanyu Li, Yushu Wu, Zhenglun Kong, Yanyue Xie, Peiyan Dong, Minghai Qin, et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
ROLLER: Fast and Efficient Tensor Compilation for Deep Learning. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)
Hongyu Zhu et al · 2022
Later among the works it cites.
Fine-tuning language models with just forward passes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model agnostic contrastive explanations for structured data
Amit Dhurandhar, Tejaswini Pedapati, Avinash Balakrishnan, Pin-Yu Chen, Karthikeyan Shanmugam, and Ruchir Puri. 2019 · 2019
Cited alongside, same era.
Performance modeling for CNN inference accelerators on FPGA
Yufei Ma, Yu Cao, Sarma Vrudhula, and Jae-Sun Seo. 2019 · 2019
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta. 2020 · 2020
Cited alongside, same era.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2020 · 2020
Cited alongside, same era.
AXI HyperConnect: A Predictable, Hypervisor-level Interconnect for Hardware Accelerators in FPGA SoC. In 2020 57th ACM/IEEE Design Automation Conference (DAC)
Francesco Restuccia, Alessandro Biondi, Mauro Marinoni, Giorgiomaria Cicero, and Giorgio Buttazzo. 2020 · 2020
Cited alongside, same era.
Pyhessian: Neural networks through the lens of the hessian. In Big data
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney. 2020 · 2020
Cited alongside, same era.
Optimal gradient checkpoint search for arbitrary computation graphs. In CVPR
Jianwei Feng and Dong Huang. 2021 · 2021
Cited alongside, same era.
Sadhika Malladi, Tianyu Gao, Eshaan Nichani, Alex Damian, Jason D Lee, Danqi Chen, and Sanjeev Arora. 2023 · 2023
Later among the works it cites.
Certified Zeroth-order Black-Box Defense with Robust UNet Denoiser
Astha Verma, Siddhesh Bangar, A Venkata Subramanyam, Naman Lal, Rajiv Ratn Shah, and Shin’ichi Satoh. 2023 · 2023
Later among the works it cites.
Dynamic sparse no training: Training-free fine-tuning for sparse llms
Yuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun, Yiwu Yao, Xingjia Han, Jared Tanner, Shiwei Liu, and Rongrong Ji. 2023 · 2023
Later among the works it cites.
Lamda: Large model fine-tuning via spectrally decomposed low-dimensional adaptation
Seyedarmin Azizi, Souvik Kundu, and Massoud Pedram. 2024 · 2024
Later among the works it cites.
TreeGRNG: Binary Tree Gaussian Random Number Generator for Efficient Probabilistic AI Hardware. In DATE
Jonas Crols, Guilherme Paim, Shirui Zhao, and Marian Verhelst. 2024 · 2024
Later among the works it cites.
A Mixed-Precision Transformer Accelerator With Vector Tiling Systolic Array for License Plate Recognition in Unconstrained Scenarios
Jie Li, Dingjiang Yan, Fangzhou He, Zhicheng Dong, and Mingfei Jiang. 2024 · 2024
Later among the works it cites.
Sparse mezo: Less parameters for better performance in zeroth-order llm fine-tuning
Yong Liu, Zirui Zhu, Chaoyu Gong, Minhao Cheng, Cho-Jui Hsieh, and Yang You. 2024 · 2024
Later among the works it cites.
Flightllm: Efficient large language model inference with a complete mapping flow on fpgas. In Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays
Shulin Zeng, Jun Liu, Guohao Dai, Xinhao Yang, Tianyu Fu, Hongyi Wang, Wenheng Ma, Hanbo Sun, Shiyao Li, Zixiao Huang, et al · 2024
Later among the works it cites.
Revisiting zeroth-order optimization for memory-efficient llm fine-tuning: A benchmark
Yihua Zhang et al · 2024
Later among the works it cites.
Second-order fine-tuning without pain for llms: A hessian informed zeroth-order optimizer
Yanjun Zhao, Sizhe Dang, Haishan Ye, Guang Dai, Yi Qian, and Ivor W Tsang. 2024 · 2024
Later among the works it cites.
Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuning
Qitao Tan, Jun Liu, Zheng Zhan, Caiwei Ding, Yanzhi Wang, Jin Lu, and Geng Yuan. 2025 · 2025
Closest in time.