Fetching the paper…
Reading the bibliography…
Vision-Language Models (VLMs) enable powerful multimodal reasoning but suffer from slow autoregressive inference, limiting their deployment in real-time applications.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M. M. A., et al · 1909
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y. W · 2006
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Earlier work this paper cites.
Consistent accelerated inference via confident adaptive transformers
Schuster, T., Fisch, A., Jaakkola, T., and Barzilay, R · 2021
Earlier work this paper cites.
GPTQ: Accurate post-training compression for generative pretrained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
Rajbhandari, S. et al · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Human-oriented Representation Learning for Robotic Manipulation
Huo, M., Ding, M., Xu, C., Tian, T., Zhu, X., Mu, Y., Sun, L., Tomizuka, M., and Zhan, W · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Earlier work this paper cites.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., et al · 2023
Cited alongside, same era.
Pass: Parallel speculative sampling
Monea, G., Joulin, A., and Grave, E · 2023
Cited alongside, same era.
Draft & verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S · 2023
Cited alongside, same era.
Distillspec: Improving speculative decoding via knowledge distillation
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A. K., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2023
Cited alongside, same era.
The Integrated Transportation Distance Between Markov Kernels: Theory, Optimization, and Applications in Risk Evaluation and Machine Learning
Lin, Z · 2024
Later among the works it cites.
Opt-tree: Speculative decoding with adaptive draft tree structure
Wang, J., Su, Y., Li, J., Xia, Q., Ye, Z., Duan, X., Wang, Z., and Zhang, M · 2024
Later among the works it cites.
Fast and lossless llm decoding via ctc-based drafting
Wen, Q., Chen, R., Yan, X., Zhao, W. X., and Wen, J.-R · 2024
Later among the works it cites.
On-device language models: A comprehensive review
Xu, J., Zhao, Y., Sun, J., Lin, J., and Han, S · 2024
Later among the works it cites.
Impact of noisy supervision in foundation model learning
Chen, H., Wang, Z., Tao, R., Wei, H., Xie, X., Sugiyama, M., Raj, B., and Wang, J · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot, 2023
Zhu, X., Tian, R., Xu, C., Huo, M., Zhan, W., Tomizuka, M., and Ding, M · 2023
Cited alongside, same era.
Hydra: Sequentially-dependent draft heads for medusa decoding
Ankner, Z., Parthasarathy, R., Nrusimha, A., Rinard, C., Ragan-Kelley, J., and Brandon, W · 2024
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Cited alongside, same era.
Sequoia: Scalable, robust, and hardware-aware speculative decoding
Chen, Z., May, A., Svirschevski, R., Huang, Y., Ryabinin, M., Jia, Z., and Chen, B · 2024
Cited alongside, same era.
AbHE: All Attention-Based Homography Estimation
Huo, M., Zhang, Z., Ren, X., Yang, X., and Ye, C · 2024
Cited alongside, same era.
Nearest neighbor speculative decoding
Li, B., Zhang, B., Xu, Z., Liu, S., Shi, H., and Li, L · 2024
Cited alongside, same era.
Joint Pedestrian Trajectory Prediction through Posterior Sampling
Lin, H., Wang, Y., Huo, M., Peng, C., Liu, Z., and Tomizuka, M · 2024
Cited alongside, same era.
Impact of noisy supervision in foundation model learning, 2025b
Chen, H., Wang, Z., Tao, R., Wei, H., Xie, X., Sugiyama, M., Raj, B., and Wang, J
Cited in the paper.
Don’t lose yourself: Boosting multimodal recommendation via reducing node-neighbor discrepancy in graph convolutional network
Chen, Z., Xu, J., and Hu, H · 2025
Closest in time.
Sca-lstm: A deep learning approach to golf swing analysis and performance enhancement
Feng, C., Bačić, B., and Li, W · 2025
Closest in time.
Leveraging optimal transport for distributed two-sample testing: An integrated transportation distance-based framework
Lin, Z. and Chen, Y · 2025
Closest in time.
Lin, Z. et al · 2025
Closest in time.
Setransformer: A hybrid attention-based architecture for robust human activity recognition
Liu, Y., Qin, X., Gao, Y., Li, X., and Feng, C · 2025
Closest in time.
Monofusion: Sparse-view 4d reconstruction via monocular fusion, 2025
Wang, Z., Tan, J., Khurana, T., Peri, N., and Ramanan, D · 2025
Closest in time.