Fetching the paper…
Reading the bibliography…
Multimodal Large Models (MLMs) are becoming a significant research focus, combining powerful large language models with multimodal learning to perform complex tasks across different data modalities.
Y. LeCun, J. Denker, and S. Solla, “Optimal brain damage,” Advances in neural information processing systems , vol. 2, 1989
1989
Earlier work this paper cites.
B. Hassibi and D. Stork, “Second order derivatives for network pruning: Optimal brain surgeon,” Advances in neural information processing systems , vol. 5, 1992
1992
Earlier work this paper cites.
G. W. Stewart, “On the early history of the singular value decomposition,” SIAM review , vol. 35, no. 4, pp. 551–566, 1993
1993
Earlier work this paper cites.
B. Hassibi, D. G. Stork, and G. J. Wolff, “Optimal brain surgeon and general network pruning,” in IEEE international conference on neural networks . IEEE, 1993, pp. 293–299
1993
Earlier work this paper cites.
R. M. Gray and D. L. Neuhoff, “Quantization,” IEEE transactions on information theory , vol. 44, no. 6, pp. 2325–2383, 1998
1998
Earlier work this paper cites.
D. Floreano, F. Mondada, A. Perez-Uribe, and D. Roggen, “Evolution of embodied intelligence,” in Embodied Artificial Intelligence: International Seminar, Dagstuhl Castle, Germany, July 7-11, 2003. Revised Papers . Springer, 2004, pp. 293–311
2004
Earlier work this paper cites.
N. Kalchbrenner and P. Blunsom, “Recurrent continuous translation models,” in Proceedings of the 2013 conference on empirical methods in natural language processing , 2013, pp. 1700–1709
2013
Earlier work this paper cites.
V. Bolón-Canedo, N. Sánchez-Maroño, and A. Alonso-Betanzos, “A review of feature selection methods on synthetic data,” Knowledge and information systems , vol. 34, pp. 483–519, 2013
2013
Earlier work this paper cites.
V. Mnih, N. Heess, A. Graves et al. , “Recurrent models of visual attention,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
C. Doersch, “Tutorial on variational autoencoders,” arXiv preprint arXiv:1606.05908 , 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in International conference on machine learning . PMLR, 2017, pp. 933–941
2017
Earlier work this paper cites.
R. Dey and F. M. Salem, “Gate-variants of gated recurrent unit (gru) neural networks,” in 2017 IEEE 60th international midwest symposium on circuits and systems (MWSCAS) . IEEE, 2017, pp. 1597–1600
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
V. Rogers, P. Meara, T. Barnett-Legh, C. Curry, and E. Davie, “Examining the llama aptitude tests,” Journal of the European Second Language Association , vol. 1, no. 1, pp. 49–60, 2017
2017
Earlier work this paper cites.
D. Ha and J. Schmidhuber, “Recurrent world models facilitate policy evolution,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Y. Yu, X. Si, C. Hu, and J. Zhang, “A review of recurrent neural networks: Lstm cells and network architectures,” Neural computation , vol. 31, no. 7, pp. 1235–1270, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al. , “Searching for mobilenetv3,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 1314–1324
2019
Earlier work this paper cites.
J. Ho, X. Chen, A. Srinivas, Y. Duan, and P. Abbeel, “Flow++: Improving flow-based generative models with variational dequantization and architecture design,” in International conference on machine learning . PMLR, 2019, pp. 2722–2730
2019
Earlier work this paper cites.
X. Han, X. Hu, W. Huang, and M. R. Scott, “Clothflow: A flow-based model for clothed person generation,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 10 471–10 480
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
D. Howard, A. E. Eiben, D. F. Kennedy, J.-B. Mouret, P. Valencia, and D. Winkler, “Evolving embodied intelligence from materials to machines,” Nature Machine Intelligence , vol. 1, no. 1, pp. 12–19, 2019
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International conference on machine learning . PMLR, 2020, pp. 5156–5165
2020
Earlier work this paper cites.
V. O. Yazici, A. Gonzalez-Garcia, A. Ramisa, B. Twardowski, and J. v. d. Weijer, “Orderless recurrent models for multi-label classification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 13 440–13 449
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Goyal, A. R. Choudhury, S. Raje, V. Chakaravarthy, Y. Sabharwal, and A. Verma, “Power-bert: Accelerating bert inference via progressive word-vector elimination,” in International Conference on Machine Learning . PMLR, 2020, pp. 3690–3699
2020
Earlier work this paper cites.
Z. Dai, G. Lai, Y. Yang, and Q. Le, “Funnel-transformer: Filtering out sequential redundancy for efficient language processing,” Advances in neural information processing systems , vol. 33, pp. 4271–4282, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 9459–9474, 2020
2020
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Earlier work this paper cites.
W. Ping, K. Peng, K. Zhao, and Z. Song, “Waveflow: A compact flow-based model for raw audio,” in International Conference on Machine Learning . PMLR, 2020, pp. 7706–7716
2020
Earlier work this paper cites.
A. Mallya, T.-C. Wang, K. Sapra, and M.-Y. Liu, “World-consistent video-to-video synthesis,” in Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VIII 16 . Springer, 2020, pp. 359–378
2020
Earlier work this paper cites.
J. Gou, B. Yu, S. J. Maybank, and D. Tao, “Knowledge distillation: A survey,” International Journal of Computer Vision , vol. 129, no. 6, pp. 1789–1819, 2021
2021
Earlier work this paper cites.
A. Rajagopal and V. Nirmala, “Convolutional gated mlp: Combining convolutions and gmlp,” in International Conference on Big Data, Machine Learning, and Applications . Springer, 2021, pp. 721–735
2021
Earlier work this paper cites.
Z. Shen, M. Zhang, H. Zhao, S. Yi, and H. Li, “Efficient attention: Attention with linear complexities,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2021, pp. 3531–3539
2021
Earlier work this paper cites.
A. Roy, M. Saffar, A. Vaswani, and D. Grangier, “Efficient content-based sparse attention with routing transformers,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 53–68, 2021
2021
Cited alongside, same era.
X. He, K. Zhao, and X. Chu, “Automl: A survey of the state-of-the-art,” Knowledge-based systems , vol. 212, p. 106622, 2021
2021
Cited alongside, same era.
T. Liang, J. Glossner, L. Wang, S. Shi, and X. Zhang, “Pruning and quantization for deep neural network acceleration: A survey,” Neurocomputing , vol. 461, pp. 370–403, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Later among the works it cites.
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang, “Diffusion models: A comprehensive survey of methods and applications,” ACM Computing Surveys , vol. 56, no. 4, pp. 1–39, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 18 392–18 402
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
D. Kingma, T. Salimans, B. Poole, and J. Ho, “Variational diffusion models,” Advances in neural information processing systems , vol. 34, pp. 21 696–21 707, 2021
2021
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International conference on machine learning . Pmlr, 2021, pp. 8821–8831
2021
Cited alongside, same era.
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99–106, 2021
2021
Cited alongside, same era.
A. Gupta, S. Savarese, S. Ganguli, and L. Fei-Fei, “Embodied intelligence via learning and evolution,” Nature communications , vol. 12, no. 1, p. 5721, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Y. LeCun, “A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27,” Open Review , vol. 62, no. 1, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4195–4205
2023
Later among the works it cites.
J. Li, K. Pan, Z. Ge, M. Gao, W. Ji, W. Zhang, T.-S. Chua, S. Tang, H. Zhang, and Y. Zhuang, “Fine-tuning multimodal llms to follow zero-shot demonstrative instructions,” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
A. Aghajanyan, L. Yu, A. Conneau, W.-N. Hsu, K. Hambardzumyan, S. Zhang, S. Roller, N. Goyal, O. Levy, and L. Zettlemoyer, “Scaling laws for generative mixed-modal language models,” in International Conference on Machine Learning . PMLR, 2023, pp. 265–279
2023
Later among the works it cites.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi et al. , “Audiolm: a language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Wang, X. Du, J. Li, R. A. Yeh, and G. Shakhnarovich, “Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 12 619–12 629
2023
Later among the works it cites.
A. Raj, S. Kaza, B. Poole, M. Niemeyer, N. Ruiz, B. Mildenhall, S. Zada, K. Aberman, M. Rubinstein, J. Barron et al. , “Dreambooth3d: Subject-driven text-to-3d generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2349–2359
2023
Later among the works it cites.
2023
Later among the works it cites.
H. Chen, J. Gu, A. Chen, W. Tian, Z. Tu, L. Liu, and H. Su, “Single-stage diffusion nerf: A unified approach to 3d generation and reconstruction,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 2416–2425
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
A. Meta, “Introducing meta llama 3: The most capable openly available llm to date,” Meta AI , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Ono and A. Morita, “Evaluating large language models: Chatgpt-4, mistral 8x7b, and google gemini benchmarked against mmlu,” Authorea Preprints , 2024
2024
Closest in time.
2024
Closest in time.
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” Advances in neural information processing systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
F. B. Baldassini, M. Shukor, M. Cord, L. Soulier, and B. Piwowarski, “What makes multimodal in-context learning work?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 1539–1550
2024
Closest in time.
Q. Sun, Y. Cui, X. Zhang, F. Zhang, Q. Yu, Y. Wang, Y. Rao, J. Liu, T. Huang, and X. Wang, “Generative multimodal models are in-context learners,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 398–14 409
2024
Closest in time.
H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan, “Videocrafter2: Overcoming data limitations for high-quality video diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 7310–7320
2024
Closest in time.
2024
Closest in time.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. N. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.