Fetching the paper…
Reading the bibliography…
Multimodal Large Language Models (MLLMs) are distinguished by their multimodal comprehensive ability and widely used in many real-world applications including GPT-4o, autonomous driving and robotics.
Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q. V.; and Salakhutdinov, R. 2019 · 1901
Earlier work this paper cites.
GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Hudson, D. A.; and Manning, C. D. 2019 · 1902
Earlier work this paper cites.
Adaptive attention span in transformers
Sukhbaatar, S.; Grave, E.; Bojanowski, P.; and Joulin, A. 2019 · 1905
Earlier work this paper cites.
Compressive Transformers for Long-Range Sequence Modelling
Rae, J. W.; Potapenko, A.; Jayakumar, S. M.; and Lillicrap, T. P. 2019 · 1911
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J.; Mohiuddin, A.; Kaiser, L.; et al. 2020 · 2009
Earlier work this paper cites.
TGIF-QA: Toward Spatio-Temporal Reasoning in Visual Question Answering
Jang, Y.; Song, Y.; Yu, Y.; Kim, Y.; and Kim, G. 2017 · 2017
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Earlier work this paper cites.
Cswin transformer: A general vision transformer backbone with cross-shaped windows
Dong, X.; Bao, J.; Chen, D.; Zhang, W.; Yu, N.; Yuan, L.; Chen, D.; and Guo, B. 2022 · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Earlier work this paper cites.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z.; Hu, H.; Lin, Y.; Yao, Z.; Xie, Z.; Wei, Y.; Ning, J.; Cao, Y.; Zhang, Z.; Dong, L.; et al. 2022 · 2022
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O.; Smith, N. A.; and Lewis, M. 2022 · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q. G.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; et al. 2023 · 2023
Cited alongside, same era.
Extending context window of large language models via positional interpolation
Chen, S.; Wong, S.; Chen, L.; and Tian, Y. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Cited alongside, same era.
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
Gao, P.; Zhang, R.; Liu, C.; Qiu, L.; Huang, S.; Lin, W.; Zhao, S.; Geng, S.; Lin, Z.; Jin, P.; et al. 2024 · 2024
Closest in time.
Chat-univi: Unified visual representation empowers large language models with image and video understanding
Jin, P.; Takanobu, R.; Zhang, C.; Cao, X.; and Yuan, L. 2024 · 2024
Closest in time.
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Kang, H.; Zhang, Q.; Kundu, S.; Jeong, G.; Liu, Z.; Krishna, T.; and Zhao, T. 2024 · 2024
Closest in time.
How Long Can Open-Source LLMs Truly Promise on Context Length?
Li*, D.; Shao*, R.; Xie, A.; Sheng, Y.; Zheng, L.; Gonzalez, J. E.; Stoica, I.; Ma, X.; and Zhang, H. 2023 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
LLaMA-VID: An image is worth 2 tokens in large language models
Li, Y.; Wang, C.; and Jia, J. 2023 · 2023
Cited alongside, same era.
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2023 · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Pope, R.; Douglas, S.; Chowdhery, A.; Devlin, J.; Bradbury, J.; Heek, J.; Xiao, K.; Agrawal, S.; and Dean, J. 2023 · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Team, G.; Anil, R.; Borgeaud, S.; Wu, Y.; Alayrac, J.-B.; Yu, J.; Soricut, R.; Schalkwyk, J.; Dai, A. M.; Hauth, A.; et al. 2023 · 2023
Cited alongside, same era.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Adnan, M.; Arunkumar, A.; Jain, G.; Nair, P.; Soloveychik, I.; and Kamath, P. 2024 · 2024
Cited alongside, same era.
SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Li, B.; Wang, R.; Wang, G.; Ge, Y.; Ge, Y.; and Shan, Y. 2023a
Cited in the paper.
Hello GPT-4o
OpenAI. 2024 · 2024
Closest in time.
Yarn: Efficient context window extension of large language models
Peng, B.; Quesnelle, J.; Fan, H.; and Shippole, E. 2024 · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J.; Ahmed, M.; Lu, Y.; Pan, S.; Bo, W.; and Liu, Y. 2024 · 2024
Closest in time.
Preparing for the era of 32K context: Early learnings and explorations
Together. 2023 · 2024
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G.; Tian, Y.; Chen, B.; Han, S.; and Lewis, M. 2024 · 2024
Closest in time.
Yu, T.; Yao, Y.; Zhang, H.; He, T.; Han, Y.; Cui, G.; Hu, J.; Liu, Z.; Zheng, H.-T.; Sun, M.; and Chua, T.-S. 2024 · 2024
Closest in time.