Fetching the paper…
Reading the bibliography…
State Space Models (SSMs) have emerged as efficient alternatives to Transformers for sequential modeling, but their inability to leverage modality-specific features limits their performance in multi-modal pretraining.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding, 2020
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2006
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I · 2013
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Glu variants improve transformer
Shazeer, N · 2020
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Ré, C · 2021
Earlier work this paper cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Earlier work this paper cites.
Theoretically principled deep rl acceleration via nearest neighbor function approximation
Shen, J. and Yang, L. F · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
Fedus, W., Zoph, B., and Shazeer, N · 2022
Earlier work this paper cites.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O., Polyak, A., Ashual, O., Sheynin, S., Parikh, D., and Taigman, Y · 2022
Earlier work this paper cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J. Y · 2022
Earlier work this paper cites.
Efficient architecture search for diverse tasks
Shen, J., Khodak, M., and Talwalkar, A · 2022
Cited alongside, same era.
NAS-bench-360: Benchmarking neural architecture search on diverse tasks
Tu, R., Roberts, N., Khodak, M., Shen, J., Sala, F., and Talwalkar, A · 2022
Cited alongside, same era.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks, 2022
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., and Wei, F · 2022
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Cited alongside, same era.
Multiway-adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
Long, Z., Killick, G., McCreadie, R., and Camarasa, G. A · 2023
Moma: Efficient early-fusion pre-training with mixture of modality-aware experts
Lin, X. V., Shrivastava, A., Luo, L., Iyer, S., Lewis, M., Gosh, G., Zettlemoyer, L., and Aghajanyan, A · 2024
Later among the works it cites.
Scaling diffusion mamba with bidirectional ssms for efficient image and video generation
Mo, S. and Tian, Y · 2024
Later among the works it cites.
Moe-mamba: Efficient selective state space models with mixture of experts, 2024
Pióro, M., Ciebiera, K., Król, K., Ludziejewski, J., and Jaszczur, S · 2024
Later among the works it cites.
Vl-mamba: Exploring state space models for multimodal learning
Qiao, Y., Yu, Z., Guo, L., Chen, S., Zhao, Z., Sun, M., Wu, Q., and Liu, J · 2024
Later among the works it cites.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cogvlm: Visual expert for pretrained language models
Wang, W., Lv, Q., Yu, W., Hong, W., Qi, J., Wang, Y., Ji, J., Yang, Z., Zhao, L., Song, X., et al · 2023
Cited alongside, same era.
Blackmamba: Mixture of experts for state-space models
Anthony, Q., Tokpanov, Y., Glorioso, P., and Millidge, B · 2024
Cited alongside, same era.
Chameleon: Mixed-modal early-fusion foundation models, 2024
Chameleon Team · 2024
Cited alongside, same era.
Scalable diffusion models with state space backbone
Fei, Z., Fan, M., Yu, C., and Huang, J · 2024
Cited alongside, same era.
Mars: Mixture of auto-regressive models for fine-grained text-to-image synthesis
He, W., Fu, S., Liu, M., Wang, X., Xiao, W., Shu, F., Wang, Y., Zhang, L., Yu, Z., Li, H., et al · 2024
Cited alongside, same era.
Zigma: A dit-style zigzag mamba diffusion model
Hu, V. T., Baumann, S. A., Gui, M., Grebenkova, O., Ma, P., Schusterbauer, J., and Ommer, B · 2024
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Cited alongside, same era.
Sukhbaatar, S., Golovneva, O., Sharma, V., Xu, H., Lin, X. V., Rozière, B., Kahn, J., Li, D., tau Yih, W., Weston, J., and Li, X · 2024
Later among the works it cites.
Learning to (learn at test time): Rnns with expressive hidden states
Sun, Y., Li, X., Dalal, K., Xu, J., Vikram, A., Zhang, G., Dubois, Y., Chen, X., Wang, X., Koyejo, S., et al · 2024
Later among the works it cites.
Specialized foundation models struggle to beat supervised baselines, 2024
Xu, Z., Gupta, R., Cheng, W., Shen, A., Shen, J., Talwalkar, A., and Khodak, M · 2024
Later among the works it cites.
Diffusion models without attention
Yan, J. N., Gu, J., and Rush, A. M · 2024
Later among the works it cites.
Cobra: Extending mamba to multi-modal large language model for efficient inference
Zhao, H., Zhang, M., Zhao, W., Ding, P., Huang, S., and Wang, D · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O · 2024
Later among the works it cites.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Zhu, L., Liao, B., Zhang, Q., Wang, X., Liu, W., and Wang, X · 2024
Later among the works it cites.
Cat: Content-adaptive image tokenization, 2025
Shen, J., Tirumala, K., Yasunaga, M., Misra, I., Zettlemoyer, L., Yu, L., and Zhou, C · 2025
Closest in time.