Fetching the paper…
Reading the bibliography…
The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity.
J. Forster, C. Schmidt, O. Koller, M. Bellgardt, and H. Ney, “Extensions of the sign language recognition and translation corpus RWTH-PHOENIX-Weather,” in International Conference on Computational Linguistics, Language Resources and Evaluation , 2014, pp. 1911–1916
1916
Earlier work this paper cites.
L. Bahl, P. Brown, P. De Souza, and R. Mercer, “Maximum mutual information estimation of hidden markov model parameters for speech recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 11. IEEE, 1986, pp. 49–52
1986
Earlier work this paper cites.
S. Satoh and T. Kanade, “Name-it: Association of face and name in video,” in IEEE Computer Society Conference on Computer Vision and Pattern Recognition . IEEE, 1997, pp. 368–373
1997
Earlier work this paper cites.
J. J. Lien, T. Kanade, J. F. Cohn, and C.-C. Li, “Automated facial expression recognition based on facs action units,” in IEEE International Conference on Automatic Face and Gesture Recognition . IEEE, 1998, pp. 390–395
1998
Earlier work this paper cites.
C. S. A. LaRocca, J. J. Morgan, and S. M. Bellinger, “On the path to 2x learning: Exploring the possibilities of advanced speech recognition,” Computer Assisted Language Instruction Consortium Journal , pp. 295–310, 1999
1999
Earlier work this paper cites.
J. Carletta, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, W. Kraaij, M. Kronenthal et al. , “The AMI meeting corpus: A pre-announcement,” in International Workshop on Machine Learning for Multimodal Interaction . Springer, 2005, pp. 28–39
2005
Earlier work this paper cites.
A. Vinciarelli, M. Pantic, H. Bourlard, and A. Pentland, “Social signal processing: state-of-the-art and future perspectives of an emerging domain,” in The 16th ACM International Conference on Multimedia , 2008, pp. 1061–1070
2008
Earlier work this paper cites.
J. Ortega-Garcia, J. Fierrez, F. Alonso-Fernandez, and J. e. a. Galbally, “The multiscenario multienvironment biosecure multimodal database,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 32, no. 6, pp. 1097–1111, 2009
2009
Earlier work this paper cites.
G. Tur, A. Stolcke, L. Voss, S. Peters, D. Hakkani-Tur, J. Dowding, B. Favre, R. Fernández, M. Frampton, M. Frandsen et al. , “The calo meeting assistant system,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 6, pp. 1601–1611, 2010
2010
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in The 28th International Conference on Machine Learning , 2011, pp. 689–696
2011
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “A better way to pretrain deep boltzmann machines,” Advances in Neural Information Processing Systems , vol. 25, 2012
2012
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in 1st International Conference on Learning Representations , 2013
2013
Earlier work this paper cites.
M. Turk, “Multimodal interaction: A review,” Pattern Recognition Letters , vol. 36, pp. 189–195, 2014
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in The 13th European Conference on Computer Vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo, “Image captioning with semantic attention,” in The IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 4651–4659
2016
Earlier work this paper cites.
J. Xu, T. Mei, T. Yao, and Y. Rui, “MSR-VTT: A large video description dataset for bridging video and language,” in The IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 5288–5296
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Information Processing Systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Aljanaki, Y.-H. Yang, and M. Soleymani, “Developing a benchmark for emotional analysis of music,” The Public Library of Science , vol. 12, no. 3, p. e0173392, 2017
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma et al. , “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” International Journal of Computer Vision , vol. 123, pp. 32–73, 2017
2017
Earlier work this paper cites.
S. Katsigiannis and N. Ramzan, “DREAMER: A database for emotion recognition through eeg and ecg signals from wireless low-cost off-the-shelf devices,” IEEE Journal of Biomedical and Health Informatics , vol. 22, no. 1, pp. 98–107, 2017
2017
Earlier work this paper cites.
F. Zenke, B. Poole, and S. Ganguli, “Continual learning through synaptic intelligence,” in International Conference on Machine Learning . PMLR, 2017, pp. 3987–3995
2017
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in The Conference of the North American Chapter of the Association for Computational Linguistics . ACL, 2018, pp. 4171–4186
2018
Earlier work this paper cites.
L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” in The AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
Z. Chen and B. Liu, Lifelong machine learning . Springer, 2018, vol. 1
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Bostrom and G. Durrett, “Byte pair encoding is suboptimal for language model pretraining,” in Findings of the Association for Computational Linguistics , 2020, pp. 4617–4624
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” The Journal of Machine Learning Research , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Cited alongside, same era.
D. Gurari, Y. Zhao, M. Zhang, and N. Bhattacharya, “Captioning images taken by people who are blind,” in The 16th European Conference on Computer Vision . Springer, 2020, pp. 417–434
2020
Cited alongside, same era.
S. Albanie, G. Varol, L. Momeni, T. Afouras, J. S. Chung, N. Fox, and A. Zisserman, “BSL-1K: Scaling up co-articulated sign language recognition using mouthing cues,” in The 16th European Conference on Computer Vision . Springer, 2020, pp. 35–53
2020
Cited alongside, same era.
A. M. H. Tiong, J. Li, B. Li, S. Savarese, and S. C. Hoi, “Plug-and-play VQA: Zero-shot vqa by conjoining large pretrained models with zero training,” in Findings of the Association for Computational Linguistics , 2022, pp. 951–967
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in Neural Information Processing Systems , vol. 35, pp. 36 479–36 494, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Prajwal, R. Mukhopadhyay, V. P. Namboodiri, and C. Jawahar, “A lip sync expert is all you need for speech to lip generation in the wild,” in The 28th ACM International Conference on Multimedia , 2020, pp. 484–492
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in The 12th Language Resources and Evaluation Conference , 2020, pp. 4218–4222
2020
Cited alongside, same era.
R. Dale, “GPT-3: What’s it good for?” Natural Language Engineering , vol. 27, no. 1, pp. 113–118, 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in The 38th International Conference on Machine Learning . PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
J. Li, R. Selvaraju, A. Gotmare, S. Joty, C. Xiong, and S. C. H. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” Advances in Neural Information Processing Systems , vol. 34, pp. 9694–9705, 2021
2021
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al. , “An image is worth 16x16 words: Transformers for image recognition at scale,” in 9th International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill, “Multimodal few-shot learning with frozen language models,” Advances in Neural Information Processing Systems , vol. 34, pp. 200–212, 2021
2021
Cited alongside, same era.
S. Zhao, G. Jia, J. Yang, G. Ding, and K. Keutzer, “Emotion recognition from multiple modalities: Fundamentals and methodologies,” IEEE Signal Processing Magazine , vol. 38, no. 6, pp. 59–73, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
2022
Later among the works it cites.
W. Gan, Z. Qi, J. Wu, and J. C. W. Lin, “Large language models in education: Vision and opportunities,” in IEEE International Conference on Big Data . IEEE, 2023, pp. 1–10
2023
Closest in time.
W. Gan, S. Wan, and P. S. Yu, “Model-as-a-service (MaaS): A survey,” in IEEE International Conference on Big Data . IEEE, 2023, pp. 1–10
2023
Closest in time.
K. Sanderson, “GPT-4 is here: What scientists think,” Nature , vol. 615, no. 7954, p. 773, 2023
2023
Closest in time.
X. Wang, G. Chen, G. Qian, P. Gao, X.-Y. Wei, Y. Wang, Y. Tian, and W. Gao, “Large-scale multi-modal pre-trained models: A comprehensive survey,” Machine Intelligence Research , pp. 1–36, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Rao, Z. Shan, L. Liu, Y. Zhou, and Y. Yang, “Retrieval-based knowledge augmented vision language pre-training,” in the 31st ACM International Conference on Multimedia , 2023, pp. 5399–5409
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International Conference on Machine Learning , 2023, pp. 19 730–19 742
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
H. Xu, Q. Ye, M. Yan, Y. Shi, J. Ye, Y. Xu, C. Li, B. Bi, Q. Qian, W. Wang et al. , “mPLUG-2: A modularized multi-modal foundation model across text, image and video,” in International Conference on Machine Learning , 2023, pp. 38 728–38 748
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
F. Zeng, W. Gan, Y. Wang, and P. S. Yu, “Distributed training of large language models,” in The 29th IEEE International Conference on Parallel and Distributed Systems . IEEE, 2023, pp. 1–8
2023
Closest in time.