Fetching the paper…
Reading the bibliography…
Although Large Language Models (LLMs) have shown promise for human-like conversations, they are primarily pre-trained on text data.
T. Brown, B. Mann, N. Ryder et al. , “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
J. Cho, J. Lei, H. Tan, and M. Bansal, “Unifying vision-and-language tasks via text generation,” in International Conference on Machine Learning . PMLR, 2021, pp. 1931–1942
1942
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL:HLT, Volume 1 , Jun. 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “LXMERT: Learning cross-modality encoder representations from transformers,” in Proceedings of EMNLP-IJCNLP , Hong Kong, China, Nov. 2019, pp. 5100–5111
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” NeurIPS , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in ICCV , 2019, pp. 7464–7473
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in ICML , 2019, pp. 2790–2799
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts et al. , “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal et al. , “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL , 2020, pp. 7871–7880
2020
Earlier work this paper cites.
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao, “Unified vision-language pre-training for image captioning and vqa,” in AAAI , vol. 34, no. 07, 2020, pp. 13 041–13 049
2020
Cited alongside, same era.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, “Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,” in ACM SIGKDD , 2020, pp. 3505–3506
2020
Cited alongside, same era.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in ICML , vol. 2, no. 3, 2021, p. 4
2021
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in EMNLP , Nov. 2021, pp. 3045–3059
2021
Cited alongside, same era.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in ACL-IJCNLP , Aug. 2021, pp. 4582–4597
Y.-L. Sung, J. Cho, and M. Bansal, “Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks,” in CVPR , 2022, pp. 5227–5237
2022
Later among the works it cites.
O. Rudovic, A. Bindal, V. Garg, P. Simha, P. Dighe, and S. Kajarekar, “Streaming on-device detection of device directed speech from voice and touch-based invocation,” in ICASSP , 2022, pp. 491–495
2022
Later among the works it cites.
S. Wang, H. Scells, B. Koopman, and G. Zuccon, “Can chatgpt write a good boolean query for systematic review literature search?” ACM SIGIR , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
M. Tsimpoukelli, J. L. Menick, S. Cabi, S. Eslami, O. Vinyals, and F. Hill, “Multimodal few-shot learning with frozen language models,” NeurIPS , vol. 34, pp. 200–212, 2021
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
M. Agrawal, S. Hegselmann, H. Lang, Y. Kim, and D. Sontag, “Large language models are few-shot clinical information extractors,” in EMNLP , Dec. 2022, pp. 1998–2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J.-B. Alayrac, J. Donahue, P. Luc et al. , “Flamingo: a visual language model for few-shot learning,” NeurIPS , vol. 35, pp. 23 716–23 736, 2022
2022
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in ICLR , 2022
2022
Cited alongside, same era.
2023
Later among the works it cites.
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang, “Clap learning audio concepts from natural language supervision,” in ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Wagner, A. Churchill, S. Sigtia, P. Georgiou, M. Mirsamadi, A. Mishra, and E. Marchi, “Multimodal data and resource efficient device-directed speech detection with large foundation models,” ICASSP , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Deshmukh, B. Elizalde, R. Singh, and H. Wang, “Pengi: An audio language model for audio tasks,” NeurIPS , vol. 36, 2024
2024
Closest in time.