Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) demonstrate remarkable proficiency in comprehending and handling text-based tasks.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Span-based localizing network for natural language video localization
Zhang, H., Sun, A., Jing, W., and Zhou, J. T · 2004
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Principal component analysis
Abdi, H. and Williams, L. J · 2010
Earlier work this paper cites.
Combining embedded accelerometers with computer vision for recognizing food preparation activities
Stein, S. and McKenna, S. J · 2013
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Kuehne, H., Arslan, A., and Serre, T · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Sequence to sequence-video to text
Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., and Saenko, K · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D · 2015
Earlier work this paper cites.
A multi-stream bi-directional recurrent neural network for fine-grained action detection
Singh, B., Marks, T. K., Jones, M., Tuzel, O., and Shao, M · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Tapaswi, M., Zhu, Y., Stiefelhagen, R., Torralba, A., Urtasun, R., and Fidler, S · 2016
Earlier work this paper cites.
Vse++: Improving visual-semantic embeddings with hard negatives
Faghri, F., Fleet, D. J., Kiros, J. R., and Fidler, S · 2017
Earlier work this paper cites.
Tall: Temporal activity localization via language query
Gao, J., Sun, C., Yang, Z., and Nevatia, R · 2017
Earlier work this paper cites.
Mask r-cnn
He, K., Gkioxari, G., Dollár, P., and Girshick, R · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., and Carlos Niebles, J · 2017
Earlier work this paper cites.
Video visual relation detection
Shang, X., Ren, T., Guo, J., Zhang, H., and Chua, T.-S · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y · 2017
Cited alongside, same era.
Pyscenedetect: Intelligent scene cut detection and video splitting tool
Castellano, B · 2018
Cited alongside, same era.
Videobert: A joint model for video and language representation learning
Sun, C., Myers, A., Vondrick, C., Murphy, K., and Schmid, C · 2019
Cited alongside, same era.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Yu, Z., Xu, D., Yu, J., Yu, T., Zhao, Z., Zhuang, Y., and Tao, D · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Vtimellm: Empower llm to grasp video moments
Huang, B., Wang, X., Chen, H., Song, Z., and Zhu, W · 2023
Later among the works it cites.
Univtg: Towards unified video-language temporal grounding
Lin, K. Q., Zhang, P., Chen, J., Pramanick, S., Gao, D., Wang, A. J., Yan, R., and Shou, M. Z · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Li, C., Yang, J., Su, H., Zhu, J., et al · 2023
Later among the works it cites.
Valley: Video assistant with large language model enhanced ability
Luo, R., Zhao, Z., Yang, M., Dong, J., Qiu, M., Lu, P., Wang, T., and Wei, Z · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Tsp: Temporally-sensitive pretraining of video encoders for localization tasks
Alwassel, H., Giancola, S., and Ghanem, B · 2021
Cited alongside, same era.
Dual encoding for video retrieval by text
Dong, J., Li, X., Xu, C., Yang, X., Yang, G., Wang, X., and Wang, M · 2021
Cited alongside, same era.
Detecting moments and highlights in videos via natural language queries
Lei, J., Berg, T. L., and Bansal, M · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Unified fully and timestamp supervised temporal action segmentation via sequence to sequence translation
Behrmann, N., Golestaneh, S. A., Kolter, Z., Gall, J., and Noroozi, M · 2022
Cited alongside, same era.
Fast and unsupervised action boundary detection for action segmentation
Du, Z., Wang, X., Zhou, G., and Wang, Q · 2022
Cited alongside, same era.
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Later among the works it cites.
Controlretriever: Harnessing the power of instructions for controllable retrieval
Pan, K., Li, J., Song, H., Fei, H., Ji, W., Zhang, S., Lin, J., Liu, X., and Tang, S · 2023
Later among the works it cites.
Timechat: A time-sensitive multimodal large language model for long video understanding
Ren, S., Yao, L., Li, S., Sun, X., and Hou, L · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H., Li, X., and Bing, L · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Fact: Teaching mllms with faithful, concise and transferable rationales
Gao, M., Chen, S., Pang, L., Yao, Y., Dang, J., Zhang, W., Li, J., Tang, S., Zhuang, Y., and Chua, T.-S · 2024
Closest in time.
Worldgpt: Empowering llm as multimodal world model
Ge, Z., Huang, H., Zhou, M., Li, J., Wang, G., Tang, S., and Zhuang, Y · 2024
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Closest in time.
Auto-encoding morph-tokens for multimodal llm, 2024
Pan, K., Tang, S., Li, J., Fan, Z., Chow, W., Yan, S., Chua, T.-S., Zhuang, Y., and Zhang, H · 2024
Closest in time.
Hyperllava: Dynamic visual and language expert tuning for multimodal large language models
Zhang, W., Lin, T., Liu, J., Shu, F., Li, H., Zhang, L., Wanggui, H., Zhou, H., Lv, Z., Jiang, H., et al · 2024
Closest in time.