Fetching the paper…
Reading the bibliography…
In this paper we introduce LifelongMemory, a new framework for accessing long-form egocentric videographic memory through natural language question answering and retrieval.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Localizing moments in video with natural language
Hendricks, L. A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., and Russell, B · 2017
Earlier work this paper cites.
Kinetics400 dataset: The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., and Zisserman, A · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D., Zhao, Z., Xiao, J., Wu, F., Zhang, H., He, X., and Zhuang, Y · 2017
Earlier work this paper cites.
Slowfast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., and He, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Just ask: Learning to answer questions from millions of narrated videos
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2021
Earlier work this paper cites.
Internvideo-ego4d: A pack of champion solutions to ego4d challenges
Chen, G., Xing, S., Chen, Z., Wang, Y., Li, K., Li, Y., Liu, Y., Wang, J., Zheng, Y.-D., Huang, B., et al · 2022
Earlier work this paper cites.
PaLM: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Earlier work this paper cites.
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Earlier work this paper cites.
Egocentric video-language pretraining@ ego4d challenge 2022
Lin, K. Q., Wang, A. J., Soldan, M., Wray, M., Yan, R., Xu, E. Z., Gao, D., Tu, R., Zhao, W., Kong, W., et al · 2022
Earlier work this paper cites.
Reler@ zju-alibaba submission to the ego4d natural language queries challenge 2022
Liu, N., Wang, X., Li, X., Yang, Y., and Zhuang, Y · 2022
Cited alongside, same era.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Cited alongside, same era.
Language models with image descriptors are strong few-shot video-language learners
Wang, Z., Li, M., Xu, R., Zhou, L., Lei, J., Lin, X., Wang, S., Yang, Z., Zhu, C., Hoiem, D., et al · 2022
Cited alongside, same era.
Zero-shot video question answering via frozen bidirectional language models
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2022
Cited alongside, same era.
OPT: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Towards reasoning in large language models: A survey
Huang, J. and Chang, K. C.-C · 2023
Closest in time.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Mangalam, K., Akshulakov, R., and Malik, J · 2023
Closest in time.
Retrieving-to-answer: Zero-shot video question answering with frozen large language models
Pan, J., Lin, Z., Ge, Y., Zhu, X., Zhang, R., Wang, Y., Qiao, Y., and Li, H · 2023
Closest in time.
NaQ: Leveraging narrations as queries to supervise episodic memory
Ramakrishnan, S. K., Al-Halah, Z., and Grauman, K · 2023
Closest in time.
Action sensitivity learning for the ego4d episodic memory challenge 2023
Shao, J., Wang, X., Quan, R., and Yang, Y · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Achiam, O. J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., et al · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Cited alongside, same era.
Large language models are visual reasoning coordinators
Chen, L., Li, B., Shen, S., Yang, J., Li, C., Keutzer, K., Darrell, T., and Liu, Z · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
PaLM-E: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Cited alongside, same era.
Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., et al · 2023
Cited alongside, same era.
An empirical study of end-to-end video-language transformers with masked visual modeling
Fu, T.-J., Li, L., Gan, Z., Lin, K., Wang, W. Y., Wang, L., and Liu, Z · 2023
Cited alongside, same era.
Vamos: Versatile action models for video understanding
Wang, S., Zhao, Q., Do, M. Q., Agarwal, N., Lee, K., and Sun, C · 2023
Closest in time.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2023
Closest in time.
Mm-react: Prompting chatgpt for multimodal reasoning and action
Yang, Z., Li, L., Wang, J., Lin, K., Azarnasab, E., Ahmed, F., Liu, Z., Liu, C., Zeng, M., and Wang, L · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., Li, C., Xu, Y., Chen, H., Tian, J., Qi, Q., Zhang, J., and Huang, F · 2023
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A., Attarian, M., brian ichter, Choromanski, K. M., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M. S., Sindhwani, V., Lee, J., Vanhoucke, V., and Florence, P · 2023
Closest in time.
Llama 3 model card
AI@Meta · 2024
Closest in time.
Wang, Y., Chen, W., Han, X., Lin, X., Zhao, H., Liu, Y., Zhai, B., Yuan, J., You, Q., and Yang, H · 2024
Closest in time.
Llm as a mastermind: A survey of strategic reasoning with large language models
Zhang, Y., Mao, S., Ge, T., Wang, X., de Wynter, A., Xia, Y., Wu, W., Song, T., Lan, M., and Wei, F · 2024
Closest in time.