Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) and Large Multi-modality Models (LMMs) have demonstrated remarkable decision masking capabilities on a variety of tasks.
Forty-five years of split-brain research and still going strong
Gazzaniga, M. S · 2005
Earlier work this paper cites.
Vision meets robotics: The kitti dataset
Geiger, A., Lenz, P., Stiller, C., and Urtasun, R · 2013
Earlier work this paper cites.
Left brain, right brain: facts and fantasies
Corballis, M. C · 2014
Earlier work this paper cites.
Sequence to sequence-video to text
Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., and Saenko, K · 2015
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Self-supervised visual planning with temporal skip connections
Ebert, F., Finn, C., Lee, A. X., and Levine, S · 2017
Earlier work this paper cites.
End-to-end learning of driving models from large-scale video datasets
Xu, H., Gao, Y., Yu, F., and Darrell, T · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Kim, J., Rohrbach, A., Darrell, T., Canny, J., and Akata, Z · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
Tulyakov, S., Liu, M.-Y., Yang, X., and Kautz, J · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
STAR: A benchmark for situated reasoning in real-world videos
Wu, B., Yu, S., Chen, Z., Tenenbaum, J. B., and Gan, C · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Later among the works it cites.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Later among the works it cites.
Videofusion: Decomposed diffusion models for high-quality video generation
Luo, Z., Chen, D., Zhang, Y., Huang, Y., Wang, L., Shen, Y., Zhao, D., Zhou, J., and Tan, T · 2023
Later among the works it cites.
Verbs in action: Improving verb understanding in video-language models
Momeni, L., Caron, M., Nagrani, A., Zisserman, A., and Schmid, C · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., et al · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2022
Cited alongside, same era.
MCVD - masked conditional video diffusion for prediction, generation, and interpolation
Voleti, V., Jolicoeur-Martineau, A., and Pal, C · 2022
Cited alongside, same era.
Internvideo: General video foundation models via generative and discriminative learning
Wang, Y., Li, K., Li, Y., He, Y., Huang, B., Zhao, Z., Zhang, H., Xu, J., Liu, Y., Wang, Z., et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Compositional foundation models for hierarchical planning
Ajay, A., Han, S., Du, Y., Li, S., Gupta, A., Jaakkola, T., Tenenbaum, J., Kaelbling, L., Srivastava, A., and Agrawal, P · 2023
Cited alongside, same era.
A dynamic multi-scale voxel flow network for video prediction
Hu, X., Huang, Z., Huang, A., Xu, J., and Zhou, S · 2023
Cited alongside, same era.
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Later among the works it cites.
Nuwa-xl: Diffusion over diffusion for extremely long video generation
Yin, S., Wu, C., Yang, H., Wang, J., Wang, X., Ni, M., Yang, Z., Li, L., Liu, S., Yang, F., et al · 2023
Later among the works it cites.
Self-chained image-language model for video localization and question answering
Yu, S., Cho, J., Yadav, P., and Bansal, M · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.