Fetching the paper…
Reading the bibliography…
Recent advancements in video generation have significantly improved the ability to synthesize videos from text instructions.
Finegym: A hierarchical video dataset for fine-grained action understanding, 2020
Shao, D., Zhao, Y., Dai, B., and Lin, D · 2004
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges, 2019
Unterthiner, T., van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Earlier work this paper cites.
One tts alignment to rule them all, 2021
Badlani, R., Łancucki, A., Shih, K. J., Valle, R., Ping, W., and Catanzaro, B · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Videoclip: Contrastive pre-training for zero-shot video-text understanding, 2021
Xu, H., Ghosh, G., Huang, P.-Y., Okhonko, D., Aghajanyan, A., Metze, F., Zettlemoyer, L., and Feichtenhofer, C · 2021
Earlier work this paper cites.
Associating objects with transformers for video object segmentation, 2021
Yang, Z., Wei, Y., and Yang, Y · 2021
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2022
Earlier work this paper cites.
Make-a-video: Text-to-video generation without text-video data, 2022
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y · 2022
Earlier work this paper cites.
Mixture-of-experts with expert choice routing, 2022
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A., Chen, Z., Le, Q., and Laudon, J · 2022
Earlier work this paper cites.
Stable video diffusion: Scaling latent video diffusion models to large datasets, 2023
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., Jampani, V., and Rombach, R · 2023
Earlier work this paper cites.
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Chen, Z., Wu, J., Wang, W., Su, W., Chen, G., Xing, S., Zhong, M., Zhang, Q., Zhu, X., Lu, L., Li, B., Luo, P., Lu, T., Qiao, Y., and Dai, J · 2023
Earlier work this paper cites.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models, 2023
Cho, J., Zala, A., and Bansal, M · 2023
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Kirstain, Y., Polyak, A., Singer, U., Matiana, S., Penna, J., and Levy, O · 2023
Cited alongside, same era.
Wu, X., Hao, Y., Sun, K., Chen, Y., Zhu, F., Zhao, R., and Li, H · 2023
Cited alongside, same era.
Training diffusion models with reinforcement learning, 2024
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S · 2024
Cited alongside, same era.
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., and et al., S. B · 2024
Later among the works it cites.
Vidgen-1m: A large-scale dataset for text-to-video generation
Tan, Z., Yang, X., Qin, L., and Li, H · 2024
Later among the works it cites.
Chameleon: Mixed-modal early-fusion foundation models, 2024
Team, C · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2024
Team, G., Anil, R., Borgeaud, S., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., Millican, K., and et al., D. S · 2024
Later among the works it cites.
Diffusion model alignment using direct preference optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cai, W., Jiang, J., Wang, F., Tang, J., Kim, S., and Huang, J · 2024
Cited alongside, same era.
Sora detector: A unified hallucination detection for large text-to-video models, 2024
Chu, Z., Zhang, L., Sun, Y., Xue, S., Wang, Z., Qin, Z., and Ren, K · 2024
Cited alongside, same era.
Fine-grained verifiers: Preference modeling as next-token prediction in vision-language alignment
Cui, C., Zhang, A., Zhou, Y., Chen, Z., Deng, G., Yao, H., and Chua, T.-S · 2024
Cited alongside, same era.
Safesora: Towards safety alignment of text2video generation via a human preference dataset, 2024
Dai, J., Chen, T., Wang, X., Yang, Z., Chen, T., Ji, J., and Yang, Y · 2024
Cited alongside, same era.
Video prediction models as rewards for reinforcement learning
Escontrela, A., Adeniji, A., Yan, W., Jain, A., Peng, X. B., Goldberg, K., Lee, Y., Hafner, D., and Abbeel, P · 2024
Cited alongside, same era.
Cogvlm2: Visual language models for image and video understanding
Hong, W., Wang, W., Ding, M., Yu, W., Lv, Q., Wang, Y., Cheng, Y., Huang, S., Ji, J., Xue, Z., et al · 2024
Cited alongside, same era.
Genai arena: An open evaluation platform for generative models
Jiang, D., Ku, M., Li, T., Ni, Y., Sun, S., Fan, R., and Chen, W · 2024
Cited alongside, same era.
A survey on long video generation: Challenges, methods, and prospects, 2024
Li, C., Huang, D., Lu, Z., Xiao, Y., Pei, Q., and Bai, L · 2024
Cited alongside, same era.
Wallace, B., Dang, M., Rafailov, R., Zhou, L., Lou, A., Purushwalkam, S., Ermon, S., Xiong, C., Joty, S., and Naik, N · 2024
Later among the works it cites.
Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models
Wang, W. and Yang, Y · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation, 2024
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Xu, J., Huang, Y., Cheng, J., Yang, Y., Xu, J., Wang, Y., Duan, W., Yang, S., Jin, Q., Li, S., et al · 2024
Later among the works it cites.
Instructvideo: instructing video diffusion models with human feedback
Yuan, H., Zhang, S., Wang, X., Wei, Y., Feng, T., Pan, Y., Zhang, Y., Liu, Z., Albanie, S., and Ni, D · 2024
Later among the works it cites.
Grape: Generalizing robot policy via preference alignment
Zhang, Z., Zheng, K., Chen, Z., Jang, J., Li, Y., Wang, C., Ding, M., Fox, D., and Yao, H · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all, March 2024
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y · 2024
Later among the works it cites.
Calibrated self-rewarding vision language models
Zhou, Y., Fan, Z., Cheng, D., Yang, S., Chen, Z., Cui, C., Wang, X., Li, Y., Zhang, L., and Yao, H · 2024
Later among the works it cites.