Ucf101: A dataset of 101 human actions classes from videos in the wild
Original
Soomro, K., Zamir, A. R., and Shah, M · 2012
Earlier work this paper cites.
A benchmark dataset and evaluation methodology for video object segmentation
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., and Sorkine-Hornung, A · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Original
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R., and Wilson, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
A short note about kinetics-600
Original
Carreira, J., Noland, E., Banki-Horvath, A., Hillier, C., and Zisserman, A · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Original
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Ranftl, R., Lasinger, K., Hafner, D., Schindler, K., and Koltun, V · 2020
Earlier work this paper cites.
Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan
Saito, M., Saito, S., Koyama, M., and Kobayashi, S · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Original
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Earlier work this paper cites.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Original
Meng, C., He, Y., Song, Y., Song, J., Wu, J., Zhu, J.-Y., and Ermon, S · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Original
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Earlier work this paper cites.
A step toward more inclusive people annotations for fairness
Schumann, C., Ricco, S., Prabhu, U., Ferrari, V., and Pantofaru, C · 2021
Earlier work this paper cites.
Godiva: Generating open-domain videos from natural descriptions
Original
Wu, C., Huang, L., Zhang, Q., Li, B., Ji, L., Yang, F., Sapiro, G., and Duan, N · 2021
Earlier work this paper cites.
Videogpt: Video generation using vq-vae and transformers
Original
Yan, W., Zhang, Y., Abbeel, P., and Srinivas, A · 2021
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Earlier work this paper cites.
Maskgit: Masked generative image transformer
Chang, H., Zhang, H., Jiang, L., Liu, C., and Freeman, W. T · 2022
Earlier work this paper cites.
PaLM: Scaling language modeling with pathways
Original
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Earlier work this paper cites.
GLaMs: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Earlier work this paper cites.