Microsoft COCO captions: Data collection and evaluation server
Original
Chen, X., Fang, H., Lin, T., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D., Irpan, A., Pastor, P., Ibarz, J., Herzog, A., Jang, E., Quillen, D., Holly, E., Kalakrishnan, M., Vanhoucke, V., et al · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Original
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Original
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Deep visual reasoning: Learning to predict action sequences for task and motion planning from an initial scene image
Driess, D., Ha, J.-S., and Toussaint, M · 2020
Earlier work this paper cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
Original
Lynch, C. and Sermanet, P · 2020
Earlier work this paper cites.
Robots that use language
Tellex, S., Gopalan, N., Kress-Gazit, H., and Matuszek, C · 2020
Earlier work this paper cites.
Unified vision-language pre-training for image captioning and vqa
Zhou, L., Palangi, H., Zhang, L., Hu, H., Corso, J., and Gao, J · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Original
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Original
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Trocr: Transformer-based optical character recognition with pre-trained models
Original
Li, M., Lv, T., Chen, J., Cui, L., Lu, Y., Florencio, D., Zhang, C., Li, Z., and Wei, F · 2021
Earlier work this paper cites.
Pretrained transformers as universal computation engines
Original
Lu, K., Grover, A., Abbeel, P., and Mordatch, I · 2021
Earlier work this paper cites.