Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., andre Matten, M., and Berg, T. L · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
World of bits: An open-domain platform for web-based agents
Shi, T., Karpathy, A., Fan, L. J., Hernández, J. Z., and Liang, P · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Zhou, B., Zhao, H., Puig, X., Fidler, S., Barriuso, A., and Torralba, A · 2017
Earlier work this paper cites.
Reinforcement learning on web interfaces using workflow-guided exploration
Original
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., and Liang, P · 2018
Earlier work this paper cites.
Measuring massive multitask language understanding
Original
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D. X., and Steinhardt, J · 2020
Earlier work this paper cites.
Flin: A flexible natural language interface for web navigation
Original
Mazumder, S. and Riva, O · 2020
Earlier work this paper cites.
Grounding ‘grounding’in nlp
Chandu, K. R., Bisk, Y., and Black, A. W · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
How much can clip benefit vision-and-language tasks?
Original
Shen, S., Li, L. H., Tan, H., Bansal, M., Rohrbach, A., Chang, K.-W., Yao, Z., and Keutzer, K · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Original
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Original
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Valter, D., Narang, S., Mishra, G., Yu, A. W., Zhao, V., Huang, Y., Dai, A. M., Yu, H., Petrov, S., hsin Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Earlier work this paper cites.
Understanding html with large language models
Gur, I., Nachum, O., Miao, Y., Safdari, M., Huang, A., Chowdhery, A., Narang, S., Fiedel, N., and Faust, A · 2022
Earlier work this paper cites.