Fetching the paper…
Reading the bibliography…
Misalignment between model predictions and intended usage can be detrimental for the deployment of computer vision models.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
Lafferty, J., McCallum, A., and Pereira, F. C · 2001
Earlier work this paper cites.
Efficient inference in fully connected crfs with gaussian edge potentials
Krähenbühl, P. and Koltun, V · 2011
Earlier work this paper cites.
Decision tree fields
Nowozin, S., Rother, C., Bagon, S., Sharp, T., Yao, B., and Kohli, P · 2011
Earlier work this paper cites.
Closed-form training of conditional random fields for large scale image segmentation
Kolesnikov, A., Guillaumin, M., Ferrari, V., and Lampert, C. H · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., and Zaremba, W · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
Minimum risk training for neural machine translation
Shen, S., Cheng, Y., He, Z., He, W., Wu, H., Sun, M., and Liu, Y · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R., Lawrence Zitnick, C., and Parikh, D · 2015
Earlier work this paper cites.
Reinforcement learning for visual object detection
Mathe, S., Pirinen, A., and Sminchisescu, C · 2016
Earlier work this paper cites.
Training deep neural networks via direct loss minimization
Song, Y., Schwing, A., Urtasun, R., et al · 2016
Earlier work this paper cites.
End-to-end training of object class detectors for mean average precision
Henderson, P. and Ferrari, V · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P · 2017
Cited alongside, same era.
Self-critical sequence training for image captioning
Rennie, S. J., Marcheret, E., Mroueh, Y., Ross, J., and Goel, V · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Can neural machine translation be improved with user feedback?
Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Metricopt: Learning to optimize black-box evaluation metrics
Huang, C., Zhai, S., Guo, P., and Susskind, J · 2021
Later among the works it cites.
Deep reinforcement learning in computer vision: a comprehensive survey
Le, N., Rathour, V. S., Yamazaki, K., Luu, K., and Savvides, M · 2021
Later among the works it cites.
Machine translation decoding beyond beam search
Leblond, R., Alayrac, J.-B., Sifre, L., Pislar, M., Jean-Baptiste, L., Antonoglou, I., Simonyan, K., and Vinyals, O · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep reinforcement learning for sequence-to-sequence models
Keneshloo, Y., Shi, T., Ramakrishnan, N., and Reddy, C. K · 2019
Cited alongside, same era.
Panoptic segmentation
Kirillov, A., He, K., Girshick, R., Rother, C., and Dollar, P · 2019
Cited alongside, same era.
Generalization in generation: A closer look at exposure bias
Schmidt, F · 2019
Cited alongside, same era.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., and Sun, J · 2019
Cited alongside, same era.
On NMT search errors and model errors: Cat got your tongue?
Stahlberg, F. and Byrne, B · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Cited alongside, same era.
Steiner, A., Kolesnikov, A., Zhai, X., Wightman, R., Uszkoreit, J., and Beyer, L · 2021
Later among the works it cites.
Big Vision
Beyer, L., Zhai, X., and Kolesnikov, A · 2022
Later among the works it cites.
Pix2seq: A language modeling framework for object detection
Chen, T., Saxena, S., Li, L., Fleet, D. J., and Hinton, G · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Glaese, A., McAleese, N., Trebacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al · 2022
Later among the works it cites.
Expansionnet v2: Block static expansion in fast end to end training for image captioning
Hu, J. C., Cavicchioli, R., and Capotondi, A · 2022
Later among the works it cites.
UVim: A unified modeling approach for vision with learned guiding codes
Kolesnikov, A., Pinto, A. S., Beyer, L., Zhai, X., Harmsen, J. J., and Houlsby, N · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
End-to-end transformer based model for image captioning
Wang, Y., Xu, J., and Sun, Y · 2022
Later among the works it cites.
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.