Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) systems are emerging from a desire to empower users to ask any natural language question about visual content and receive a valid answer in response.
L. Breiman, “Random forests,” Machine Learning , vol. 45, no. 1, pp. 5–32, 2001
2001
Earlier work this paper cites.
V. S. Sheng, F. Provost, and P. G. Ipeirotis, “Get another label? Improving data quality and data mining using multiple, noisy labelers,” in International Conference on Knowledge Discovery and Data Mining (KDD) , 2008, pp. 614–622
2008
Earlier work this paper cites.
J. P. Bigham, C. Jayant, H. Ji, G. Little, A. Miller, R. C. Miller, R. Miller, A. Tatarowicz, B. White, S. White, and T. Yeh, “Vizwiz: Nearly real-time answers to visual questions,” in ACM symposium on User interface software and technology (UIST) , 2010, pp. 333–342
2010
Earlier work this paper cites.
P. Welinder, S. Branson, S. Belongie, and P. Perona, “The multidimensional wisdom of crowds,” in Advances in Neural Information Processing Systems (NIPS) , 2010, pp. 2424–2432
2010
Earlier work this paper cites.
B. Settles, “Active learning literature survey,” University of Wisconsin, Madison, Tech. Rep., 2010
2010
Earlier work this paper cites.
S. Vijayanarasimhan and K. Grauman, “Cost-sensitive active visual category learning,” in International Journal of Computer Vision (IJCV) , vol. 91, no. 1, 2011, pp. 24–44
2011
Earlier work this paper cites.
M. A. Burton, E. Brady, R. Brewer, C. Neylan, J. P. Bigham, and A. Hurst, “Crowdsourcing subjective fashion advice using VizWiz: Challenges and opportunities,” in ACM SIGACCESS conference on Computers and accessibility (ASSETS) , 2012, pp. 135–142
2012
Earlier work this paper cites.
A. Sheshadri and M. Lease, “SQUARE: A benchmark for research on computing crowd consensus,” in AAAI Conference on Human Computation and Crowdsourcing (HCOMP) , 2013
2013
Earlier work this paper cites.
S. D. Jain and K. Grauman, “Predicting sufficient annotation strength for interactive foreground segmentation,” in IEEE International Conference on Computer Vision (ICCV) , 2013, pp. 1313–1320
2013
Earlier work this paper cites.
W. S. Lasecki, P. Thiha, Y. Zhong, E. Brady, and J. P. Bigham, “Answering visual questions with conversational crowd assistants,” in ACM SIGACCESS Conference on Computers and Accessibility (ASSETS) , 2013
2013
Cited alongside, same era.
M. Malinowski and M. Fritz, “A multi-world approach to question answering about real-world scenes based on uncertain input,” in Advances in Neural Information Processing Systems (NIPS) , 2014, pp. 1682–1690
2014
Cited alongside, same era.
C. H. Lin, Mausam, and D. S. Weld, “To re(label), or not to re(label),” in AAAI Conference on Human Computation and Crowdsourcing (HCOMP) , 2014
2014
Cited alongside, same era.
T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollar, and C. L. Zitnick, “Microsoft COCO: Common objects in context,” in IEEE European Conference on Computer Vision (ECCV) , 2014, pp. 740–755
2014
Cited alongside, same era.
E. Amid and A. Ukkonen, “Multiview triplet embedding: Learning attributes in multiple maps,” in International Conference on Machine Learning (ICML) , 2015, pp. 1472–1480
2015
Later among the works it cites.
A. T. Nguyen, B. C. Wallace, and M. Lease, “Combining crowd and expert labels using decision theoretic active learning,” in AAAI Conference on Human Computation and Crowdsourcing (HCOMP) , 2015
2015
Later among the works it cites.
G. Patterson, G. V. Horn, S. Belongie, P. Perona, and J. Hays, “Tropel: Crowdsourcing detectors with minimal training,” in AAAI Conference on Human Computation and Crowdsourcing (HCOMP) , 2015
2015
Later among the works it cites.
M. Malinowski, M. Rohrbach, and M. Fritz, “Ask your neurons: A neural-based approach to answering questions about images,” in IEEE European Conference on Computer Vision (ECCV) , 2015, pp. 1–9
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 2425–2433
2015
Cited alongside, same era.
L. Yu, E. Park, A. C. Berg, and T. L. Berg, “Visual madlibs: Fill in the blank image generation and question answering,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 2461–2469
2015
Cited alongside, same era.
M. Jas and D. Parikh, “Image specificity,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 2727–2736
2015
Cited alongside, same era.
http://www.bemyeyes.org/, “Be my eyes.”
Cited in the paper.
J. Zhang, S. Ma, M. Sameki, S. Sclaroff, M. Betke, Z. Lin, X. Shen, B. Price, and R. Mech, “Salient object subitizing,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 4045–4054
2015
Later among the works it cites.
J. Lu, X. Lin, D. Batra, and D. Parikh, “Deeper lstm and normalized cnn visual question answering model,” https://github.com/VT-vision-lab/VQA_LSTM_CNN
2015
Later among the works it cites.
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein, “Learning to compose neural networks for question answering,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT) , 2016, pp. 1545—1554
2016
Closest in time.
D. Gurari, S. D. Jain, M. Betke, and K. Grauman, “Pull the plug? predicting if computers or humans should segment images,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 382–391
2016
Closest in time.