Fetching the paper…
Reading the bibliography…
Many structured prediction problems (particularly in vision and language domains) are ambiguous, with multiple outputs being correct for an input - e.g.
Semi-supervised learning using gaussian fields and harmonic functions
Zhu, Xiaojin, Ghahramani, Zoubin, Lafferty, John, et al · 2003
Earlier work this paper cites.
Semi-supervised learning literature survey
Zhu, Xiaojin · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, Jia, Dong, Wei, Socher, Richard, Li, Li-Jia, Li, Kai, and Fei-Fei, Li · 2009
Earlier work this paper cites.
Multi-label learning with incomplete class assignments
Bucak, Serhat Selcuk, Jin, Rong, and Jain, Anil K · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Wah, Catherine, Branson, Steve, Welinder, Peter, Perona, Pietro, and Belongie, Serge · 2011
Earlier work this paper cites.
Diverse M-Best Solutions in Markov Random Fields
Batra, Dhruv, Yadollahpour, Payman, Guzman-Rivera, Abner, and Shakhnarovich, Gregory · 2012
Earlier work this paper cites.
Multiple Choice Learning: Learning to Produce Multiple Structured Outputs
Guzman-Rivera, Abner, Batra, Dhruv, and Kohli, Pushmeet · 2012
Earlier work this paper cites.
Deep learning via semi-supervised embedding
Weston, Jason, Ratle, Frédéric, Mobahi, Hossein, and Collobert, Ronan · 2012
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Hodosh, Micah, Young, Peter, and Hockenmaier, Julia · 2013
Earlier work this paper cites.
Exploring svm for image annotation in presence of confusing labels
Verma, Yashaswi and Jawahar, CV · 2013
Earlier work this paper cites.
Large-scale object classification using label relation graphs
Deng, Jia, Ding, Nan, Jia, Yangqing, Frome, Andrea, Murphy, Kevin, Bengio, Samy, Li, Yuan, Neven, Hartmut, and Adam, Hartwig · 2014
Earlier work this paper cites.
Efficiently enforcing diversity in multi-output structured prediction
Guzman-Rivera, Abner, Kohli, Pushmeet, Batra, Dhruv, and Rutenbar, Rob · 2014
Earlier work this paper cites.
Deepwalk: Online learning of social representations
Perozzi, Bryan, Al-Rfou, Rami, and Skiena, Steven · 2014
Earlier work this paper cites.
Submodular meets structured: Finding diverse subsets in exponentially-large structured item sets
Prasad, Adarsh, Jegelka, Stefanie, and Batra, Dhruv · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, Peter, Lai, Alice, Hodosh, Micah, and Hockenmaier, Julia · 2014
Earlier work this paper cites.
Large-scale multi-label learning with missing labels
Yu, Hsiang-Fu, Jain, Prateek, Kar, Purushottam, and Dhillon, Inderjit · 2014
Cited alongside, same era.
Exploring nearest neighbor approaches for image captioning
Devlin, Jacob, Gupta, Saurabh, Girshick, Ross, Mitchell, Margaret, and Zitnick, C Lawrence · 2015
Cited alongside, same era.
Image specificity
Jas, Mainak and Parikh, Devi · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Karpathy, Andrej and Fei-Fei, Li · 2015
Cited alongside, same era.
ADAM: A method for stochastic optimization
Kingma, Diederik P and Adam, Jimmy Ba · 2015
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
Li, Jiwei, Galley, Michel, Brockett, Chris, Gao, Jianfeng, and Dolan, Bill · 2015
Reference based lstm for image captioning
Chen, Minghai, Ding, Guiguang, Zhao, Sicheng, Chen, Hui, Liu, Qiang, and Han, Jungong · 2017
Later among the works it cites.
Towards better decoding and language model integration in sequence to sequence models
Chorowski, Jan and Jaitly, Navdeep · 2017
Later among the works it cites.
Towards diverse and natural image descriptions via a conditional gan
Dai, Bo, Lin, Dahua, Urtasun, Raquel, and Fidler, Sanja · 2017
Later among the works it cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, Hakan, Khosravi, Khashayar, and Socher, Richard · 2017
Later among the works it cites.
Creativity: Generating diverse questions using variational autoencoders
Jain, Unnat, Zhang, Ziyu, and Schwing, Alexander · 2017
Later among the works it cites.
Knowing when to look: Adaptive attention via a visual sentinel for image captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Szegedy, Christian, Vanhoucke, Vincent, Ioffe, Sergey, Shlens, Jonathon, and Wojna, Zbigniew · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
Vedantam, Ramakrishna, Lawrence Zitnick, C, and Parikh, Devi · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron, Salakhudinov, Ruslan, Zemel, Rich, and Bengio, Yoshua · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Anderson, Peter, Fernando, Basura, Johnson, Mark, and Gould, Stephen · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian · 2016
Cited alongside, same era.
Stochastic multiple choice learning for training diverse deep ensembles
Lee, Stefan, Purushwalkam, Senthil, Cogswell, Michael, Ranjan, Viresh, Crandall, David J., and Batra, Dhruv · 2016
Cited alongside, same era.
Lu, Jiasen, Xiong, Caiming, Parikh, Devi, and Socher, Richard · 2017
Later among the works it cites.
Text-guided attention model for image captioning
Mun, Jonghwan, Cho, Minsu, and Han, Bohyung · 2017
Later among the works it cites.
Regularizing neural networks by penalizing confident output distributions
Pereyra, Gabriel, Tucker, George, Chorowski, Jan, Kaiser, Łukasz, and Hinton, Geoffrey · 2017
Later among the works it cites.
Self-critical sequence training for image captioning
Rennie, Steven J, Marcheret, Etienne, Mroueh, Youssef, Ross, Jarret, and Goel, Vaibhava · 2017
Later among the works it cites.
Speaking the same language: Matching machine to human captions by adversarial training
Shetty, Rakshith, Rohrbach, Marcus, Hendricks, Lisa Anne, Fritz, Mario, and Schiele, Bernt · 2017
Later among the works it cites.
Stochastic segmentation trees for multiple ground truths
Snell, Jake and Zemel, Richard S · 2017
Later among the works it cites.
Zero-shot learning-a comprehensive evaluation of the good, the bad and the ugly
Xian, Yongqin, Lampert, Christoph H, Schiele, Bernt, and Akata, Zeynep · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol · 2017
Later among the works it cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
Vijayakumar, Ashwin K, Cogswell, Michael, Selvaraju, Ramprasath R, Sun, Qing, Lee, Stefan, Crandall, David, and Batra, Dhruv · 2018
Closest in time.