Fetching the paper…
Reading the bibliography…
Attention mechanism has gained huge popularity due to its effectiveness in achieving high accuracy in different domains.
Malinowski, Mateusz, and Mario Fritz. ”A multi-world approach to question answering about real-world scenes based on uncertain input.” In Advances in neural information processing systems, pp. 1682-1690. 2014
2014
Earlier work this paper cites.
Xu, Kelvin, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. ”Show, attend and tell: Neural image caption generation with visual attention.” In International conference on machine learning, pp. 2048-2057. 2015
2015
Earlier work this paper cites.
Donahue, Jeffrey, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. ”Long-term recurrent convolutional networks for visual recognition and description.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2625-2634. 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Yang, Zichao, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. ”Stacked attention networks for image question answering.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 21-29. 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Vinyals, Oriol, Alexander Toshev, Samy Bengio, and Dumitru Erhan. ”Show and tell: Lessons learned from the 2015 mscoco image captioning challenge.” IEEE transactions on pattern analysis and machine intelligence 39, no. 4 (2016): 652-663
2016
Cited alongside, same era.
Xu, Huijuan, and Kate Saenko. ”Ask, attend and answer: Exploring question-guided spatial attention for visual question answering.” In European Conference on Computer Vision, pp. 451-466. Springer, Cham, 2016
2016
Cited alongside, same era.
Lu, Jiasen, Jianwei Yang, Dhruv Batra, and Devi Parikh. ”Hierarchical question-image co-attention for visual question answering.” In Advances in neural information processing systems, pp. 289-297. 2016
2016
Cited alongside, same era.
Nam, Hyeonseob, Jung-Woo Ha, and Jeonghee Kim. ”Dual attention networks for multimodal reasoning and matching.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 299-307. 2017
2017
Cited alongside, same era.
Anderson, Peter, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. ”Bottom-up and top-down attention for image captioning and visual question answering.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6077-6086. 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
Nguyen, Duy-Kien, and Takayuki Okatani. ”Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6087-6096. 2018
2018
Later among the works it cites.
Kim, Jin-Hwa, Jaehyun Jun, and Byoung-Tak Zhang. ”Bilinear attention networks.” In Advances in Neural Information Processing Systems, pp. 1564-1574. 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. ”Attention is all you need.” In Advances in neural information processing systems, pp. 5998-6008. 2017
2017
Cited alongside, same era.
Yu, Zhou, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao. ”Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering.” IEEE transactions on neural networks and learning systems 29, no. 12 (2018): 5947-5959
2018
Cited alongside, same era.
Zhao, Zhou, Zhu Zhang, Shuwen Xiao, Zhou Yu, Jun Yu, Deng Cai, Fei Wu, and Yueting Zhuang. ”Open-Ended Long-form Video Question Answering via Adaptive Hierarchical Reinforced Networks.” In IJCAI, pp. 3683-3689. 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Later among the works it cites.
Lee, Juho, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. ”Set transformer: A framework for attention-based permutation-invariant neural networks.” In International Conference on Machine Learning, pp. 3744-3753. 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Yu, Zhou, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. ”Deep modular co-attention networks for visual question answering.” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6281-6290. 2019
2019
Later among the works it cites.
Sur, Chiranjib. ”RBN: enhancement in language attribute prediction using global representation of natural language transfer learning technology like Google BERT.” SN Applied Sciences 2, no. 1 (2020): 22
2020
Closest in time.