Fetching the paper…
Reading the bibliography…
In this paper, we study the compositional learning of images and texts for image retrieval.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Multimodal residual learning for visual qa
Jin-Hwa Kim, Sang-Woo Lee, Donghyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Image question answering using convolutional neural network with dynamic parameter prediction
Hyeonwoo Noh, Paul Hongsuck Seo, and Bohyung Han · 2016
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
In defense of the triplet loss for person re-identification
Alexander Hermans, Lucas Beyer, and Bastian Leibe · 2017
Earlier work this paper cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao · 2017
Cited alongside, same era.
Memory-augmented attribute manipulation networks for interactive fashion search
Bo Zhao, Jiashi Feng, Xiao Wu, and Shuicheng Yan · 2017
Cited alongside, same era.
Learning attribute representations with localization for flexible fashion search
Kenan E Ak, Ashraf A Kassim, Joo Hwee Lim, and Jo Yew Tham · 2018
Cited alongside, same era.
Dialog-based interactive image retrieval
Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, Gerald Tesauro, and Rogerio Schmidt Feris · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Cited alongside, same era.
A hierarchical graph network for 3d object detection on point clouds
Jintai Chen, Biwen Lei, Qingyu Song, Haochao Ying, Danny Z Chen, and Jian Wu · 2020
Later among the works it cites.
Learning joint visual semantic matching embeddings for language-guided retrieval
Yanbei Chen and Loris Bazzani · 2020
Later among the works it cites.
Image search with text feedback by visiolinguistic attention learning
Yanbei Chen, Shaogang Gong, and Loris Bazzani · 2020
Later among the works it cites.
Skeleton-based action recognition with shift graph convolutional network
Ke Cheng, Yifan Zhang, Xiangyu He, Weihan Chen, Jian Cheng, and Hanqing Lu · 2020
Later among the works it cites.
Modality-agnostic attention fusion for visual search with text feedback
Eric Dodds, Jack Culpepper, Simao Herdade, Yang Zhang, and Kofi Boakye · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao · 2018
Cited alongside, same era.
Block: Bilinear superdiagonal fusion for visual question answering and visual relationship detection
Hedi Ben-Younes, Remi Cadene, Nicolas Thome, and Matthieu Cord · 2019
Cited alongside, same era.
Multi-label image recognition with graph convolutional networks
Zhao-Min Chen, Xiu-Shen Wei, Peng Wang, and Yanwen Guo · 2019
Cited alongside, same era.
Neural naturalist: generating fine-grained image comparisons
Maxwell Forbes, Christine Kaeser-Chen, Piyush Sharma, and Serge Belongie · 2019
Cited alongside, same era.
Fashion iq: A new dataset towards retrieving images by natural language feedback
Xiaoxiao Guo, Hui Wu, Yupeng Gao, Steven Rennie, and Rogerio Feris · 2019
Cited alongside, same era.
Guided similarity separation for image retrieval
Chundi Liu, Guangwei Yu, Maksims Volkovs, Cheng Chang, Himanshu Rai, Junwei Ma, and Satya Krishna Gorti · 2019
Cited alongside, same era.
Semi-supervised feature-level attribute manipulation for fashion image retrieval
Minchul Shin, Sanghyuk Park, and Taeksoo Kim · 2019
Cited alongside, same era.
Surgan Jandial, Ayush Chopra, Pinkesh Badjatiya, Pranit Chawla, Mausoom Sarkar, and Balaji Krishnamurthy · 2020
Later among the works it cites.
Od-gcn: Object detection boosted by knowledge gcn
Zheng Liu, Zidong Jiang, Wei Feng, and Hui Feng · 2020
Later among the works it cites.
Fashion-iq 2020 challenge 2nd place team’s solution
Minchul Shin, Yoonjae Cho, and Seongwuk Hong · 2020
Later among the works it cites.
View-gcn: View-based graph convolutional network for 3d shape analysis
Xin Wei, Ruixuan Yu, and Jian Sun · 2020
Later among the works it cites.
L2-gcn: Layer-wise and learned efficient training of graph convolutional networks
Yuning You, Tianlong Chen, Zhangyang Wang, and Yang Shen · 2020
Later among the works it cites.
Curlingnet: Compositional learning between images and text for fashion iq data
Youngjae Yu, Seunghwan Lee, Yuncheol Choi, and Gunhee Kim · 2020
Later among the works it cites.
Cycled compositional learning between images and text
Jongseok Kim, Youngjae Yu, Seunghwan Lee, et al · 2021
Closest in time.
Cosmo: Content-style modulation for image retrieval with text feedback
Seungmin Lee, Dongwan Kim, and Bohyung Han · 2021
Closest in time.
Image retrieval on real-life images with pre-trained vision-and-language models
Zheyuan Liu, Cristian Rodriguez-Opazo, Damien Teney, and Stephen Gould · 2021
Closest in time.