Fetching the paper…
Reading the bibliography…
Cognitive grammar suggests that the acquisition of language grammar is grounded within visual structures.
Syntactic pattern recognition and applications
K. Fu · 1968
Earlier work this paper cites.
Hierarchy theory: The challenge of complex systems
HA Simon and Howard H Pattee · 1973
Earlier work this paper cites.
Syntactic shape recognition using attributed grammars
K. C. You and K. Fu · 1978
Earlier work this paper cites.
Trainable grammars for speech recognition
J. Baker · 1979
Earlier work this paper cites.
Foundations of cognitive grammar
Ronald W. Langacker · 1983
Earlier work this paper cites.
Asymptotic methods in statistical decision theory
L. L. Cam · 1986
Earlier work this paper cites.
An introduction to cognitive grammar
Ronald W. Langacker · 1986
Earlier work this paper cites.
The body in the mind: the bodily basis of meaning
M. Johnson · 1987
Earlier work this paper cites.
Defining and parsing visual languages with layered graph grammars 1
J. R EKERS and A. S CHÜRR · 1997
Earlier work this paper cites.
A generative constituent-context model for improved grammar induction
D. Klein and Christopher D. Manning · 2002
Earlier work this paper cites.
The estimation of stochastic context-free grammars using the inside-outside algorithm
Vladimir Solmon · 2003
Earlier work this paper cites.
Image parsing: unifying segmentation, detection, and recognition
Zhuowen Tu, Xiangrong Chen, A. Yuille, and S. Zhu · 2003
Earlier work this paper cites.
A stochastic grammar of images
Song-Chun Zhu and David Mumford · 2006
Earlier work this paper cites.
The shared logistic normal distribution for grammar induction
Shay B. Cohen and Noah A. Smith · 2008
Earlier work this paper cites.
Bottom-up/top-down image parsing with attribute grammar
F. Han and S. Zhu · 2009
Earlier work this paper cites.
Viterbi training improves unsupervised dependency parsing
Valentin I. Spitkovsky, H. Alshawi, Dan Jurafsky, and Christopher D. Manning · 2010
Earlier work this paper cites.
A numerical study of the bottom-up and top-down inference processes in and-or graphs
Tianfu Wu and S. Zhu · 2010
Earlier work this paper cites.
Literal and metaphorical sense identification through concrete and abstract context
Peter D. Turney, Y. Neuman, Dan Assaf, and Yohai Cohen · 2011
Earlier work this paper cites.
Learning and-or templates for object recognition and detection
Zhangzhang Si and S. Zhu · 2013
Cited alongside, same era.
Unsupervised structure learning of stochastic and-or grammars
K. Tu, Maria Pavlovskaia, and S. Zhu · 2013
Cited alongside, same era.
Deep fragment embeddings for bidirectional image sentence mapping
Andrej Karpathy, Armand Joulin, and Li Fei-Fei · 2014
Cited alongside, same era.
Improving multi-modal representations using image dispersion: Why less is sometimes more
Douwe Kiela, Felix Hill, A. Korhonen, and Stephen Clark · 2014
Cited alongside, same era.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Unsupervised latent tree induction with deep inside-outside recursive autoencoders
Andrew Drozdov, Pat Verga, Mohit Yadav, Mohit Iyyer, and A. McCallum · 2019
Later among the works it cites.
Compound probabilistic context-free grammars for grammar induction
Yoon Kim, Chris Dyer, and Alexander M. Rush · 2019
Later among the works it cites.
Unsupervised recurrent neural network grammars
Yoon Kim, Alexander M. Rush, L. Yu, Adhiguna Kuncoro, Chris Dyer, and Gábor Melis · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Multimodal convolutional neural networks for matching image and sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Inside-outside and forward-backward algorithms are just backprop (tutorial paper)
Jason Eisner · 2016
Cited alongside, same era.
Hierarchical co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E. Hinton · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Structurenet: Hierarchical graph networks for 3d shape generation
Kaichun Mo, P. Guerrero, L. Yi, H. Su, Peter Wonka, Niloy Mitra, and L. Guibas · 2019
Later among the works it cites.
Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding
Kaichun Mo, Shilin Zhu, Angel X Chang, Li Yi, Subarna Tripathi, Leonidas J Guibas, and Hao Su · 2019
Later among the works it cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron C. Courville · 2019
Later among the works it cites.
Visually grounded neural syntax acquisition
Haoyue Shi, Jiayuan Mao, Kevin Gimpel, and Karen Livescu · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Partnet: A recursive part decomposition network for fine-grained and hierarchical shape segmentation
Fenggen Yu, Kun Liu, Yan Zhang, Chenyang Zhu, and K. Xu · 2019
Later among the works it cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Later among the works it cites.
Scan: Learning to classify images without labels
Wouter Van Gansbeke, Simon Vandenhende, S. Georgoulis, M. Proesmans, and L. Gool · 2020
Later among the works it cites.
Closed loop neural-symbolic learning via integrating neural perception, grammar parsing, and symbolic reasoning
Qing Li, Siyuan Huang, Yining Hong, Yixin Chen, Ying Nian Wu, and Song-Chun Zhu · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Later among the works it cites.
Visually grounded compound pcfgs
Yanpeng Zhao and Ivan Titov · 2020
Later among the works it cites.
How to represent part-whole hierarchies in a neural network
Geoffrey E. Hinton · 2021
Closest in time.