Fetching the paper…
Reading the bibliography…
Designed for tracking user goals in dialogues, a dialogue state tracker is an essential component in a dialogue system.
Dialog state tracking: A neural reading comprehension approach
Shuyang Gao, Abhishek Sethi, Sanchit Aggarwal, Tagyoung Chung, and Dilek Hakkani-Tur. 2019 · 1908
Earlier work this paper cites.
Visual objects in context
Moshe Bar. 2004 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Multimodal dialogue state tracking by qa approach with data augmentation
Xiangyang Mou, Brandyn Sigouin, Ian Steenstra, and Hui Su. 2020 · 2007
Earlier work this paper cites.
Selecting and perceiving multiple visual objects
Yaoda Xu and Marvin M Chun. 2009 · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth. 2010 · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
The hidden information state model: A practical framework for pomdp-based spoken dialogue management
Steve Young, Milica Gašić, Simon Keizer, François Mairesse, Jost Schatzmann, Blaise Thomson, and Kai Yu. 2010 · 2010
Earlier work this paper cites.
Robust dialog state tracking using delexicalised recurrent neural networks and unsupervised adaptation
Matthew Henderson, Blaise Thomson, and Steve J. Young. 2014b · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederick P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
A dataset for movie description
Anna Rohrbach, Marcus Rohrbach, Niket Tandon, and Bernt Schiele. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Recurrent support vector machines for slot tagging in spoken language understanding
Yangyang Shi, Kaisheng Yao, Hu Chen, Dong Yu, Yi-Cheng Pan, and Mei-Yuh Hwang. 2016 · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Earlier work this paper cites.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Evaluating visual conversational agents via cooperative human-ai games
Prithvijit Chattopadhyay, Deshraj Yadav, Viraj Prabhu, Arjun Chandrasekaran, Abhishek Das, Stefan Lee, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Learning cooperative visual dialog agents with deep reinforcement learning
Abhishek Das, Satwik Kottur, José M. F. Moura, Stefan Lee, and Dhruv Batra. 2017b · 2017
Earlier work this paper cites.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm De Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville. 2017 · 2017
Cited alongside, same era.
Key-value retrieval networks for task-oriented dialogue
Mihail Eric, Lakshmi Krishnan, Francois Charette, and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
Tall: Temporal activity localization via language query
Jiyang Gao, Chen Sun, Zhenheng Yang, and Ram Nevatia. 2017 · 2017
Cited alongside, same era.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. 2017 · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Later among the works it cites.
Vision-and-dialog navigation
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
Transferable multi-domain state generator for task-oriented dialogue systems
Chien-Sheng Wu, Andrea Madotto, Ehsan Hosseini-Asl, Caiming Xiong, Richard Socher, and Pascale Fung. 2019 · 2019
Later among the works it cites.
Neural symbolic reader: Scalable integration of distributed and symbolic representations for reading comprehension
Xinyun Chen, Chen Liang, Adams Wei Yu, Denny Zhou, Dawn Song, and Quoc V. Le. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. 2017 · 2017
Cited alongside, same era.
Neural belief tracker: Data-driven dialogue state tracking
Nikola Mrkšić, Diarmuid Ó Séaghdha, Tsung-Hsien Wen, Blaise Thomson, and Steve Young. 2017 · 2017
Cited alongside, same era.
Scalable multi-domain dialogue state tracking
Abhinav Rastogi, Dilek Z. Hakkani-Tür, and Larry P. Heck. 2017 · 2017
Cited alongside, same era.
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A network-based end-to-end trainable task-oriented dialogue system
Tsung-Hsien Wen, David Vandyke, Nikola Mrkšić, Milica Gašić, Lina M. Rojas-Barahona, Pei-Hao Su, Stefan Ultes, and Steve Young. 2017 · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Context information supports serial dependence of multiple visual objects across memory episodes
Cora Fischer, Stefan Czoschke, Benjamin Peters, Benjamin Rahm, Jochen Kaiser, and Christoph Bledowski. 2020 · 2020
Later among the works it cites.
Cater: A diagnostic dataset for compositional actions and temporal reasoning
Rohit Girdhar and Deva Ramanan. 2020 · 2020
Later among the works it cites.
Modality-balanced models for visual dialogue
Hyounghun Kim, Hao Tan, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
BiST: Bi-directional spatio-temporal reasoning for video-grounded dialogues
Hung Le, Doyen Sahoo, Nancy Chen, and Steven C.H. Hoi. 2020a · 2020
Later among the works it cites.
UniConv: A unified conversational neural architecture for multi-domain task-oriented dialogues
Hung Le, Doyen Sahoo, Chenghao Liu, Nancy Chen, and Steven C.H. Hoi. 2020b · 2020
Later among the works it cites.
Visual dialogue state tracking for question generation
Wei Pang and Xiaojie Wang. 2020 · 2020
Later among the works it cites.
Two causal principles for improving visual dialog
Jiaxin Qi, Yulei Niu, Jianqiang Huang, and Hanwang Zhang. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Learning object permanence from video
Aviv Shamsian, Ofri Kleinfeld, Amir Globerson, and Gal Chechik. 2020 · 2020
Later among the works it cites.
Find or classify? dual strategy for slot-value predictions on multi-domain dialog state tracking
Jianguo Zhang, Kazuma Hashimoto, Chien-Sheng Wu, Yao Wang, Philip Yu, Richard Socher, and Caiming Xiong. 2020 · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Preview, attend and review: Schema-aware curriculum learning for multi-domain dialogue state tracking
Yinpei Dai, Hangyu Li, Yongbin Li, Jian Sun, Fei Huang, Luo Si, and Xiaodan Zhu. 2021 · 2021
Later among the works it cites.
SIMMC 2.0: A task-oriented dialog dataset for immersive multimodal conversations
Satwik Kottur, Seungwhan Moon, Alborz Geramifard, and Babak Damavandi. 2021 · 2021
Later among the works it cites.
Bridging text and video: A universal multimodal transformer for video-audio scene-aware dialog
Zekang Li, Zongjia Li, Jinchao Zhang, Yang Feng, and Jie Zhou. 2021 · 2021
Later among the works it cites.
Knowledge-aware graph-enhanced GPT-2 for dialogue state tracking
Weizhe Lin, Bo-Hsiang Tseng, and Bill Byrne. 2021 · 2021
Later among the works it cites.
A contextual attention network for multimodal emotion recognition in conversation
Tana Wang, Yaqing Hou, Dongsheng Zhou, and Qiang Zhang. 2021 · 2021
Later among the works it cites.
Factor graph attention
Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander G Schwing. 2019 · 2048
Closest in time.
Leveraging sentence-level information with encoder LSTM for semantic slot filling
Gakuto Kurata, Bing Xiang, Bowen Zhou, and Mo Yu. 2016 · 2083
Closest in time.