Fetching the paper…
Reading the bibliography…
Existing benchmarks for grounding language in interactive environments either lack real-world linguistic elements, or prove difficult to scale up due to substantial human involvement in the collection of data or feedback signals.
Flask API, 2010
A. Ronacher · 2010
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, Çaglar Gülçehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
ScraperAPI, 2015
D. Ni · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Improving Information Extraction by Acquiring External Evidence with Reinforcement Learning
K. Narasimhan, A. Yala, and R. Barzilay · 2016
Earlier work this paper cites.
End-to-End Goal-Driven Web Navigation
R. Nogueira and K. Cho · 2016
Earlier work this paper cites.
Bidirectional attention flow for machine comprehension
M. Seo, A. Kembhavi, A. Farhadi, and H. Hajishirzi · 2016
Earlier work this paper cites.
Task-Oriented Query Reformulation with Reinforcement Learning
R. Nogueira and K. Cho · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
World of Bits: An Open-Domain platform for web-based agents
T. Shi, A. Karpathy, L. Fan, J. Hernandez, and P. Liang · 2017
Earlier work this paper cites.
Building Natural Language Interfaces to Web APIs
Y. Su, A. H. Awadallah, M. Khabsa, P. Pantel, M. Gamon, and M. Encarnacion · 2017
Earlier work this paper cites.
Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
P. Budzianowski, T.-H. Wen, B.-H. Tseng, I. Casanueva, S. Ultes, O. Ramadan, and M. Gašić · 2018
Earlier work this paper cites.
I. Gur, U. Rueckert, A. Faust, and D. Hakkani-Tur · 2018
Earlier work this paper cites.
Reinforcement Learning on Web Interfaces using Workflow-Guided Exploration
E. Z. Liu, K. Guu, P. Pasupat, T. Shi, and P. Liang · 2018
Earlier work this paper cites.
Natural Language Interfaces with Fine-Grained User Interaction: A Case Study on Web APIs
Y. Su, A. Hassan Awadallah, M. Wang, and R. W. White · 2018
Earlier work this paper cites.
Unsupervised predictive memory in a goal-directed agent
G. Wayne, C.-C. Hung, D. Amos, M. Mirza, A. Ahuja, A. Grabska-Barwinska, J. Rae, P. Mirowski, J. Z. Leibo, A. Santoro, et al · 2018
Earlier work this paper cites.
Go-explore: a new approach for hard-exploration problems
A. Ecoffet, J. Huizinga, J. Lehman, K. O. Stanley, and J. Clune · 2019
Earlier work this paper cites.
Generalization of reinforcement learners with working and episodic memory
M. Fortunato, M. Tan, R. Faulkner, S. Hansen, A. Puigdomènech Badia, G. Buttimore, C. Deck, J. Z. Leibo, and C. Blundell · 2019
Cited alongside, same era.
Dom-q-net: Grounded RL on Structured Language
S. Jia, J. Kiros, and J. Ba · 2019
Cited alongside, same era.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2019
Cited alongside, same era.
A survey of reinforcement learning informed by natural language
J. Luketina, N. Nardelli, G. Farquhar, J. N. Foerster, J. Andreas, E. Grefenstette, S. Whiteson, and T. Rocktäschel · 2019
Cited alongside, same era.
Automatic Task Completion Flows from Web APIs
K. Williams, S. H. Hashemi, and I. Zitouni · 2019
Cited alongside, same era.
Htlm: Hyper-text pre-training and prompting of language models
A. Aghajanyan, D. Okhonko, M. Lewis, M. Joshi, H. Xu, G. Ghosh, and L. Zettlemoyer · 2021
Later among the works it cites.
Adversarial Environment Generation for Learning to Navigate the Web
I. Gur, N. Jaques, K. Malta, M. Tiwari, H. Lee, and A. Faust · 2021
Later among the works it cites.
The Klarna Product Page Dataset: A RealisticBenchmark for Web Representation Learning
A. Hotti, R. S. Risuleo, S. Magureanu, A. Moradi, and J. Lagergren · 2021
Later among the works it cites.
Internet-augmented dialogue generation
M. Komeili, K. Shuster, and J. Weston · 2021
Later among the works it cites.
Towards mental time travel: a hierarchical memory for reinforcement learning agents
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Task-oriented dialogue as dataflow synthesis
J. Andreas, J. Bufe, D. Burkett, C. Chen, J. Clausman, J. Crawford, K. Crim, J. DeLoach, L. Dorner, J. Eisner, et al · 2020
Cited alongside, same era.
Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data
E. M. Bender and A. Koller · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
X. Guo, M. Yu, Y. Gao, C. Gan, M. Campbell, and S. Chang · 2020
Cited alongside, same era.
Interactive fiction games: A colossal adventure
M. Hausknecht, P. Ammanabrolu, M.-A. Côté, and X. Yuan · 2020
Cited alongside, same era.
Multi-agent Communication meets Natural Language: Synergies between Functional and Structural Language Learning
A. Lazaridou, A. Potapenko, and O. Tieleman · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei, et al · 2020
Cited alongside, same era.
A. Lampinen, S. Chan, A. Banino, and F. Hill · 2021
Later among the works it cites.
J. Lin, X. Ma, S.-C. Lin, J.-H. Yang, R. Pradeep, and R. Nogueira · 2021
Later among the works it cites.
WebGPT: Browser-Assisted Question-Answering with Human Feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Later among the works it cites.
AndroidEnv: A Reinforcement Learning Platform for Android
D. Toyama, P. Hamel, A. Gergely, G. Comanici, A. Glaese, Z. Ahmed, T. Jackson, S. Mourad, and D. Precup · 2021
Later among the works it cites.
Survey on reinforcement learning for language processing
V. Uc-Cetina, N. Navarro-Guerrero, A. Martin-Gonzalez, C. Weber, and S. Wermter · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao · 2021
Later among the works it cites.
Reading and acting while blindfolded: The need for semantics in text game agents
S. Yao, K. Narasimhan, and M. Hausknecht · 2021
Later among the works it cites.
Silg: The multi-domain symbolic interactive language grounding benchmark
V. Zhong, A. W. Hanjie, S. Wang, K. Narasimhan, and L. Zettlemoyer · 2021
Later among the works it cites.
Interactive Mobile App Navigation with Uncertain or Under-specified Natural Language Commands
A. Burns, D. Arsan, S. Agrawal, R. Kumar, K. Saenko, and B. A. Plummer · 2022
Closest in time.
A data-driven approach for learning to control computers
P. C. Humphreys, D. Raposo, T. Pohlen, G. Thornton, R. Chhaparia, A. Muldal, J. Abramson, P. Georgiev, A. Goldin, A. Santoro, et al · 2022
Closest in time.
Internet-augmented language models through few-shot prompting for open-domain question answering
A. Lazaridou, E. Gribovskaya, W. Stokowiec, and N. Grigorev · 2022
Closest in time.
K. Shuster, M. Komeili, L. Adolphs, S. Roller, A. D. Szlam, and J. Weston · 2022
Closest in time.
Multi-stage episodic control for strategic exploration in text games
J. Tuyls, S. Yao, S. Kakade, and K. Narasimhan · 2022
Closest in time.
S. Zhuang, H. Ren, L. Shou, J. Pei, M. Gong, G. Zuccon, and D. Jiang · 2022
Closest in time.