Fetching the paper…
Reading the bibliography…
We present ScienceWorld, a benchmark to test agents' scientific reasoning abilities in a new interactive text environment at the level of a standard elementary school science curriculum.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
UNIFIEDQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020 · 1907
Earlier work this paper cites.
Xusen Yin and Jonathan May. 2019 · 1908
Earlier work this paper cites.
Zork: a computerized fantasy simulation game
P David Lebling, Marc S Blank, and Timothy A Anderson. 1979 · 1979
Earlier work this paper cites.
Learning zil
Infocom. 1989 · 1989
Earlier work this paper cites.
Enhancing text-based reinforcement learning agents with commonsense knowledge
Keerthiram Murugesan, Mattia Atzeni, Pushkar Shukla, Mrinmaya Sachan, Pavan Kapanipathi, and Kartik Talamadupula. 2020b · 2005
Earlier work this paper cites.
How to avoid being eaten by a grue: Structured exploration strategies for textual worlds
Prithviraj Ammanabrolu, Ethan Tien, Matthew Hausknecht, and Mark O Riedl. 2020 · 2006
Earlier work this paper cites.
The structure and function of explanations
Tania Lombrozo. 2006 · 2006
Earlier work this paper cites.
Natural language, semantic analysis, and interactive fiction
Graham Nelson. 2006 · 2006
Earlier work this paper cites.
Text-based rl agents with commonsense knowledge: New challenges, environments and baselines
Keerthiram Murugesan, Mattia Atzeni, Pavan Kapanipathi, Pushkar Shukla, Sadhana Kumaravel, Gerald Tesauro, Kartik Talamadupula, Mrinmaya Sachan, and Murray Campbell. 2020a · 2010
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. 2020 · 2010
Earlier work this paper cites.
A study of the knowledge base requirements for passing an elementary science test
Peter Clark, Philip Harrison, and Niranjan Balasubramanian. 2013 · 2013
Earlier work this paper cites.
The z-machine standards document version 1.1
Graham Nelson. 2014 · 2014
Earlier work this paper cites.
Leveraging linguistic structure for open domain information extraction
Gabor Angeli, Melvin Jose Johnson Premkumar, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Ceptre: A language for modeling generative interactive systems
Chris Martens. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning with a combinatorial action space for predicting popular Reddit threads
Ji He, Mari Ostendorf, Xiaodong He, Jianshu Chen, Jianfeng Gao, Lihong Li, and Li Deng. 2016 · 2016
Earlier work this paper cites.
What’s in an explanation? characterizing knowledge and inference requirements for elementary science exams
Peter Jansen, Niranjan Balasubramanian, Mihai Surdeanu, and Peter Clark. 2016 · 2016
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A systematic classification of knowledge, reasoning, and context within the ARC dataset
Michael Boratko, Harshit Padigela, Divyendra Mikkilineni, Pritish Yuvraj, Rajarshi Das, Andrew McCallum, Maria Chang, Achille Fokoue-Nkoutche, Pavan Kapanipathi, Nicholas Mattei, Ryan Musa, Kartik Talamadupula, and Michael Witbrock. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Cited alongside, same era.
Textworld: A learning environment for text-based games
Marc-Alexandre Côté, Ákos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et al. 2018 · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Cited alongside, same era.
WorldTree: A corpus of explanation graphs for elementary science questions supporting multi-hop inference
Peter Jansen, Elizabeth Wainwright, Steven Marmorstein, and Clayton Morrison. 2018 · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021 · 2021
Later among the works it cites.
How to motivate your dragon: Teaching goal-driven agents to speak and act in fantasy worlds
Prithviraj Ammanabrolu, Jack Urbanek, Margaret Li, Arthur Szlam, Tim Rocktäschel, and Jason Weston. 2021 · 2021
Later among the works it cites.
Case-based reasoning for better generalization in text-adventure games
Mattia Atzeni, Shehzaad Dhuliawala, Keerthiram Murugesan, and Mrinmaya Sachan. 2021 · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. 2021 · 2021
Later among the works it cites.
Explaining answers with entailment trees
Bhavana Dalvi, Peter Jansen, Oyvind Tafjord, Zhengnan Xie, Hannah Smith, Leighanna Pipatanangkura, and Peter Clark. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone. 2018 · 2018
Cited alongside, same era.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Graph constrained reinforcement learning for natural language action spaces
Prithviraj Ammanabrolu and Matthew Hausknecht. 2020 · 2020
Cited alongside, same era.
From ‘f’to ‘a’on the ny regents science exams: An overview of the aristo project
Peter Clark, Oren Etzioni, Tushar Khot, Daniel Khashabi, Bhavana Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, Niket Tandon, et al. 2020 · 2020
Cited alongside, same era.
Interactive fiction games: A colossal adventure
Matthew Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre Côté, and Xingdi Yuan. 2020 · 2020
Cited alongside, same era.
R4C: A benchmark for evaluating RC systems to get the right answer for the right reason
Naoya Inoue, Pontus Stenetorp, and Kentaro Inui. 2020 · 2020
Cited alongside, same era.
On the challenges of evaluating compositional explanations in multi-hop inference: Relevance, completeness, and expert ratings
Peter Jansen, Kelly J. Smith, Dan Moreno, and Huitzilin Ortiz. 2021 · 2021
Later among the works it cites.
A systematic survey of text worlds as embodied natural language environments
Peter A Jansen. 2021 · 2021
Later among the works it cites.
QED: A framework and dataset for explanations in question answering
Matthew Lamm, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins. 2021 · 2021
Later among the works it cites.
Inherently explainable reinforcement learning in natural language
Xiangyu Peng, Mark O Riedl, and Prithviraj Ammanabrolu. 2021 · 2021
Later among the works it cites.
General-purpose question-answering with Macaw
Oyvind Tafjord and Peter Clark. 2021 · 2021
Later among the works it cites.
Process-level representation of scientific protocols with interactive annotation
Ronen Tamari, Fan Bai, Alan Ritter, and Gabriel Stanovsky. 2021 · 2021
Later among the works it cites.
Unification-based reconstruction of multi-hop explanations for science questions
Marco Valentino, Mokanarangan Thayaparan, and André Freitas. 2021 · 2021
Later among the works it cites.
Exploiting reasoning chains for multi-hop science question answering
Weiwen Xu, Yang Deng, Huihui Zhang, Deng Cai, and Wai Lam. 2021a · 2021
Later among the works it cites.
Dynamic semantic graph construction and reasoning for explainable multi-hop science question answering
Weiwen Xu, Huihui Zhang, Deng Cai, and Wai Lam. 2021b · 2021
Later among the works it cites.
Reading and acting while blindfolded: The need for semantics in text game agents
Shunyu Yao, Karthik Narasimhan, and Matthew Hausknecht. 2021 · 2021
Later among the works it cites.
Situated dialogue learning through procedural environment generation
Prithviraj Ammanabrolu, Renee Jia, and Mark O Riedl. 2022 · 2022
Closest in time.
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022 · 2022
Closest in time.
Designing effective sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022 · 2022
Closest in time.