Fetching the paper…
Reading the bibliography…
In recent years, large pre-trained language models (LLMs) have demonstrated the ability to follow instructions and perform novel tasks from a few examples.
Speech understanding systems: Report of a steering committee
Medress, M. F., Cooper, F. S., Forgie, J. W., Green, C., Klatt, D. H., O’Malley, M. H., Neuburg, E. P., Newell, A., Reddy, D., Ritea, B., et al · 1977
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Lstm recurrent networks learn simple context-free and context-sensitive languages
Gers, F. A. and Schmidhuber, E · 2001
Earlier work this paper cites.
Optimal ordered problem solver
Schmidhuber, J · 2004
Earlier work this paper cites.
Gödel machines: Fully self-referential optimal universal self-improvers
Schmidhuber, J · 2007
Earlier work this paper cites.
Schmidhuber, J · 2015
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
URL https://openai.com/blog/openai-api/
Brockman, G., Murati, M., and Welinder, P., Sept 2018 · 2018
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B. et al · 2020
Earlier work this paper cites.
On measuring and mitigating biased inferences of word embeddings
Dev, S., Li, T., Phillips, J. M., and Srikumar, V · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Unsupervised question decomposition for question answering
Perez, E., Lewis, P., Yih, W.-t., Cho, K., and Kiela, D · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Csordás, R., Irie, K., and Schmidhuber, J · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Addressing some limitations of transformers with feedback memory, 2021
Fan, A., Lavril, T., Grave, E., Joulin, A., and Sukhbaatar, S · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J · 2021
Earlier work this paper cites.
Which linguist invented the lightbulb? presupposition verification for question-answering
Kim, N., Pavlick, E., Ayan, B. K., and Ramachandran, D · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Prompting contrastive explanations for commonsense reasoning tasks
Paranjape, B., Michael, J., Ghazvininejad, M., Hajishirzi, H., and Zettlemoyer, L · 2021
Earlier work this paper cites.
Societal biases in language generation: Progress and challenges
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N · 2021
Earlier work this paper cites.
Beyond goldfish memory: Long-term open-domain conversation
Xu, J., Szlam, A. D., and Weston, J · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al · 2022
Cited alongside, same era.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A. J., Lewkowycz, A., Misra, V., Ramasesh, V. V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Faithful reasoning using large language models
Creswell, A. and Shanahan, M · 2022
Cited alongside, same era.
Towards teachable reasoning systems
Dalvi, B., Tafjord, O., and Clark, P · 2022
Talking about large language models
Shanahan, M · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Later among the works it cites.
Galactica: A large language model for science
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J., et al · 2022
Cited alongside, same era.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2022
Cited alongside, same era.
Improving zero and few-shot generalization in dialogue through instruction tuning
Gupta, P., Jiao, C., Yeh, Y.-T., Mehri, S., Eskenazi, M., and Bigham, J. P · 2022
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., Chebotar, Y., Sermanet, P., Jackson, T., Brown, N., Luu, L., Levine, S., Hausman, K., and brian ichter · 2022
Cited alongside, same era.
Block-recurrent transformers
Hutchins, D., Schlag, I., Wu, Y., Dyer, E., and Neyshabur, B · 2022
Cited alongside, same era.
Karpas, E., Abend, O., Belinkov, Y., Lenz, B., Lieber, O., Ratner, N., Shoham, Y., Bata, H., Levine, Y., Leyton-Brown, K., et al · 2022
Cited alongside, same era.
Large language models still can’t plan (a benchmark for LLMs on planning and reasoning about change)
Valmeekam, K., Olmo, A., Sreedharan, S., and Kambhampati, S · 2022
Later among the works it cites.
Super-NaturalInstructions: Generalization via declarative instructions on 1600+ NLP tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H., Purohit, I., Mondal, I., Anderson, J., Kuznia, K., Doshi, K., Pal, K. K., Patel, M., Moradshahi, M., Parmar, M., Purohit, M., Varshney, N., Kaza, P. R., Verma, P., Puri, R. S., Karia, R., Doshi, S., Sampat, S. K., Mishra, S., Reddy A, S., Patro, S., Dixit, T., and Shen, X · 2022
Later among the works it cites.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts
Wu, T., Terry, M., and Cai, C. J · 2022
Later among the works it cites.
SEQZERO: Few-shot compositional semantic parsing with sequential prompts and zero-shot models
Yang, J., Jiang, H., Yin, Q., Zhang, D., Yin, B., and Yang, D · 2022
Later among the works it cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Yao, S., Chen, H., Yang, J., and Narasimhan, K. R · 2022
Later among the works it cites.
STar: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Langchain
Chase, H · 2023
Closest in time.
Crawling the internal knowledge-base of language models
Cohen, R., Geva, M., Berant, J., and Globerson, A · 2023
Closest in time.
Neural networks and the chomsky hierarchy
Deletang, G., Ruoss, A., Grau-Moya, J., Genewein, T., Wenliang, L. K., Catt, E., Cundy, C., Hutter, M., Legg, S., Veness, J., and Ortega, P. A · 2023
Closest in time.
Looped transformers as programmable computers
Giannou, A., Rajput, S., Sohn, J.-y., Lee, K., Lee, J. D., and Papailiopoulos, D · 2023
Closest in time.
Decomposed prompting: A modular approach for solving complex tasks
Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A · 2023
Closest in time.
Internet-augmented language models through few-shot prompting for open-domain question answering, 2023
Lazaridou, A., Gribovskaya, E., Stokowiec, W. J., and Grigorev, N · 2023
Closest in time.
Transformers learn shortcuts to automata
Liu, B., Ash, J. T., Goel, S., Krishnamurthy, A., and Zhang, C · 2023
Closest in time.
Iterated decomposition: Improving science q&a by supervising reasoning processes
Reppert, J., Rachbach, B., George, C., Byun, L. S. J., Appleton, M., and Stuhlmüller, A · 2023
Closest in time.
Language models can (kind of) reason: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H · 2023
Closest in time.
Memory augmented large language models are computationally universal
Schuurmans, D · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A., Attarian, M., brian ichter, Choromanski, K. M., Wong, A., Welker, S., Tombari, F., Purohit, A., Ryoo, M. S., Sindhwani, V., Lee, J., Vanhoucke, V., and Florence, P · 2023
Closest in time.