Fetching the paper…
Reading the bibliography…
Language is typically modelled with discrete sequences.
The origin of speech
Charles F Hockett and Charles D Hockett · 1960
Earlier work this paper cites.
Language and nature
Noam Chomsky · 1995
Earlier work this paper cites.
Long short-term memory
S Hochreiter · 1997
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher D Manning · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
How did language go discrete
Michael Studdert-Kennedy · 2005
Earlier work this paper cites.
Continuous space language models
Holger Schwenk · 2006
Earlier work this paper cites.
A scalable hierarchical distributed language model
Andriy Mnih and Geoffrey E Hinton · 2008
Earlier work this paper cites.
Lstm neural networks for language modeling
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney · 2012
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
Vector space models for phrase-based machine translation
Tamer Alkhouli, Andreas Guta, and Hermann Ney · 2014
Earlier work this paper cites.
Word representations via gaussian embedding
Luke Vilnis and Andrew McCallum · 2014
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew M Dai, Rafal Jozefowicz, and Samy Bengio · 2015
Earlier work this paper cites.
Linguistics: An introduction to language and communication
Adrian Akmajian, Ann K Farmer, Lee Bickmore, Richard A Demers, and Robert M Harnish · 2017
Earlier work this paper cites.
Continuous multilinguality with language vectors
Robert Östling and Jörg Tiedemann · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Neural ordinary differential equations, 2019
Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David Duvenaud · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Aaron Baier-Reinio and Hans De Sterck · 2020
Cited alongside, same era.
Continuous spatiotemporal transformers, 2023
Antonio Fonseca, Emanuele Zappala, Josue Ortega Caro, and David van Dijk · 2023
Later among the works it cites.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Later among the works it cites.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, et al · 2020
Cited alongside, same era.
Neural controlled differential equations for irregular time series
Patrick Kidger, James Morrill, James Foster, and Terry Lyons · 2020
Cited alongside, same era.
Redesigning the transformer architecture with insights from multi-particle dynamical systems
Subhabrata Dutta, Tanya Gautam, Soumen Chakrabarti, and Tanmoy Chakraborty · 2021
Cited alongside, same era.
Neural rough differential equations for long time series
James Morrill, Cristopher Salvi, Patrick Kidger, and James Foster · 2021
Cited alongside, same era.
Continuous language generative flow
Zineng Tang, Shiyue Zhang, Hyounghun Kim, and Mohit Bansal · 2021
Cited alongside, same era.
Continuous self-attention models with neural ode networks
Jing Zhang, Peng Zhang, Baiwen Kong, Junqiu Wei, and Xin Jiang · 2021
Cited alongside, same era.
Diffuseq: Sequence to sequence text generation with diffusion models
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and LingPeng Kong · 2022
Cited alongside, same era.
Eric Todd, Millicent L Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Anthropic · 2024
Later among the works it cites.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda · 2024
Later among the works it cites.
Contiformer: Continuous-time transformer for irregular time series modeling
Yuqi Chen, Kan Ren, Yansen Wang, Yuchen Fang, Weiwei Sun, and Dongsheng Li · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Likelihood-based diffusion language models
Ishaan Gulrajani and Tatsunori B Hashimoto · 2024
Later among the works it cites.
Graph-enhanced large language models in asynchronous plan reasoning
Fangru Lin, Emanuele La Malfa, Valentin Hofmann, Elle Michelle Yang, Anthony Cohn, and Janet B Pierrehumbert · 2024
Later among the works it cites.
Rough transformers: Lightweight continuous-time sequence modelling with path signatures, 2024
Fernando Moreno-Pino, Álvaro Arroyo, Harrison Waldon, Xiaowen Dong, and Álvaro Cartea · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Planner: generating diversified paragraph via latent language diffusion model
Yizhe Zhang, Jiatao Gu, Zhuofeng Wu, Shuangfei Zhai, Joshua Susskind, and Navdeep Jaitly · 2024
Later among the works it cites.