Fetching the paper…
Reading the bibliography…
The uniform information density (UID) hypothesis, which posits that speakers behaving optimally tend to distribute information uniformly across a linguistic signal, has gained traction in psycholinguistics as an explanation for certain syntactic, morphological, and prosodic choices.
A mathematical theory of communication
Claude E. Shannon. 1948 · 1948
Earlier work this paper cites.
Information Theory and Statistical Mechanics
Edwin T. Jaynes. 1957 · 1957
Earlier work this paper cites.
Konstanz im kurzzeitgedächtnis-konstanz im sprachlichen informationsfluß
August Fenk and Gertraud Fenk. 1980 · 1980
Earlier work this paper cites.
An Introduction to the Bootstrap
Bradley Efron and Robert J. Tibshirani. 1994 · 1994
Earlier work this paper cites.
The dependency locality theory: A distance-based theory of linguistic complexity
Edward Gibson. 2000 · 2000
Earlier work this paper cites.
A probabilistic Earley parser as a psycholinguistic model
John Hale. 2001 · 2001
Earlier work this paper cites.
Probabilistic relations between words: Evidence from reduction in lexical production
Daniel Jurafsky, Alan Bell, Michelle Gregory, and William D. Raymond. 2001 · 2001
Earlier work this paper cites.
Effects of disfluencies, predictability, and utterance position on word form variation in English conversation
Alan Bell, Daniel Jurafsky, Eric Fosler-Lussier, Cynthia Girand, Michelle Gregory, and Daniel Gildea. 2003 · 2003
Earlier work this paper cites.
The smooth signal redundancy hypothesis: A functional explanation for relationships between redundancy, prosodic prominence, and duration in spontaneous speech
Matthew Aylett and Alice Turk. 2004 · 2004
Earlier work this paper cites.
Knowledge of grammar, knowledge of usage: Syntactic probabilities affect pronunciation variation
Susanne Gahl and Susan M. Garnsey. 2004 · 2004
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Lexical frequency and acoustic reduction in spoken dutch
Mark Pluymaekers, Mirjam Ernestus, and R. Harald Baayen. 2005 · 2005
Earlier work this paper cites.
Speakers optimize information density through syntactic reduction
Roger P. Levy and T. F. Jaeger. 2007 · 2007
Earlier work this paper cites.
Speaking rationally: Uniform information density as an optimal strategy for language production
Austin F. Frank and T. Florian Jaeger. 2008 · 2008
Earlier work this paper cites.
Predictability effects on durations of content and function words in conversational English
Alan Bell, Jason M. Brenier, Michelle Gregory, Cynthia Girand, and Dan Jurafsky. 2009 · 2009
Earlier work this paper cites.
Redundancy and reduction: Speakers manage syntactic information density
T. Florian Jaeger. 2010 · 2010
Cited alongside, same era.
On language ‘utility’: Processing complexity and communicative efficiency
T. Florian Jaeger and Harry Tily. 2011 · 2011
Cited alongside, same era.
Syntactic surprisal affects spoken word duration in conversational contexts
Vera Demberg, Asad Sayeed, Philip Gorinski, and Nikolaos Engonopoulos. 2012 · 2012
Cited alongside, same era.
The effects of construction probability on word durations during spontaneous incremental sentence production
Victor Kuperman and Joan Bresnan. 2012 · 2012
Cited alongside, same era.
Syntactic predictability influences duration
Claire Moore-Cantwell. 2013 · 2013
Cited alongside, same era.
Information density and dependency length as complementary cognitive models
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Later among the works it cites.
The role of predictability in shaping phonological patterns
Kathleen Currie Hall, Elizabeth Hume, T. Florian Jaeger, and Andrew Wedel. 2018 · 2018
Later among the works it cites.
Uniform Information Density effects on syntactic choice in Hindi
Ayush Jain, Vishal Singh, Sidharth Ranjan, Rajakrishnan Rajkumar, and Sumeet Agarwal. 2018 · 2018
Later among the works it cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Later among the works it cites.
Importance of search and evaluation strategies in neural dialogue modeling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Xavier Collins. 2014 · 2014
Cited alongside, same era.
Word informativity influences acoustic duration: Effects of contextual predictability on lexical representation
Scott Seyfarth. 2014 · 2014
Cited alongside, same era.
Testing the processing hypothesis of word order variation using a probabilistic language model
Jelke Bloem. 2016 · 2016
Cited alongside, same era.
The length of words reflects their conceptual complexity
Molly L. Lewis and Michael C. Frank. 2016 · 2016
Cited alongside, same era.
Information density and quality estimation features as translationese indicators for human translation classification
Raphael Rubino, Ekaterina Lapshinova-Koltunski, and Josef van Genabith. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019 · 2019
Later among the works it cites.
What kind of language is hard to language-model?
Sabrina J. Mielke, Ryan Cotterell, Kyle Gorman, Brian Roark, and Jason Eisner. 2019 · 2019
Later among the works it cites.
When does label smoothing help?
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Towards a better understanding of label smoothing in neural machine translation
Yingbo Gao, Weiyue Wang, Christian Herold, Zijian Yang, and Hermann Ney. 2020 · 2020
Later among the works it cites.
Wiki-40B: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou. 2020 · 2020
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Later among the works it cites.
If beam search is the answer, what was the question?
Clara Meister, Ryan Cotterell, and Tim Vieira. 2020a · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Closest in time.