Fetching the paper…
Reading the bibliography…
Authorship Analysis, also known as stylometry, has been an essential aspect of Natural Language Processing (NLP) for a long time.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019 · 1908
Earlier work this paper cites.
Accent mobility: A model and some data
Howard Giles. 1973 · 1973
Earlier work this paper cites.
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. 1975 · 1975
Earlier work this paper cites.
On the differences between spoken and written language
F Niyi Akinnaso. 1982 · 1982
Earlier work this paper cites.
Discourse analysis
Gillian Brown, Gillian D Brown, George Yule, Gillian R Brown, and Brown Gillian. 1983 · 1983
Earlier work this paper cites.
Variation across speech and writing
Douglas Biber. 1991 · 1991
Earlier work this paper cites.
Discrimination of authorship using visualization
Bradley Kjell, W Addison Woods, and Ophir Frieder. 1994 · 1994
Earlier work this paper cites.
Longman grammar of spoken and written English
Douglas Biber, Stig Johansson, Geoffrey Leech, Susan Conrad, and Edward Finegan. 2000 · 2000
Earlier work this paper cites.
Dialogue act modeling for automatic tagging and recognition of conversational speech
Andreas Stolcke, Klaus Ries, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates, Daniel Jurafsky, Paul Taylor, Rachel Martin, Carol Van Ess-Dykema, and Marie Meteer. 2000 · 2000
Earlier work this paper cites.
Linguistic inquiry and word count: Liwc 2001
James W Pennebaker, Martha E Francis, and Roger J Booth. 2001 · 2001
Earlier work this paper cites.
The british academic spoken english (base) corpus project
Paul Thompson and Hilary Nesi. 2001 · 2001
Earlier work this paper cites.
The oyez project: Us supreme court multimedia database
Melvin I Urofsky. 2001 · 2001
Earlier work this paper cites.
Using compression-based language models for text categorization
William J Teahan and David J Harper. 2003 · 2003
Earlier work this paper cites.
Who was student and why do we care so much about his t-test? 1
Edward H Livingston. 2004 · 2004
Earlier work this paper cites.
Do writing and speaking employ the same syntactic representations?
Alexandra A Cleland and Martin J Pickering. 2006 · 2006
Earlier work this paper cites.
British national corpus
BNC Consortium et al. 2007 · 2007
Earlier work this paper cites.
How language works: How babies babble, words change meaning, and languages live or die
David Crystal. 2007 · 2007
Earlier work this paper cites.
Author verification by linguistic profiling: An exploration of the parameter space
Hans Van Halteren. 2007 · 2007
Earlier work this paper cites.
Writeprints: A stylometric approach to identity-level identification and similarity detection in cyberspace
Ahmed Abbasi and Hsinchun Chen. 2008 · 2008
Earlier work this paper cites.
Human communication in everyday life: Explanations and applications
Jason S Wrench, James C McCroskey, and Virginia P Richmond. 2008 · 2008
Earlier work this paper cites.
Confessions of a public speaker
Scott Berkun. 2009 · 2009
Earlier work this paper cites.
Characteristics of speaking style and implications for speech recognition
Takahiro Shinozaki, Mari Ostendorf, and Les Atlas. 2009 · 2009
Earlier work this paper cites.
A survey of modern authorship attribution methods
Efstathios Stamatatos. 2009 · 2009
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011 · 2011
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Corpus of Contemporary American English (COCA)
Mark Davies. 2015 · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Which spoken language markers identify deception in high-stakes settings? evidence from earnings conference calls
Judee Burgoon, William J Mayew, Justin Scott Giboney, Aaron C Elkins, Kevin Moffitt, Bradley Dorn, Michael Byrd, and Lee Spitzley. 2016 · 2016
Earlier work this paper cites.
Authorship verification for different languages, genres and topics
Oren Halvani, Christian Winter, and Anika Pflug. 2016 · 2016
Cited alongside, same era.
Tie-breaker: Using language models to quantify gender bias in sports journalism
Fu Liye, C Danescu, and Lillian Lee. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Conversational flow in oxford-style debates
Justine Zhang, Ravi Kumar, Sujith Ravi, and Cristian Danescu-Niculescu-Mizil. 2016 · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
The lj speech dataset
Keith Ito and Linda Johnson. 2017 · 2017
Cited alongside, same era.
Bridging the gap between pre-training and fine-tuning for end-to-end speech translation
Chengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang, and Ming Zhou. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang*, Varsha Kishore*, Felix Wu*, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Gruen for evaluating linguistic quality of generated text
Wanzheng Zhu and Suma Bhat. 2020 · 2020
Later among the works it cites.
All that’s’ human’is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Later among the works it cites.
A survey of speaker recognition: Fundamental theories, recognition methods and opportunities
Muhammad Mohsin Kabir, Muhammad F Mridha, Jungpil Shin, Israt Jahan, and Abu Quwsar Ohi. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The spoken bnc2014: Designing and building a spoken corpus of everyday conversations
Robbie Love, Claire Dembry, Andrew Hardie, Vaclav Brezina, and Tony McEnery. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Voxceleb: a large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman. 2017 · 2017
Cited alongside, same era.
Surveying stylometry techniques and applications
Tempestt Neal, Kalaivani Sundararajan, Aneez Fatima, Yiming Yan, Yingfei Xiang, and Damon Woodard. 2017 · 2017
Cited alongside, same era.
Asking too much? the rhetorical role of questions in political discourse
Justine Zhang, Arthur Spirling, and Cristian Danescu-Niculescu-Mizil. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Katarzyna Kredens, Arja Heini, and Peter Pezik. 2021 · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Later among the works it cites.
VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux. 2021 · 2021
Later among the works it cites.
Feature vector difference based authorship verification for open-world settings
Janith Weerasinghe, Rhia Singh, and Rachel Greenstadt. 2021 · 2021
Later among the works it cites.
Authorship identification using ensemble learning
Ahmed Abbasi, Abdul Rehman Javed, Farkhund Iqbal, Zunera Jalil, Thippa Reddy Gadekallu, and Natalia Kryvinska. 2022 · 2022
Later among the works it cites.
Overview of pan 2022: Authorship verification, profiling irony and stereotype spreaders, style change detection, and trigger detection
Janek Bevendorff, Berta Chulvi, Elisabetta Fersini, Annina Heini, Mike Kestemont, Krzysztof Kredens, Maximilian Mayerl, Reyner Ortega-Bueno, Piotr Pęzik, Martin Potthast, et al. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
On the state of the art in authorship attribution and authorship verification
Jacob Tyo, Bhuwan Dhingra, and Zachary C Lipton. 2022 · 2022
Later among the works it cites.
Real-time end-to-end speech emotion recognition with cross-domain adaptation
Konlakorn Wongpatikaseree, Sattaya Singkul, Narit Hnoohom, and Sumeth Yuenyong. 2022 · 2022
Later among the works it cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al. 2023 · 2023
Closest in time.
Chatalpaca: A multi-turn dialogue corpus based on alpaca instructions
Ning Bian, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, and Ben He. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E Gonzalez, et al. 2023 · 2023
Closest in time.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023 · 2023
Closest in time.
MGTBench: Benchmarking Machine-Generated Text Detection
Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2023 · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023 · 2023
Closest in time.
New ai classifier for indicating ai-written text
J Hendrik Kirchner, L Ahmad, S Aaronson, and J Leike. 2023 · 2023
Closest in time.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023 · 2023
Closest in time.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023 · 2023
Closest in time.
Deepfake text detection: Limitations and opportunities
Jiameng Pu, Zain Sarwar, Sifat Muhammad Abdullah, Abdullah Rehman, Yoonjin Kim, Parantapa Bhattacharya, Mobin Javed, and Bimal Viswanath. 2023 · 2023
Closest in time.
Gptzero
Edward Tian. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Closest in time.
Attribution and obfuscation of neural text authorship: A data mining perspective
Adaku Uchendu, Thai Le, and Dongwon Lee. 2023 · 2023
Closest in time.