Fetching the paper…
Reading the bibliography…
Language modeling studies the probability distributions over strings of texts.
Note on the general case of the bayes-laplace formula for inductive or a posteriori probabilities
George James Lidstone · 1920
Earlier work this paper cites.
Probability: The deductive and inductive problems
William Ernest Johnson · 1932
Earlier work this paper cites.
Generalized iterative scaling for log-linear models
John N Darroch and Douglas Ratcliff · 1972
Earlier work this paper cites.
Design of a linguistic statistical decoder for the recognition of continuous speech
Frederick Jelinek, Lalit Bahl, and Robert Mercer · 1975
Earlier work this paper cites.
Continuous speech recognition by statistical methods
Frederick Jelinek · 1976
Earlier work this paper cites.
Interpolated estimation of markov source parameters from sparse data
Frederick Jelinek · 1980
Earlier work this paper cites.
A maximum likelihood approach to continuous speech recognition
Lalit R Bahl, Frederick Jelinek, and Robert L Mercer · 1983
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Slava Katz · 1987
Earlier work this paper cites.
A statistical approach to machine translation
Peter F Brown, John Cocke, Stephen A Della Pietra, Vincent J Della Pietra, Frederick Jelinek, John Lafferty, Robert L Mercer, and Paul S Roossin · 1990
Earlier work this paper cites.
Self-organized language modeling for speech recognition
Fred Jelinek, B Merialdo, S Roukos, M Strauss, et al · 1990
Earlier work this paper cites.
A comparison of the enhanced good-turing and deleted estimation methods for estimating probabilities of english bigrams
Kenneth W Church and William A Gale · 1991
Earlier work this paper cites.
Class-based n-gram models of natural language
Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer · 1992
Earlier work this paper cites.
Adaptive language modeling using minimum discriminant estimation
Stephen A Della Pietra, Vincent J Della Pietra, Robert L Mercer, and Salim Roukos · 1992
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Towards better language models for spontaneous speech
Bernhard Suhm and Alex Waibel · 1994
Earlier work this paper cites.
Ergodic hidden markov models and polygrams for language modeling
Thomas Kuhn, Heinrich Niemann, and Ernst Günter Schukat-Talamazzini · 1994
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Reinhard Kneser and Hermann Ney · 1995
Earlier work this paper cites.
Bayesian estimation methods for n-gram language model adaptation
Marcello Federico · 1996
Earlier work this paper cites.
A variable-length category-based n-gram language model
Thomas R Niesler and Philip C Woodland · 1996
Earlier work this paper cites.
A maximum entropy approach to natural language processing
Adam Berger, Stephen A Della Pietra, and Vincent J Della Pietra · 1996
Earlier work this paper cites.
A maximum entropy approach to adaptive statistical language modeling
Roni Rosenfeld · 1996
Earlier work this paper cites.
Class phrase models for language modeling
Klaus Ries, Finn Dag Buo, and Alex Waibel · 1996
Earlier work this paper cites.
A whole sentence maximum entropy language model
Ronald Rosenfeld · 1997
Earlier work this paper cites.
Exploiting syntactic structure for language modeling
Ciprian Chelba and Frederick Jelinek · 1998
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Data-driven determination of appropriate dictionary units for korean lvcsr
Daniel Kiecza, Tanja Schultz, and Alex Waibel · 1999
Earlier work this paper cites.
Efficient sampling and feature selection in whole sentence maximum entropy language models
Stanley F Chen and Ronald Rosenfeld · 1999
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman · 1999
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Structured language modeling
Ciprian Chelba and Frederick Jelinek · 2000
Earlier work this paper cites.
Variable n-grams and extensions for conversational speech language modeling
Manhung Siu and Mari Ostendorf · 2000
Earlier work this paper cites.
An efficient a* search algorithm for statistical machine translation
Franz Josef Och, Nicola Ueffing, and Hermann Ney · 2001
Earlier work this paper cites.
Data-driven approach to designing compound words for continuous speech recognition
George Saon and Mukund Padmanabhan · 2001
Earlier work this paper cites.
Whole-sentence exponential language models: a vehicle for linguistic-statistical integration
Ronald Rosenfeld, Stanley F Chen, and Xiaojin Zhu · 2001
Earlier work this paper cites.
A decoder for syntax-based statistical mt
Kenji Yamada and Kevin Knight · 2002
Earlier work this paper cites.
Unsupervised morpheme segmentation and morphology induction from text corpora using Morfessor 1.0
Mathias Creutz and Krista Lagus · 2005
Earlier work this paper cites.
Training neural network language models on very large corpora
Holger Schwenk and Jean-Luc Gauvain · 2005
Earlier work this paper cites.
Morph-based speech recognition and modeling of out-of-vocabulary words across languages
Mathias Creutz, Teemu Hirsimäki, Mikko Kurimo, Antti Puurula, Janne Pylkkönen, Vesa Siivola, Matti Varjokallio, Ebru Arisoy, Murat Saraçlar, and Andreas Stolcke · 2007
Earlier work this paper cites.
Large vocabulary continuous speech recognition of an inflected language using stems and endings
Tomaž Rotovnik, Mirjam Sepesy Maučec, and Zdravko Kačič · 2007
Earlier work this paper cites.
Joint morphological-lexical language modeling (jmllm) for arabic
Ruhi Sarikaya, Mohamed Afify, and Yuqing Gao · 2007
Earlier work this paper cites.
Gaussian mixture language models for speech recognition
Mohamed Afify, Olivier Siohan, and Ruhi Sarikaya · 2007
Earlier work this paper cites.
Continuous space language models
Holger Schwenk · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Morphology-based and sub-word language modeling for turkish speech recognition
Haşim Sak, Murat Saraclar, and Tunga Güngör · 2010
Earlier work this paper cites.
Uyghur morpheme-based language models and asr
Mijit Ablimit, Graham Neubig, Masato Mimura, Shinsuke Mori, Tatsuya Kawahara, and Askar Hamdulla · 2010
Earlier work this paper cites.
Generating text with recurrent neural networks
Ilya Sutskever, James Martens, and Geoffrey E Hinton · 2011
Earlier work this paper cites.
Mandarin word-character hybrid-input neural network language model
Moonyoung Kang, Tim Ng, and Long Nguyen · 2011
Earlier work this paper cites.
Extensions of recurrent neural network language model
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2011
Earlier work this paper cites.
Recurrent neural network based language modeling in meeting recognition
Stefan Kombrink, Tomas Mikolov, Martin Karafiát, and Lukás Burget · 2011
Earlier work this paper cites.
Subword language modeling with neural networks
Tomáš Mikolov, Ilya Sutskever, Anoop Deoras, Hai-Son Le, Stefan Kombrink, and Jan Cernocky · 2012
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
Deep neural network language models
Ebru Arisoy, Tara N Sainath, Brian Kingsbury, and Bhuvana Ramabhadran · 2012
Earlier work this paper cites.
Lstm neural networks for language modeling
Martin Sundermeyer, Ralf Schlüter, and Hermann Ney · 2012
Earlier work this paper cites.
Dependency language models for sentence completion
Joseph Gubbins and Andreas Vlachos · 2013
Earlier work this paper cites.
Converting neural network language models into back-off language models for efficient decoding in automatic speech recognition
Ebru Arısoy, Stanley F Chen, Bhuvana Ramabhadran, and Abhinav Sethy · 2013
Earlier work this paper cites.
Word-phrase-entity language models: Getting more mileage out of n-grams
Michael Levit, Sarangarajan Parthasarathy, Shuangyu Chang, Andreas Stolcke, and Benoit Dumoulin · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Cache based recurrent neural network language model inference for first pass speech recognition
Zhiheng Huang, Geoffrey Zweig, and Benoit Dumoulin · 2014
Earlier work this paper cites.
Dependency recurrent neural language models for sentence completion
Piotr Mirowski and Andreas Vlachos · 2015
Earlier work this paper cites.
Bidirectional recurrent neural network language models for automatic speech recognition
Ebru Arisoy, Abhinav Sethy, Bhuvana Ramabhadran, and Stanley Chen · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2015
Earlier work this paper cites.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M Rush · 2016
Earlier work this paper cites.
Gated word-character recurrent language model
Yasumasa Miyamoto and Kyunghyun Cho · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
A simple, fast diverse decoding algorithm for neural generation
Jiwei Li, Will Monroe, and Dan Jurafsky · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Character-level language modeling with hierarchical recurrent neural networks
Kyuyeon Hwang and Wonyong Sung · 2017
Earlier work this paper cites.
Character-word lstm language models
Lyan Verwimp, Joris Pelemans, Patrick Wambacq, et al · 2017
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learning discourse-level diversity for neural dialog models using conditional variational autoencoders
Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi · 2017
Earlier work this paper cites.
Neural response generation via gan with an approximate embedding layer
Zhen Xu, Bingquan Liu, Baoxun Wang, Cheng-Jie Sun, Xiaolong Wang, Zhuoran Wang, and Chao Qi · 2017
Earlier work this paper cites.
Trainable greedy decoding for neural machine translation
Jiatao Gu, Kyunghyun Cho, and Victor OK Li · 2017
Earlier work this paper cites.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2017
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Cited alongside, same era.
Automatic detection of machine generated text: A critical survey
Ganesh Jawahar, Muhammad Abdul-Mageed, and VS Laks Lakshmanan · 2020
Later among the works it cites.
The cost of training nlp models: A concise overview
Or Sharir, Barak Peleg, and Yoav Shoham · 2020
Later among the works it cites.
Pre-training transformers as energy-based cloze models
Kevin Clark, Minh-Thang Luong, Quoc Le, and Christopher D Manning · 2020
Later among the works it cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Later among the works it cites.
Structured pruning of large language models
Ziheng Wang, Jeremy Wohlwend, and Tao Lei · 2020
Later among the works it cites.
Tinybert: Distilling bert for natural language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Cited alongside, same era.
Diverse beam search for improved description of complex scenes
Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra · 2018
Cited alongside, same era.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Cited alongside, same era.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhijeng Chen, and Rohit Prabhavalkar · 2018
Cited alongside, same era.
A comparison of techniques for language model integration in encoder-decoder speech recognition
Shubham Toshniwal, Anjuli Kannan, Chung-Cheng Chiu, Yonghui Wu, Tara N Sainath, and Karen Livescu · 2018
Cited alongside, same era.
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2020
Later among the works it cites.
Leveraging pre-trained checkpoints for sequence generation tasks
Sascha Rothe, Shashi Narayan, and Aliaksei Severyn · 2020
Later among the works it cites.
Knowledge graph-augmented abstractive summarization with semantic-driven cloze reward
Luyang Huang, Lingfei Wu, and Lu Wang · 2020
Later among the works it cites.
How context affects language models’ factual predictions
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel · 2020
Later among the works it cites.
Green ai
Roy Schwartz, Jesse Dodge, Noah A Smith, and Oren Etzioni · 2020
Later among the works it cites.
Train big, then compress: Rethinking model size for efficient training and inference of transformers
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joey Gonzalez · 2020
Later among the works it cites.
E-bert: A phrase and product knowledge enhanced language model for e-commerce
Denghui Zhang, Zixuan Yuan, Yanchi Liu, Fuzhen Zhuang, Haifeng Chen, and Hui Xiong · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Later among the works it cites.
Pre-trained models: Past, present and future
Xu Han, Zhengyan Zhang, Ning Ding, Yuxian Gu, Xiao Liu, Yuqi Huo, Jiezhong Qiu, Yuan Yao, Ao Zhang, Liang Zhang, et al · 2021
Later among the works it cites.
Adapterfusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych · 2021
Later among the works it cites.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze · 2021
Later among the works it cites.
Few-shot text generation with natural language instructions
Timo Schick and Hinrich Schütze · 2021
Later among the works it cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2021
Later among the works it cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Later among the works it cites.
Factual probing is [mask]: Learning vs. learning to recall
Zexuan Zhong, Dan Friedman, and Danqi Chen · 2021
Later among the works it cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Later among the works it cites.
Learning how to ask: Querying lms with mixtures of soft prompts
Guanghui Qin and Jason Eisner · 2021
Later among the works it cites.
Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning
Colin Wei, Sang Michael Xie, and Tengyu Ma · 2021
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Later among the works it cites.
Persistent anti-muslim bias in large language models
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Later among the works it cites.
Robustness gym: Unifying the nlp evaluation landscape
Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré · 2021
Later among the works it cites.
Textflint: Unified multilingual robustness evaluation toolkit for natural language processing
Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, et al · 2021
Later among the works it cites.
Tweepfake: About detecting deepfake tweets
Tiziano Fagni, Fabrizio Falchi, Margherita Gambini, Antonio Martella, and Maurizio Tesconi · 2021
Later among the works it cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, et al · 2021
Later among the works it cites.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Xiaocheng Li, and Samuel Bowman · 2021
Later among the works it cites.
Parameter-efficient transfer learning with diff pruning
Demi Guo, Alexander M Rush, and Yoon Kim · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Later among the works it cites.
Non-autoregressive text generation with pre-trained language models
Yixuan Su, Deng Cai, Yan Wang, David Vandyke, Simon Baker, Piji Li, and Nigel Collier · 2021
Later among the works it cites.
Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training
Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou · 2021
Later among the works it cites.
KLMo: Knowledge graph enhanced pretrained language model with fine-grained relationships
Lei He, Suncong Zheng, Tao Yang, and Feng Zhang · 2021
Later among the works it cites.
KEPLER: A unified model for knowledge embedding and pre-trained language representation
Xiaozhi Wang, Tianyu Gao, Zhaocheng Zhu, Zhengyan Zhang, Zhiyuan Liu, Juanzi Li, and Jian Tang · 2021
Later among the works it cites.
QA-GNN: Reasoning with language models and knowledge graphs for question answering
Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec · 2021
Later among the works it cites.
Mind the gap: Assessing temporal generalization in neural language models
Angeliki Lazaridou, Adhi Kuncoro, Elena Gribovskaya, Devang Agrawal, Adam Liska, Tayfun Terzi, Mai Gimenez, Cyprien de Masson d’Autume, Tomas Kocisky, Sebastian Ruder, et al · 2021
Later among the works it cites.
Inductive learning on commonsense knowledge graph completion
Bin Wang, Guangtao Wang, Jing Huang, Jiaxuan You, Jure Leskovec, and C-C Jay Kuo · 2021
Later among the works it cites.
A survey on green deep learning
Jingjing Xu, Wangchunshu Zhou, Zhiyi Fu, Hao Zhou, and Lei Li · 2021
Later among the works it cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon · 2021
Later among the works it cites.
Relational sentence embedding for flexible semantic matching
Bin Wang and Haizhou Li · 2022
Later among the works it cites.
Task-specific dependency-based word embedding methods
Chengwei Wei, Bin Wang, and C.-C. Jay Kuo · 2022
Later among the works it cites.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2022
Later among the works it cites.
Sustainable ai: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al · 2022
Later among the works it cites.
Robust natural language processing: Recent advances, challenges, and future directions
Marwan Omar, Soohyeon Choi, DaeHun Nyang, and David Mohaisen · 2022
Later among the works it cites.
On the explainability of natural language processing deep models
Julia El Zini and Mariette Awad · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Later among the works it cites.
Recent advances in deep learning based dialogue systems: A systematic survey
Jinjie Ni, Tom Young, Vlad Pandelea, Fuzhao Xue, and Erik Cambria · 2022
Later among the works it cites.
A focused study on sequence length for dialogue summarization
Bin Wang, Chen Zhang, Chengwei Wei, and Haizhou Li · 2022
Later among the works it cites.
Analyzing and evaluating faithfulness in dialogue summarization
Bin Wang, Chen Zhang, Yan Zhang, Yiming Chen, and Haizhou Li · 2022
Later among the works it cites.
Detecting computer-generated disinformation
Harald Stiff and Fredrik Johansson · 2022
Later among the works it cites.
Elmer: A non-autoregressive pre-trained language model for efficient and effective text generation
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2022
Later among the works it cites.
Kgboost: A classification-based knowledge base completion method with negative sampling
Yun-Cheng Wang, Xiou Ge, Bin Wang, and C-C Jay Kuo · 2022
Later among the works it cites.
Compounde: Knowledge graph embedding with translation, rotation and scaling compound operations
Xiou Ge, Yun-Cheng Wang, Bin Wang, and C-C Jay Kuo · 2022
Later among the works it cites.
Deep bidirectional language-knowledge graph pretraining
Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christopher D Manning, Percy S Liang, and Jure Leskovec · 2022
Later among the works it cites.
GreaseLM: Graph REASoning enhanced language models
Xikun Zhang, Antoine Bosselut, Michihiro Yasunaga, Hongyu Ren, Percy Liang, Christopher D Manning, and Jure Leskovec · 2022
Later among the works it cites.
Empowering language models with knowledge graph reasoning for open-domain question answering
Ziniu Hu, Yichong Xu, Wenhao Yu, Shuohang Wang, Ziyi Yang, Chenguang Zhu, Kai-Wei Chang, and Yizhou Sun · 2022
Later among the works it cites.
Time waits for no one! analysis and challenges of temporal misalignment
Kelvin Luu, Daniel Khashabi, Suchin Gururangan, Karishma Mandyam, and Noah A. Smith · 2022
Later among the works it cites.
Lifelong pretraining: Continually adapting language models to emerging corpora
Xisen Jin, Dejiao Zhang, Henghui Zhu, Wei Xiao, Shang-Wen Li, Xiaokai Wei, Andrew Arnold, and Xiang Ren · 2022
Later among the works it cites.
Green learning: Introduction, examples and outlook
C.-C. Jay Kuo and Azad M Madni · 2022
Later among the works it cites.
GreenKGC: A lightweight knowledge graph completion method
Yun-Cheng Wang, Xiou Ge, Bin Wang, and C-C Jay Kuo · 2022
Later among the works it cites.
A domain knowledge enhanced pre-trained language model for vertical search: Case study on medicinal products
Kesong Liu, Jianhui Jiang, and Feifei Lyu · 2022
Later among the works it cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu · 2022
Later among the works it cites.
Knowledge graph contrastive learning for recommendation
Yuhao Yang, Chao Huang, Lianghao Xia, and Chenliang Li · 2022
Later among the works it cites.
Synwmd: Syntax-aware word mover’s distance for sentence similarity evaluation
Chengwei Wei, Bin Wang, and C-C Jay Kuo · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Closest in time.
Automatic prompt optimization with" gradient descent" and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng · 2023
Closest in time.
An overview on generative ai at scale with edge-cloud computing
Yun-Cheng Wang, Jintang Xue, Chengwei Wei, and C-C Jay Kuo · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.