Fetching the paper…
Reading the bibliography…
The drastic increase in language models' parameters has led to a new trend of deploying models in cloud servers, raising growing concerns about private inference for Transformer-based models.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N.; Hinton, G. E.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
Automatic Evaluation of Machine Translation Quality Using Longest Common Subsequence and Skip-Bigram Statistics
Lin, C.; and Och, F. J. 2004 · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Earlier work this paper cites.
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
Li, Y.; Su, H.; Shen, X.; Li, W.; Cao, Z.; and Niu, S. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Popovic, M. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Deep Learning using Rectified Linear Units (ReLU)
Agarap, A. F. 2018 · 2018
Cited alongside, same era.
Findings of the E2E NLG Challenge
Dusek, O.; Novikova, J.; and Rieser, V. 2018 · 2018
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Cited alongside, same era.
MultiWOZ 2.1: A Consolidated Multi-Domain Dialogue Dataset with State Corrections and State Tracking Baselines
Eric, M.; Goel, R.; Paul, S.; Sethi, A.; Agarwal, S.; Gao, S.; Kumar, A.; Goyal, A. K.; Ku, P.; and Hakkani-Tür, D. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020 · 2020
Later among the works it cites.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2020
Later among the works it cites.
CrypTen: Secure Multi-Party Computation Meets Machine Learning
Knott, B.; Venkataraman, S.; Hannun, A. Y.; Sengupta, S.; Ibrahim, M.; and van der Maaten, L. 2021 · 2021
Later among the works it cites.
BARTScore: Evaluating Generated Text as Text Generation
Yuan, W.; Neubig, G.; and Liu, P. 2021 · 2021
Later among the works it cites.
THE-X: Privacy-Preserving Transformer Inference with Homomorphic Encryption
Chen, T.; Bao, H.; Huang, S.; Dong, L.; Jiao, B.; Jiang, D.; Zhou, H.; Li, J.; and Wei, F. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; and Zettlemoyer, L. 2020 · 2020
Cited alongside, same era.
CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning
Lin, B. Y.; Zhou, W.; Shen, M.; Zhou, P.; Bhagavatula, C.; Choi, Y.; and Ren, X. 2020 · 2020
Cited alongside, same era.
Delphi: A Cryptographic Inference Service for Neural Networks
Mishra, P.; Lehmkuhl, R.; Srinivasan, A.; Zheng, W.; and Popa, R. A. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
Iron: Private Inference on Transformers
Hao, M.; Li, H.; Chen, H.; Xing, P.; Xu, G.; and Zhang, T. 2022 · 2022
Later among the works it cites.
How Much Does Attention Actually Attend? Questioning the Importance of Attention in Pretrained Transformers
Hassid, M.; Peng, H.; Rotem, D.; Kasai, J.; Montero, I.; Smith, N. A.; and Schwartz, R. 2022 · 2022
Later among the works it cites.
MPCFormer: fast, performant and private Transformer inference with MPC
Li, D.; Shao, R.; Wang, H.; Guo, H.; Xing, E. P.; and Zhang, H. 2022 · 2022
Later among the works it cites.