Fetching the paper…
Reading the bibliography…
Large transformer models can highly improve Answer Sentence Selection (AS2) tasks, but their high computational costs prevent their use in many real-world applications.
A compare-aggregate model with latent clustering for answer selection
Seunghyun Yoon, Franck Dernoncourt, Doo Soon Kim, Trung Bui, and Kyomin Jung. 2019 · 1905
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Well-read students learn better: On the importance of pre-training compact models
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 1908
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
A compare-aggregate model with dynamic-clip attention for answer selection
Weijie Bian, Si Li, Zhao Yang, Guang Chen, and Zhiqing Lin. 2017 · 1990
Earlier work this paper cites.
Hydra: Preserving ensemble diversity for model distillation
Linh Tran, Bastiaan S Veeling, Kevin Roth, Jakub Swiatkowski, Joshua V Dillon, Jasper Snoek, Stephan Mandt, Tim Salimans, Sebastian Nowozin, and Rodolphe Jenatton. 2020 · 2001
Earlier work this paper cites.
Zhuohan Li, Eric Wallace, Sheng Shen, Kevin Lin, Kurt Keutzer, Dan Klein, and Joseph E Gonzalez. 2020 · 2002
Earlier work this paper cites.
Improving BERT Fine-Tuning via Self-Ensemble and Self-Distillation
Yige Xu, Xipeng Qiu, Ligao Zhou, and Xuanjing Huang. 2020 · 2002
Earlier work this paper cites.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2003
Earlier work this paper cites.
What is the jeopardy model? a quasi-synchronous grammar for qa
Mengqiu Wang, Noah A Smith, and Teruko Mitamura. 2007 · 2007
Earlier work this paper cites.
Tradeoffs in sentence selection techniques for open-domain question answering
Shih-ting Lin and Greg Durrett. 2020 · 2009
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li. 2020 · 2012
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Learning to rank short text pairs with convolutional deep neural networks
Aliaksei Severyn and Alessandro Moschitti. 2015 · 2015
Earlier work this paper cites.
Query distillation: BERT-based distillation for ensemble ranking
Wangshu Zhang, Junhong Liu, Zujie Wen, Yafang Wang, and Gerard de Melo. 2020 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
A compare-aggregate model for matching text sequences
Shuohang Wang and Jing Jiang. 2016 · 2016
Earlier work this paper cites.
Inter-weighted alignment network for sentence pair modeling
Gehui Shen, Yunlun Yang, and Zhi-Hong Deng. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A compare-aggregate model for matching text sequences
Shuohang Wang and Jing Jiang. 2017 · 2017
Cited alongside, same era.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Integrating Question Classification and Deep Learning for improved Answer Selection
Harish Tayyar Madabushi, Mark Lee, and John Barnden. 2018 · 2018
Cited alongside, same era.
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan Alistarh. 2018 · 2018
Cited alongside, same era.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2019 · 2019
Later among the works it cites.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019 · 2019
Later among the works it cites.
TANDA: Transfer and adapt pre-trained transformer models for answer sentence selection
Siddhant Garg, Thuy Vu, and Alessandro Moschitti. 2020 · 2020
Later among the works it cites.
TinyBERT: Distilling BERT for natural language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Multi-cast attention networks for retrieval-based question answering and response prediction
Yi Tay, Luu Anh Tuan, and Siu Cheung Hui. 2018 · 2018
Cited alongside, same era.
The context-dependent additive recurrent neural net
Quan Hung Tran, Tuan Lai, Gholamreza Haffari, Ingrid Zukerman, Trung Bui, and Hung Bui. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek. 2015 · 2018
Cited alongside, same era.
Optuna: A next-generation hyperparameter optimization framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019 · 2019
Cited alongside, same era.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Later among the works it cites.
Adaptive knowledge distillation based on entropy
Kisoo Kwon, Hwidong Na, Hoshik Lee, and Nam Soo Kim. 2020 · 2020
Later among the works it cites.
MixKD: Towards Efficient Distillation of Large-scale Language Models
Kevin J Liang, Weituo Hao, Dinghan Shen, Yufan Zhou, Weizhu Chen, Changyou Chen, and Lawrence Carin. 2020 · 2020
Later among the works it cites.
Rikinet: Reading wikipedia pages for natural question answering
Dayiheng Liu, Yeyun Gong, Jie Fu, Yu Yan, J. Chen, Daxin Jiang, J. Lv, and N. Duan. 2020 · 2020
Later among the works it cites.
TwinBERT: Distilling Knowledge to Twin-Structured Compressed BERT Models for Large-Scale Retrieval
Wenhao Lu, Jian Jiao, and Ruofei Zhang. 2020 · 2020
Later among the works it cites.
Reranking for Efficient Transformer-based Answer Selection
Yoshitomo Matsubara, Thuy Vu, and Alessandro Moschitti. 2020 · 2020
Later among the works it cites.
XtremeDistil: Multi-stage Distillation for Massive Multilingual Models
Subhabrata Mukherjee and Ahmed Hassan Awadallah. 2020 · 2020
Later among the works it cites.
The cascade transformer: an application for efficient answer sentence selection
Luca Soldaini and Alessandro Moschitti. 2020 · 2020
Later among the works it cites.
MobileBERT: a compact task-agnostic BERT for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Model compression with two-stage multi-teacher knowledge distillation for web question answering system
Ze Yang, Linjun Shou, Ming Gong, Wutao Lin, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Modeling Context in Answer Sentence Selection Systems on a Latency Budget
Rujun Han, Luca Soldaini, and Alessandro Moschitti. 2021 · 2021
Later among the works it cites.
Answer Sentence Selection Using Local and Global Context in Transformer Models
Ivano Lauriola and Alessandro Moschitti. 2021 · 2021
Later among the works it cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2021 · 2021
Later among the works it cites.