Fetching the paper…
Reading the bibliography…
It has become a popular paradigm to transfer the knowledge of large-scale pre-trained models to various downstream tasks via fine-tuning the entire model parameters.
Language Models Are Few-Shot Learners. In Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin (Eds.), Vol. 33. 1877–1901
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Graph-Bert: Only Attention Is Needed for Learning Graph Representations
Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun. 2020 · 2001
Earlier work this paper cites.
Molecule Attention Transformer
Łukasz Maziarka, Tomasz Danel, Sławomir Mucha, Krzysztof Rataj, Jacek Tabor, and Stanisław Jastrzębski. 2020 · 2002
Earlier work this paper cites.
Open Graph Benchmark: Datasets for Machine Learning on Graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. 2021b · 2005
Earlier work this paper cites.
ChEMBL: A Large-Scale Bioactivity Database for Drug Discovery
A. Gaulton, L. J. Bellis, A. P. Bento, J. Chambers, M. Davies, A. Hersey, Y. Light, S. McGlinchey, D. Michalovich, B. Al-Lazikani, and J. P. Overington. 2012 · 2012
Earlier work this paper cites.
ZINC 15 – Ligand Discovery for Everyone
Teague Sterling and John J. Irwin. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Representations
Thomas N. Kipf and Max Welling. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need. In Advances in Neural Information Processing Systems , Vol. 30
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
MetStabOn—Online Platform for Metabolic Stability Predictions
Sabina Podlewska and Rafał Kafel. 2018 · 2018
Earlier work this paper cites.
Improving Language Understanding by Generative Pre-Training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
MoleculeNet: A Benchmark for Molecular Machine Learning
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP. In Proceedings of the 36th International Conference on Machine Learning . 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Language Models Are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning. In Proceedings of the 36th International Conference on Machine Learning . 5986–5995
Asa Cooper Stickland and Iain Murray. 2019 · 2019
Earlier work this paper cites.
How Powerful Are Graph Neural Networks?. In International Conference on Learning Representations
Keyulu Xu*, Weihua Hu*, Jure Leskovec, and Stefanie Jegelka. 2019 · 2019
Earlier work this paper cites.
Strategies for Pre-training Graph Neural Networks. In International Conference on Learning Representations
Weihua Hu*, Bowen Liu*, Joseph Gomes, Marinka Zitnik, Percy Liang, Vijay Pande, and Jure Leskovec. 2020 · 2020
Earlier work this paper cites.
GPT-GNN: Generative Pre-Training of Graph Neural Networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’20) . 1857–1867
Ziniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang, and Yizhou Sun. 2020 · 2020
Earlier work this paper cites.
SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized Optimization. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 2177–2190
Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Tuo Zhao. 2020 · 2020
Earlier work this paper cites.
Interpretable Rumor Detection in Microblogs by Attending to User Interactions. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 8783–8790
Ling Min Serena Khoo, Hai Leong Chieu, Zhong Qian, and Jing Jiang. 2020 · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
AdapterHub: A Framework for Adapting Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations . 46–54
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020 · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Self-Supervised Graph Transformer on Large-Scale Molecular Data. In Advances in Neural Information Processing Systems , Vol. 33. 12559–12571
Yu Rong, Yatao Bian, Tingyang Xu, Weiyang Xie, Ying WEI, Wenbing Huang, and Junzhou Huang. 2020 · 2020
Cited alongside, same era.
Towards a Unified View of Parameter-Efficient Transfer Learning. In International Conference on Learning Representations
Junxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick, and Graham Neubig. 2022 · 2022
Later among the works it cites.
FacT: Factor-Tuning for Lightweight Adaptation on Vision Transformer
Shibo Jie and Zhi-Hong Deng. 2022 · 2022
Later among the works it cites.
Pure Transformers Are Powerful Graph Learners. In Advances in Neural Information Processing Systems
Jinwoo Kim, Dat Tien Nguyen, Seonwoo Min, Sungjun Cho, Moontae Lee, Honglak Lee, and Seunghoon Hong. 2022 · 2022
Later among the works it cites.
Rethinking Graph Transformers with Spectral Attention. In Advances in Neural Information Processing Systems
Devin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau, and Prudencio Tossou. 2022 · 2022
Later among the works it cites.
Scaling & Shifting Your Features: A New Baseline for Efficient Model Tuning. In Advances in Neural Information Processing Systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Intrinsic Dimensionality Explains the Effectiveness of Language Model Fine-Tuning. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 7319–7328
Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. 2021 · 2021
Cited alongside, same era.
Parameter-Efficient Multi-task Fine-tuning for Transformers via Shared Hypernetworks. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 565–576
Rabeeh Karimi Mahabadi, Sebastian Ruder, Mostafa Dehghani, and James Henderson. 2021 · 2021
Cited alongside, same era.
The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 3045–3059
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Cited alongside, same era.
Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 4582–4597
Xiang Lisa Li and Percy Liang. 2021 · 2021
Cited alongside, same era.
Mesh Graphormer. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . 12919–12928
Kevin Lin, Lijuan Wang, and Zicheng Liu. 2021 · 2021
Cited alongside, same era.
Compacter: Efficient Low-Rank Hypercomplex Adapter Layers. In Advances in Neural Information Processing Systems
Rabeeh Karimi Mahabadi, James Henderson, and Sebastian Ruder. 2021 · 2021
Cited alongside, same era.
GraphiT: Encoding Graph Structure in Transformers
Grégoire Mialon, Dexiong Chen, Margot Selosse, and Julien Mairal. 2021 · 2021
Cited alongside, same era.
AdapterFusion: Non-Destructive Task Composition for Transfer Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume . 487–503
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
Dongze Lian, Zhou Daquan, Jiashi Feng, and Xinchao Wang. 2022 · 2022
Later among the works it cites.
P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) . 61–68
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022 · 2022
Later among the works it cites.
UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning
Yuning Mao, Lambert Mathias, Rui Hou, Amjad Almahairi, Hao Ma, Jiawei Han, Wen-tau Yih, and Madian Khabsa. 2022 · 2022
Later among the works it cites.
Transformer for Graphs: An Overview from Architecture Perspective
Erxue Min, Runfa Chen, Yatao Bian, Tingyang Xu, Kangfei Zhao, Wenbing Huang, Peilin Zhao, Junzhou Huang, Sophia Ananiadou, and Yu Rong. 2022 · 2022
Later among the works it cites.
SPoT: Better Frozen Model Adaptation through Soft Prompt Transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 5039–5059
Tu Vu, Brian Lester, Noah Constant, Rami Al-Rfou’, and Daniel Cer. 2022 · 2022
Later among the works it cites.
AdaMix: Mixture-of-Adapter for Parameter-efficient Tuning of Large Language Models
Yaqing Wang, Subhabrata Mukherjee, Xiaodong Liu, Jing Gao, Ahmed Hassan Awadallah, and Jianfeng Gao. 2022 · 2022
Later among the works it cites.
A Survey of Pretraining on Graphs: Taxonomy, Methods, and Applications
Jun Xia, Yanqiao Zhu, Yuanqi Du, and Stan Z. Li. 2022 · 2022
Later among the works it cites.
Towards a Unified View on Visual Parameter-Efficient Transfer Learning
Bruce X. B. Yu, Jianlong Chang, Lingbo Liu, Qi Tian, and Chang Wen Chen. 2022 · 2022
Later among the works it cites.
Specformer: Spectral Graph Neural Networks Meet Transformers. In The Eleventh International Conference on Learning Representations
Deyu Bo, Chuan Shi, Lele Wang, and Renjie Liao. 2023 · 2023
Closest in time.
Relational Attention: Generalizing Transformers for Graph-Structured Tasks. In The Eleventh International Conference on Learning Representations
Cameron Diao and Ricky Loynd. 2023 · 2023
Closest in time.
Rethinking Efficient Tuning Methods from a Unified Perspective
Zeyinzi Jiang, Chaojie Mao, Ziyuan Huang, Yiliang Lv, Deli Zhao, and Jingren Zhou. 2023 · 2023
Closest in time.
Edgeformers: Graph-Empowered Transformers for Representation Learning on Textual-Edge Networks. In The Eleventh International Conference on Learning Representations
Bowen Jin, Yu Zhang, Yu Meng, and Jiawei Han. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Multitask Prompt Tuning Enables Parameter-Efficient Transfer Learning. In The Eleventh International Conference on Learning Representations
Zhen Wang, Rameswar Panda, Leonid Karlinsky, Rogerio Feris, Huan Sun, and Yoon Kim. 2023 · 2023
Closest in time.
Monocular Scene Reconstruction with 3D SDF Transformers. In The Eleventh International Conference on Learning Representations
Weihao Yuan, Xiaodong Gu, Heng Li, Zilong Dong, and Siyu Zhu. 2023 · 2023
Closest in time.
Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning. In The Eleventh International Conference on Learning Representations
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023 · 2023
Closest in time.
AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning
Han Zhou, Xingchen Wan, Ivan Vulić, and Anna Korhonen. 2023 · 2023
Closest in time.