Fetching the paper…
Reading the bibliography…
Recent advancement of large-scale pretrained models such as BERT, GPT-3, CLIP, and Gopher, has shown astonishing achievements across various task domains.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; et al. 2020 · 1901
Earlier work this paper cites.
Neural networks for optimal approximation of smooth and analytic functions
Mhaskar, H. N. 1996 · 1996
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; et al. 2020 · 2001
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Hadsell, R.; Chopra, S.; and LeCun, Y. 2006 · 2006
Earlier work this paper cites.
The Kendall rank correlation coefficient
Abdi, H. 2007 · 2007
Earlier work this paper cites.
Factorization machines
Rendle, S. 2010 · 2010
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R.; Mikolov, T.; and Bengio, Y. 2013 · 2013
Earlier work this paper cites.
DeepFM: a factorization-machine based neural network for CTR prediction
Guo, H.; Tang, R.; Ye, Y.; Li, Z.; and He, X. 2017 · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Kipf, T. N.; and Welling, M. 2017 · 2017
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Nsml: A machine learning platform that enables you to focus on your models
Sung, N.; Kim, M.; Jo, H.; Yang, Y.; Kim, J.; Lausen, L.; Kim, Y.; et al. 2017 · 2017
Earlier work this paper cites.
Nsml: Meet the mlaas platform with a real-world case study
Kim, H.; Kim, M.; Seo, D.; Kim, J.; Park, H.; Park, S.; et al. 2018 · 2018
Earlier work this paper cites.
Mixed precision training
Micikevicius, P.; Narang, S.; Alben, J.; Diamos, G.; et al. 2018 · 2018
Earlier work this paper cites.
Behavior sequence transformer for e-commerce recommendation in alibaba
Chen, Q.; Zhao, H.; et al. 2019 · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; et al. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Ni, J.; Li, J.; and McAuley, J. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; et al. 2019 · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Cited alongside, same era.
Recommending what video to watch next: a multitask ranking system
Zhao, Z.; Hong, L.; Wei, L.; Chen, J.; Nath, A.; Andrews, S.; Kumthekar, A.; Sathiamoorthy, M.; Yi, X.; and Chi, E. 2019 · 2019
Cited alongside, same era.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Baevski, A.; Zhou, Y.; Mohamed, A.; and Auli, M. 2020 · 2020
Cited alongside, same era.
Unsupervised learning of visual features by contrasting cluster assignments
Scaling deep contrastive learning batch size under memory limited setup
Gao, L.; Zhang, Y.; Han, J.; and Callan, J. 2021 · 2021
Closest in time.
SimCSE: Simple Contrastive Learning of Sentence Embeddings
Gao, T.; Yao, X.; and Chen, D. 2021 · 2021
Closest in time.
Self-supervised pretraining of visual features in the wild
Goyal, P.; Caron, M.; Lefaudeux, B.; Xu, M.; Wang, P.; Pai, V.; et al. 2021 · 2021
Closest in time.
Exploiting behavioral consistence for universal user representation
Gu, J.; Wang, F.; Sun, Q.; Ye, Z.; Xu, X.; Chen, J.; and Zhang, J. 2021 · 2021
Closest in time.
Hutter, M. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Caron, M.; Misra, I.; Mairal, J.; Goyal, P.; Bojanowski, P.; and Joulin, A. 2020 · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
Chen, T.; Kornblith, S.; Norouzi, M.; and Hinton, G. 2020 · 2020
Cited alongside, same era.
Don’t stop pretraining: adapt language models to domains and tasks
Gururangan, S.; Marasović, A.; Swayamdipta, S.; Lo, K.; Beltagy, I.; Downey, D.; and Smith, N. A. 2020 · 2020
Cited alongside, same era.
Lightgcn: Simplifying and powering graph convolution network for recommendation
He, X.; Deng, K.; Wang, X.; Li, Y.; Zhang, Y.; and Wang, M. 2020 · 2020
Cited alongside, same era.
div2vec: Diversity-Emphasized Node Embedding
Jeong, J.; Yun, J.-M.; Keam, H.; et al. 2020 · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S.; Rasley, J.; et al. 2020 · 2020
Cited alongside, same era.
Neural machine translation with byte-level subwords
Wang, C.; Cho, K.; and Gu, J. 2020 · 2020
Cited alongside, same era.
What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers
Kim, B.; Kim, H.; Lee, S.-W.; et al. 2021 · 2021
Closest in time.
Representation learning via invariant causal mechanisms
Mitrovic, J.; McWilliams, B.; Walker, J.; et al. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; et al. 2021 · 2021
Closest in time.
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Rae, J. W.; Borgeaud, S.; Cai, T.; Millican, K.; et al. 2021 · 2021
Closest in time.
One4all User Representation for Recommender Systems in E-commerce
Shin, K.; Kwak, H.; Kim, K.-M.; Kim, M.; Park, Y.-J.; Jeong, J.; and Jung, S. 2021 · 2021
Closest in time.
Interest-oriented Universal User Representation via Contrastive Learning
Sun, Q.; Gu, J.; Yang, B.; Xu, X.; et al. 2021 · 2021
Closest in time.
One person, one model, one world: Learning continual user representation without forgetting
Yuan, F.; Zhang, G.; Karatzoglou, A.; et al. 2021 · 2021
Closest in time.
Zhai, X.; Kolesnikov, A.; Houlsby, N.; and Beyer, L. 2021 · 2021
Closest in time.
Towards Universal Sequence Representation Learning for Recommender Systems
Hou, Y.; Mu, S.; Zhao, W. X.; Li, Y.; Ding, B.; and Wen, J.-R. 2022 · 2022
Closest in time.
Simvlm: Simple visual language model pretraining with weak supervision
Wang, Z.; Yu, J.; Yu, A. W.; Dai, Z.; Tsvetkov, Y.; and Cao, Y. 2022 · 2022
Closest in time.
UserBERT: Pre-Training User Model with Contrastive Self-Supervision
Wu, C.; Wu, F.; Qi, T.; and Huang, Y. 2022 · 2092
Closest in time.