Fetching the paper…
Reading the bibliography…
Social media data exhibits severe redundancy caused by its noisy nature.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Reinforced training data selection for domain adaptation
Miaofeng Liu, Yan Song, Hongbin Zou, and Tong Zhang. 2019a · 1968
Earlier work this paper cites.
On the resemblance and containment of documents
Andrei Z Broder. 1997 · 1997
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, et al. 2020 · 2001
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Detecting near-duplicates for web crawling
Gurmeet Singh Manku, Arvind Jain, and Anish Das Sarma. 2007 · 2007
Earlier work this paper cites.
Adaptive near-duplicate detection via similarity learning
Hannaneh Hajishirzi, Wen-tau Yih, and Aleksander Kolcz. 2010 · 2010
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis. 2010 · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Groundhog day: near-duplicate detection on twitter
Ke Tao, Fabian Abel, Claudia Hauff, Geert-Jan Houben, and Ujwal Gadiraju. 2013 · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 · 2015
Earlier work this paper cites.
Data selection strategies for multi-domain sentiment analysis
Sebastian Ruder, Parsa Ghaffari, and John G Breslin. 2017 · 2017
Earlier work this paper cites.
Abstractive summarization of reddit posts with multi-level memory networks
Byeongchang Kim, Hyunwoo Kim, and Gunhee Kim. 2018 · 2018
Earlier work this paper cites.
Topic memory networks for short text classification
Jichuan Zeng, Jing Li, Yan Song, Cuiyun Gao, Michael R. Lyu, and Irwin King. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
TweetEval: Unified benchmark and comparative evaluation for tweet classification
Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020 · 2020
Cited alongside, same era.
Selection via proxy: Efficient data selection for deep learning
Cody Coleman, Christopher Yeh, et al. 2020 · 2020
Cited alongside, same era.
GoEmotions: A Dataset of Fine-Grained Emotions
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020 · 2020
Scaling laws and interpretability of learning from repeated data
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. 2022 · 2022
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Later among the works it cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Sören Mindermann, Jan M Brauner, et al. 2022 · 2022
Later among the works it cites.
Nlp from scratch without large-scale pretraining: A simple and efficient framework
Xingcheng Yao, Yanan Zheng, Xiaocong Yang, and Zhilin Yang. 2022 · 2022
Later among the works it cites.
Can ChatGPT understand causal language in science claims?
Yuheun Kim, Lu Guo, Bei Yu, and Yingya Li. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Keybert: Minimal keyword extraction with bert
Maarten Grootendorst. 2020 · 2020
Cited alongside, same era.
BERTweet: A pre-trained language model for English tweets
Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020 · 2020
Cited alongside, same era.
Dynamic online conversation recommendation
Xingshan Zeng, Jing Li, Lu Wang, Zhiming Mao, and Kam-Fai Wong. 2020 · 2020
Cited alongside, same era.
Simcse: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Cited alongside, same era.
Stance detection in COVID-19 tweets
Kyle Glandt, Sarthak Khanal, Yingjie Li, Doina Caragea, and Cornelia Caragea. 2021 · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Angle-optimized text embeddings
Xianming Li and Jing Li. 2023 · 2023
Later among the works it cites.
Data selection for fine-tuning large language models using transferred shapley values
Stephanie Schoch, Ritwick Mishra, and Yangfeng Ji. 2023 · 2023
Later among the works it cites.
Hicl: Hashtag-driven in-context learning for social media natural language understanding
Hanzhuo Tan, Chunpu Xu, Jing Li, Yuqun Zhang, Zeyang Fang, Zeyu Chen, and Baohua Lai. 2023 · 2023
Later among the works it cites.
D4: Improving LLM pretraining via document de-duplication and diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari S. Morcos. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. 2023 · 2023
Later among the works it cites.
Cold-start data selection for better few-shot language model fine-tuning: A prompt-based uncertainty propagation approach
Yue Yu, Rongzhi Zhang, Ran Xu, Jieyu Zhang, Jiaming Shen, and Chao Zhang. 2023 · 2023
Later among the works it cites.
VIBE: Topic-driven temporal adaptation for Twitter classification
Yuji Zhang, Jing Li, and Wenjie Li. 2023 · 2023
Later among the works it cites.
BeLLM: Backward dependency enhanced large language model for sentence embeddings
Xianming Li and Jing Li. 2024b · 2024
Closest in time.
Less: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024 · 2024
Closest in time.