Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated great capabilities in natural language understanding and generation, largely attributed to the intricate alignment process using human feedback.
Language models are few-shot learners
1901
Earlier work this paper cites.
How propaganda works
2015
Earlier work this paper cites.
Question answer system for online feedable new born chatbot
2017
Earlier work this paper cites.
Toxic comment classification challenge, 2017
2017
Earlier work this paper cites.
Superagent: A customer service chatbot for e-commerce websites
2017
Earlier work this paper cites.
Dailydialog: A manually labelled multi-turn dialogue dataset
2017
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
2018
Earlier work this paper cites.
Chatbot as an intelligent personal assistant for mobile language learning
2018
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
2019
Earlier work this paper cites.
The curious case of neural text degeneration
2019
Earlier work this paper cites.
Chatbot and conversational analysis to promote collaborative learning in distance education
2019
Earlier work this paper cites.
The woman worked as a babysitter: On biases in language generation
2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing NLP
2019
Earlier work this paper cites.
Semeval-2019 task 6: Identifying and categorizing offensive language in social media (offenseval)
2019
Earlier work this paper cites.
Gender bias in contextualized word embeddings
2019
Earlier work this paper cites.
Tweeteval: Unified benchmark and comparative evaluation for tweet classification
2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
2020
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models
2020
Cited alongside, same era.
Onion: A simple and effective defense against textual backdoor attacks
2020
Cited alongside, same era.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
2020
Cited alongside, same era.
Trl: Transformer reinforcement learning
2020
Cited alongside, same era.
Concealed data poisoning attacks on nlp models
2020
Cited alongside, same era.
Training language models to follow instructions with human feedback
2022
Later among the works it cites.
Why so toxic? measuring and triggering toxic behavior in open-domain chatbots
2022
Later among the works it cites.
Education in the era of generative artificial intelligence (ai): Understanding the potential benefits of chatgpt in promoting teaching and learning
2023
Later among the works it cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
2023
Later among the works it cites.
More than a feeling: Accuracy and application of sentiment analysis
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, Mar. 2021
2021
Cited alongside, same era.
Extracting training data from large language models
2021
Cited alongside, same era.
Badpre: Task-agnostic backdoor attacks to pre-trained nlp foundation models
2021
Cited alongside, same era.
Badnl: Backdoor attacks against nlp models with semantic-preserving improvements
2021
Cited alongside, same era.
Hidden backdoors in human-centric language models
2021
Cited alongside, same era.
Mind the style of text! adversarial and backdoor attacks based on text style transfer
2021
Cited alongside, same era.
Turn the combination lock: Learnable textual backdoor attacks via word substitution
2021
Cited alongside, same era.
2023
Later among the works it cites.
Notable: Transferable backdoor attacks against prompt-based nlp models
2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
2023
Later among the works it cites.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt
2023
Later among the works it cites.
Defending against backdoor attacks in natural language generation
2023
Later among the works it cites.
Poisoning language models during instruction tuning
2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
2023
Later among the works it cites.
How do you use personal data in model training?
2024
Closest in time.
Best-of-venom: Attacking rlhf by injecting poisoned preference data
2024
Closest in time.
How your data is used to improve model performance
2024
Closest in time.