Fetching the paper…
Reading the bibliography…
We investigate a new threat to neural sequence-to-sequence (seq2seq) models: training-time attacks that cause models to "spin" their outputs so as to support an adversary-chosen sentiment or point of view -- but only when the input contains adversary-chosen trigger words.
E. H. Henderson, “Toward a definition of propaganda,” The Journal of Social Psychology , 1943
1943
Earlier work this paper cites.
F. R. Hampel, “The influence curve and its role in robust estimation,” JASA , 1974
1974
Earlier work this paper cites.
S. Hidi and V. Anderson, “Producing written summaries: Task demands, cognitive operations, and implications for instruction,” Review of Educational Research , 1986
1986
Earlier work this paper cites.
P. J. Rousseeuw and C. Croux, “Alternatives to the median absolute deviation,” JASA , 1993
1993
Earlier work this paper cites.
A. Ratnaparkhi, “A maximum entropy model for part-of-speech tagging,” in EMNLP , 1996
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , 1997
1997
Earlier work this paper cites.
I. Gaber, “Government by spin: An analysis of the process,” Media, Culture & Society , 2000
2000
Earlier work this paper cites.
J. A. Maltese, Spin control: The White House Office of Communications and the management of presidential news . Univ of North Carolina Press, 2000
2000
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “BLEU: A method for automatic evaluation of machine translation,” in ACL , 2002
2002
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” in ACL Workshop , 2004
2004
Earlier work this paper cites.
C. Bannard and C. Callison-Burch, “Paraphrasing with bilingual parallel corpora,” in ACL , 2005
2005
Earlier work this paper cites.
C. W. Tindale, Fallacies and Argument Appraisal . Cambridge University Press, 2007
2007
Earlier work this paper cites.
D. Miller and W. Dinan, A century of spin: How public relations became the cutting edge of corporate power . Pluto Press, 2008
2008
Earlier work this paper cites.
E. S. Herman and N. Chomsky, Manufacturing consent: The political economy of the mass media . Random House, 2010
2010
Earlier work this paper cites.
R. R. Wilcox, Introduction to Robust Estimation and Hypothesis Testing . Academic Press, 2011
2011
Earlier work this paper cites.
B. Biggio, B. Nelson, and P. Laskov, “Poisoning attacks against support vector machines,” in ICML , 2012
2012
Earlier work this paper cites.
Z. Chu, S. Gianvecchio, H. Wang, and S. Jajodia, “Detecting automation of Twitter accounts: Are you a human, bot, or cyborg?” IEEE Trans. Dependable and Secure Computing , 2012
2012
Earlier work this paper cites.
J.-A. Désidéri, “Multiple-gradient descent algorithm (MGDA) for multiobjective optimization,” Comptes Rendus Mathématique , 2012
2012
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NIPS , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in ICLR , 2015
2015
Earlier work this paper cites.
K. M. Hermann, T. Kocisky, E. Grefenstette, L. Espeholt, W. Kay, M. Suleyman, and P. Blunsom, “Teaching machines to read and comprehend,” in NIPS , 2015
2015
Earlier work this paper cites.
J. Stanley, How Propaganda Works . Princeton University Press, 2015
2015
Earlier work this paper cites.
X. Zhang, J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” in NIPS , 2015
2015
Earlier work this paper cites.
O. Bojar, R. Chatterjee, C. Federmann, Y. Graham, B. Haddow, M. Huck, A. Jimeno Yepes, P. Koehn, V. Logacheva, C. Monz, M. Negri, A. Neveol, M. Neves, M. Popel, M. Post, R. Rubino, C. Scarton, L. Specia, M. Turchi, K. Verspoor, and M. Zampieri, “Findings of the 2016 conference on machine translation,” in WMT , 2016
2016
Earlier work this paper cites.
M. Gabielkov, A. Ramachandran, A. Chaintreau, and A. Legout, “Social clicks: What and who gets read on twitter?” in SIGMETRICS , 2016
2016
Earlier work this paper cites.
A. Caliskan, J. J. Bryson, and A. Narayanan, “Semantics derived automatically from language corpora contain human-like biases,” Science , 2017
2017
Earlier work this paper cites.
J. W. Schwieter, A. Ferreira, and J. Wiley, The Handbook of Translation and Cognition . Wiley Online Library, 2017
2017
Earlier work this paper cites.
A. See, P. J. Liu, and C. D. Manning, “Get to the point: Summarization with pointer-generator networks,” in ACL , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Earlier work this paper cites.
S. Volkova, K. Shaffer, J. Y. Jang, and N. Hodas, “Separating facts from fiction: Linguistic models to classify suspicious and trusted news posts on Twitter,” in ACL , 2017
2017
Earlier work this paper cites.
M. Alzantot, Y. Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang, “Generating natural language adversarial examples,” in EMNLP , 2018
2018
Earlier work this paper cites.
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou, “HotFlip: White-box adversarial examples for text classification,” in ACL , 2018
2018
Earlier work this paper cites.
M. Grusky, M. Naaman, and Y. Artzi, “Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies,” in NAACL , 2018
2018
Cited alongside, same era.
M. Junczys-Dowmunt, R. Grundkiewicz, T. Dwojak, H. Hoang, K. Heafield, T. Neckermann, F. Seide, U. Germann, A. F. Aji, N. Bogoychev, A. F. T. Martins, and A. Birch, “Marian: Fast neural machine translation in C++,” in ACL System Demonstrations , 2018
2018
Cited alongside, same era.
K. Liu, B. Dolan-Gavitt, and S. Garg, “Fine-pruning: Defending against backdooring attacks on deep neural networks,” in RAID , 2018
2018
Cited alongside, same era.
S. Narayan, S. B. Cohen, and M. Lapata, “Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization,” in EMNLP , 2018
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” OpenAI Blog , 2018
E. Durmus, H. He, and M. Diab, “FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization,” in ACL , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. T. Hancock, M. Naaman, and K. Levy, “AI-mediated communication: Definition, research agenda, and ethical considerations,” J. Computer-Mediated Communication , 2020
2020
Later among the works it cites.
L. Hanu and Unitary team, “Detoxify,” https://github.com/unitaryai/detoxify , 2020
2020
Later among the works it cites.
J. Hohenstein and M. Jung, “AI as a moral crumple zone: The effects of AI-mediated communication on attribution and trust,” Computers in Human Behavior , 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
O. Sener and V. Koltun, “Multi-task learning as multi-objective optimization,” in NIPS , 2018
2018
Cited alongside, same era.
B. Tran, J. Li, and A. Madry, “Spectral signatures in backdoor attacks,” in NIPS , 2018
2018
Cited alongside, same era.
A. Weston, A Rulebook for Arguments . Hackett Publishing, 2018
2018
Cited alongside, same era.
A. Williams, N. Nangia, and S. Bowman, “A broad-coverage challenge corpus for sentence understanding through inference,” in NAACL , 2018
2018
Cited alongside, same era.
B. Chen, W. Carvalho, N. Baracaldo, H. Ludwig, B. Edwards, T. Lee, I. Molloy, and B. Srivastava, “Detecting backdoor attacks on deep neural networks by activation clustering,” in SafeAI@AAAI , 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Cited alongside, same era.
Y. Gao, C. Xu, D. Wang, S. Chen, D. C. Ranasinghe, and S. Nepal, “STRIP: A defence against trojan attacks on deep neural networks,” in ACSAC , 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
K. Kurita, P. Michel, and G. Neubig, “Weight poisoning attacks on pre-trained models,” in ACL , 2020
2020
Later among the works it cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL , 2020
2020
Later among the works it cites.
Y. Li, B. Wu, Y. Jiang, Z. Li, and S.-T. Xia, “Backdoor learning: A survey,” arXiv:2007.08745 , 2020
2020
Later among the works it cites.
J. Mackenzie, R. Benham, M. Petri, J. R. Trippas, J. S. Culpepper, and A. Moffat, “CC-News-En: A large English news corpus,” in CIKM , 2020
2020
Later among the works it cites.
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela, “Adversarial NLI: A new benchmark for natural language understanding,” in ACL , 2020
2020
Later among the works it cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” JMLR , 2020
2020
Later among the works it cites.
R. Schuster, T. Schuster, Y. Meri, and V. Shmatikov, “Humpty dumpty: Controlling word meanings via corpus poisoning,” in S&P , 2020
2020
Later among the works it cites.
S. Tan, S. Joty, M.-Y. Kan, and R. Socher, “It’s morphin’ time! Combating linguistic discrimination with inflectional perturbations,” in ACL , 2020
2020
Later among the works it cites.
A. Toral, “Reassessing claims of human parity and super-human performance in machine translation at WMT 2019,” in EAMT , 2020
2020
Later among the works it cites.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush, “Transformers: State-of-the-art natural language processing,” in EMNLP: System Demonstrations , 2020
2020
Later among the works it cites.
J. Zhang, Y. Zhao, M. Saleh, and P. Liu, “PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization,” in ICML , 2020
2020
Later among the works it cites.
E. Bagdasaryan and V. Shmatikov, “Blind backdoors in deep learning models,” in USENIX Security , 2021
2021
Closest in time.
B. Buchanan, A. Lohn, M. Musser, and K. Sedova, “Truth, lies, and automation,” Center for Security and Emerging Technology , 2021
2021
Closest in time.
A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev, “SummEval: Re-evaluating Summarization Evaluation,” TACL , 2021
2021
Closest in time.
2021
Closest in time.
S. Li, H. Liu, T. Dong, B. Z. H. Zhao, M. Xue, H. Zhu, and J. Lu, “Hidden backdoors in human-centric language models,” in CCS , 2021
2021
Closest in time.
J. Rae, G. Irving, and L. Weidinger, “Language modelling at scale: Gopher, ethical considerations, and retrieval,” in DeepMind Blog , 2021
2021
Closest in time.
P. Remy, “Name dataset,” https://github.com/philipperemy/name-dataset , 2021
2021
Closest in time.
R. Schuster, C. Song, E. Tromer, and V. Shmatikov, “You autocomplete me: Poisoning vulnerabilities in neural code completion,” in USENIX Security , 2021
2021
Closest in time.
E. Wallace, T. Z. Zhao, S. Feng, and S. Singh, “Customizing triggers with concealed data poisoning,” in NAACL , 2021
2021
Closest in time.
J. Wang, C. Xu, F. Guzmán, A. El-Kishky, Y. Tang, B. Rubinstein, and T. Cohn, “Putting words into the system’s mouth: A targeted attack on neural machine translation using monolingual data poisoning,” in ACL-IJCNLP , 2021
2021
Closest in time.
E. Winer, “Funny Names,” https://ethanwiner.com/funnames.html , 2021
2021
Closest in time.
W. Yang, L. Li, Z. Zhang, X. Ren, X. Sun, and B. He, “Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in NLP models,” in NAACL-HLT , 2021
2021
Closest in time.
X. Zhang, Z. Zhang, S. Ji, and T. Wang, “Trojaning language models for fun and profit,” in EuroS&P , 2021
2021
Closest in time.
2021
Closest in time.
K. Chen, Y. Meng, X. Sun, S. Guo, T. Zhang, J. Li, and C. Fan, “BadPre: Task-agnostic backdoor attacks to pre-trained NLP foundation models,” in ICLR , 2022
2022
Closest in time.
J. Jia, Y. Liu, and N. Z. Gong, “BadEncoder: Backdoor attacks to pre-trained encoders in self-supervised learning,” in S&P , 2022
2022
Closest in time.