Fetching the paper…
Reading the bibliography…
The trustworthiness of machine learning has emerged as a critical topic in the field, encompassing various applications and research areas such as robustness, security, interpretability, and fairness.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 1901
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Maynez, J., Narayan, S., Bohnet, B., and McDonald, R · 1919
Earlier work this paper cites.
A survey of race, racism, and anti-racism in NLP
Field, A., Blodgett, S. L., Waseem, Z., and Tsvetkov, Y · 1925
Earlier work this paper cites.
Tariff on animal and vegetable oils
Wright, P. G · 1928
Earlier work this paper cites.
Generalist vision foundation models for medical imaging: A case study of segment anything model on zero-shot medical segmentation
Shi, P., Qiu, J., Abaxi, S. M. D., Wei, H., Lo, F. P.-W., and Yuan, W · 1947
Earlier work this paper cites.
Generalist vision foundation models for medical imaging: A case study of segment anything model on zero-shot medical segmentation
Shi, P., Qiu, J., Abaxi, S. M. D., Wei, H., Lo, F. P.-W., and Yuan, W · 1947
Earlier work this paper cites.
Estimation of the parameters of a single equation in a complete system of stochastic equations
Anderson, T. W., and Rubin, H · 1949
Earlier work this paper cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A · 1965
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R · 1967
Earlier work this paper cites.
How sensitive are translation systems to extra contexts? mitigating gender bias in neural machine translation models through relevant contexts
Sharma, S., Dey, M., and Sinha, K · 1984
Earlier work this paper cites.
Advanced econometrics
Amemiya, T · 1985
Earlier work this paper cites.
Distributionally robust finetuning bert for covariate drift in spoken language understanding
Broscheit, S., Do, Q., and Gaspers, J · 1985
Earlier work this paper cites.
Estimation and simultaneous correlation in complete equation systems
Theil, H · 1992
Earlier work this paper cites.
Causal diagrams for empirical research
Pearl, J · 1995
Earlier work this paper cites.
Cycada: Cycle-consistent adversarial domain adaptation
Hoffman, J., Tzeng, E., Park, T., Zhu, J.-Y., Isola, P., Saenko, K., Efros, A., and Darrell, T · 1998
Earlier work this paper cites.
Confounding and collapsibility in causal inference
Greenland, S., Pearl, J., and Robins, J. M · 1999
Earlier work this paper cites.
Models, reasoning and inference
Pearl, J., et al · 2000
Earlier work this paper cites.
Causal models as minimal descriptions of multivariate systems, 2006
Lemeire, J., and Dirkx, E · 2006
Earlier work this paper cites.
Analysis of representations for domain adaptation
Ben-David, S., et al · 2007
Earlier work this paper cites.
Biographies, Bollywood, boom-boxes and blenders: Domain adaptation for sentiment classification
Blitzer, J., Dredze, M., and Pereira, F · 2007
Earlier work this paper cites.
Frustratingly easy domain adaptation
Daumé III, H · 2007
Earlier work this paper cites.
Employment tests and discriminatory hiring
GUION, R · 2008
Earlier work this paper cites.
Robust regression and lasso, 2008
Xu, H., Caramanis, C., and Mannor, S · 2008
Earlier work this paper cites.
Understanding confounding and mediation
Babyak, M. A · 2009
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Erhan, D., Bengio, Y., Courville, A., and Vincent, P · 2009
Earlier work this paper cites.
Covariate shift by kernel mean matching
Gretton, A., Smola, A. J., Huang, J., Schmittfull, M., Borgwardt, K. M., and Schölkopf, B · 2009
Earlier work this paper cites.
Domain adaptation: Learning bounds and algorithms
Mansour, Y., Mohri, M., and Rostamizadeh, A · 2009
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Pearl, J · 2009
Earlier work this paper cites.
When training and test sets are different: characterizing learning transfer
Storkey, A · 2009
Earlier work this paper cites.
Explaining instance classifications with interactions of subsets of feature values
Štrumbelj, E., Kononenko, I., and Šikonja, M. R · 2009
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W · 2010
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J., and Yang, Q · 2010
Earlier work this paper cites.
Adapting visual category models to new domains
Saenko, K., Kulis, B., Fritz, M., and Darrell, T · 2010
Earlier work this paper cites.
An efficient explanation of individual classifications using game theory
Strumbelj, E., and Kononenko, I · 2010
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A., and Efros, A. A · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A., and Efros, A. A · 2011
Earlier work this paper cites.
Undoing the damage of dataset bias
Khosla, A., Zhou, T., Malisiewicz, T., Efros, A. A., and Torralba, A · 2012
Earlier work this paper cites.
How to control confounding effects by statistical analysis
Pourhoseingholi, M. A., Baghestani, A. R., and Vahedi, M · 2012
Earlier work this paper cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Earlier work this paper cites.
Domain generalization via invariant feature representation
Muandet, K., Balduzzi, D., and Schölkopf, B · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., et al · 2013
Earlier work this paper cites.
On the definition of a confounder
VanderWeele, T. J., and Shpitser, I · 2013
Earlier work this paper cites.
Learning fair representations
Zemel, R., Wu, Y., Swersky, K., Pitassi, T., and Dwork, C · 2013
Earlier work this paper cites.
Domain adaptation under target and conditional shift
Zhang, K., Schölkopf, B., Muandet, K., and Wang, Z · 2013
Earlier work this paper cites.
Domain-adversarial neural networks
Ajakan, H., Germain, P., Larochelle, H., Laviolette, F., and Marchand, M · 2014
Earlier work this paper cites.
Instrumental variable methods for causal inference
Baiocchi, M., Cheng, J., and Small, D. S · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
Donahue, J., Jia, Y., Vinyals, O., Hoffman, J., Zhang, N., Tzeng, E., and Darrell, T · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation, 2014
Girshick, R., Donahue, J., Darrell, T., and Malik, J · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Towards deep neural network architectures robust to adversarial examples, 2014
Gu, S., and Rigazio, L · 2014
Earlier work this paper cites.
Confounding variables
McDonald, J · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Zeiler, M. D., and Fergus, R · 2014
Earlier work this paper cites.
Vqa: Visual question answering, 2015
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Batra, D., and Parikh, D · 2015
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., and Samek, W · 2015
Earlier work this paper cites.
Censoring representations with an adversary
Edwards, H., and Storkey, A · 2015
Earlier work this paper cites.
Fast r-cnn
Girshick, R · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples (2014)
Goodfellow, I. J., et al · 2015
Earlier work this paper cites.
Robust convolutional neural networks under adversarial noise
Jin, J., Dundar, A., and Culurciello, E · 2015
Earlier work this paper cites.
Manifold regularized deep neural networks using adversarial examples
Lee, T., Choi, M., and Yoon, S · 2015
Earlier work this paper cites.
Foveation-based mechanisms alleviate adversarial examples
Luo, Y., Boix, X., Roig, G., Poggio, T., and Zhao, Q · 2015
Earlier work this paper cites.
A unified gradient regularization family for adversarial examples
Lyu, C., Huang, K., and Liang, H.-N · 2015
Earlier work this paper cites.
Understanding deep image representations by inverting them
Mahendran, A., and Vedaldi, A · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., and Sun, J · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Distributionally robust logistic regression, 2015
Shafieezadeh-Abadeh, S., Esfahani, P. M., and Kuhn, D · 2015
Earlier work this paper cites.
Transitive transfer learning
Tan, B., Song, Y., Zhong, E., and Yang, Q · 2015
Earlier work this paper cites.
A deeper look at dataset bias, 2015
Tommasi, T., Patricia, N., Caputo, B., and Tuytelaars, T · 2015
Earlier work this paper cites.
Simultaneous deep transfer across domains and tasks
Tzeng, E., Hoffman, J., Darrell, T., and Saenko, K · 2015
Earlier work this paper cites.
Spurious correlations
Vigen, T · 2015
Earlier work this paper cites.
Hyper-class augmented and regularized deep learning for fine-grained image classification
Xie, S., Yang, T., Wang, X., and Lin, Y · 2015
Earlier work this paper cites.
Understanding neural networks through deep visualization
Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., and Lipson, H · 2015
Earlier work this paper cites.
Supervised representation learning: Transfer learning with deep autoencoders
Zhuang, F., Cheng, X., Luo, P., Pan, S. J., and He, Q · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L · 2016
Earlier work this paper cites.
Domain separation networks
Bousmalis, K., Trigeorgis, G., Silberman, N., Krishnan, D., and Erhan, D · 2016
Earlier work this paper cites.
A new pac-bayesian perspective on domain adaptation
Germain, P., Habrard, A., Laviolette, F., and Morvant, E · 2016
Earlier work this paper cites.
Causal inference in statistics: A primer
Glymour, M., Pearl, J., and Jewell, N. P · 2016
Earlier work this paper cites.
Satisfying real-world goals with dataset constraints
Goh, G., Cotter, A., Gupta, M., and Friedlander, M. P · 2016
Earlier work this paper cites.
Robust text classification in the presence of confounding bias
Landeiro, V., and Culotta, A · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Lei, T., Barzilay, R., and Jaakkola, T · 2016
Earlier work this paper cites.
Deepfool: a simple and accurate method to fool deep neural networks
Moosavi-Dezfooli, S.-M., et al · 2016
Earlier work this paper cites.
Stochastic gradient methods for distributionally robust optimization with f-divergences
Namkoong, H., and Duchi, J. C · 2016
Earlier work this paper cites.
Simple black-box adversarial perturbations for deep networks, 2016
Narodytska, N., and Kasiviswanathan, S. P · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J · 2016
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
Papernot, N., McDaniel, P., and Goodfellow, I · 2016
Earlier work this paper cites.
The limitations of deep learning in adversarial settings
Papernot, N., McDaniel, P., Jha, S., Fredrikson, M., Celik, Z. B., and Swami, A · 2016
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks
Papernot, N., McDaniel, P., Wu, X., Jha, S., and Swami, A · 2016
Earlier work this paper cites.
” why should i trust you?” explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
Adversarial diversity and hard positive generation
Rozsa, A., Rudd, E. M., and Boult, T. E · 2016
Earlier work this paper cites.
Regularization with stochastic transformations and perturbations for deep semi-supervised learning, 2016
Sajjadi, M., Javanmardi, M., and Tasdizen, T · 2016
Earlier work this paper cites.
Using non-invertible data transformations to build adversary-resistant deep neural networks
Wang, Q., Guo, W., Alexander, G., Ororbia, I., Xing, X., Lin, L., Giles, C. L., Liu, X., Liu, P., and Xiong, G · 2016
Earlier work this paper cites.
Learning adversary-resistant deep neural networks
Wang, Q., Guo, W., Zhang, K., Ororbia, I., Alexander, G., Xing, X., Liu, X., and Giles, C. L · 2016
Earlier work this paper cites.
A survey of transfer learning
Weiss, K., Khoshgoftaar, T. M., and Wang, D · 2016
Earlier work this paper cites.
Improving the robustness of deep neural networks via stability training, 2016
Zheng, S., Song, Y., Leung, T., and Goodfellow, I · 2016
Earlier work this paper cites.
Learning deep features for discriminative localization
Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Torralba, A · 2016
Earlier work this paper cites.
An actor-critic algorithm for sequence prediction
Bahdanau, D., Brakel, P., Xu, K., Goyal, A., Lowe, R., Pineau, J., Courville, A., and Bengio, Y · 2017
Earlier work this paper cites.
Adversarial transformation networks: Learning to generate adversarial examples
Baluja, S., and Fischer, I · 2017
Earlier work this paper cites.
A convex framework for fair regression
Berk, R., Heidari, H., Jabbari, S., Joseph, M., Kearns, M., Morgenstern, J., Neel, S., and Roth, A · 2017
Earlier work this paper cites.
Data decisions and theoretical implications when adversarially learning fair representations
Beutel, A., Chen, J., Zhao, Z., and Chi, E. H · 2017
Earlier work this paper cites.
Unsupervised pixel-level domain adaptation with generative adversarial networks
Bousmalis, K., Silberman, N., Dohan, D., Erhan, D., and Krishnan, D · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Carlini, N., and Wagner, D · 2017
Earlier work this paper cites.
Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models
Chen, P.-Y., Zhang, H., Sharma, Y., Yi, J., and Hsieh, C.-J · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Houdini: Fooling deep structured prediction models
Cisse, M., Adi, Y., Neverova, N., and Keshet, J · 2017
Earlier work this paper cites.
Domain adaptation for visual applications: A comprehensive survey
Csurka, G · 2017
Earlier work this paper cites.
Real time image saliency for black box classifiers
Dabkowski, P., and Gal, Y · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F., and Kim, B · 2017
Earlier work this paper cites.
Dermatologist-level classification of skin cancer with deep neural networks
Esteva, A., Kuprel, B., Novoa, R. A., Ko, J., Swetter, S. M., Blau, H. M., and Thrun, S · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Measuring the tendency of cnns to learn surface statistical regularities
Jo, J., and Bengio, Y · 2017
Earlier work this paper cites.
An analysis of visual question answering algorithms
Kafle, K., and Kanan, C · 2017
Earlier work this paper cites.
Mitigating fooling with competitive overcomplete output layer neural networks
Kardan, N., and Stanley, K. O · 2017
Earlier work this paper cites.
Learning to discover cross-domain relations with generative adversarial networks
Kim, T., Cha, M., Kim, H., Lee, J. K., and Kim, J · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2017
Earlier work this paper cites.
Dense associative memory is robust to adversarial inputs
Krotov, D., and Hopfield, J. J · 2017
Earlier work this paper cites.
Adversarial examples in the physical world
Kurakin, A., Goodfellow, I., and Bengio, S · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollar, P · 2017
Earlier work this paper cites.
Enhanced attacks on defensively distilled deep neural networks, 2017
Liu, Y., Zhang, W., Li, S., and Yu, N · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S. M., and Lee, S.-I · 2017
Earlier work this paper cites.
Universal adversarial perturbations
Moosavi-Dezfooli, S.-M., Fawzi, A., Fawzi, O., and Frossard, P · 2017
Earlier work this paper cites.
Unified deep supervised domain adaptation and generalization
Motiian, S., Piccirilli, M., Adjeroh, D. A., and Doretto, G · 2017
Earlier work this paper cites.
Cascade adversarial machine learning regularized with a unified embedding
Na, T., Ko, J. H., and Mukhopadhyay, S · 2017
Earlier work this paper cites.
Biologically inspired protection of deep networks from adversarial attacks
Nayebi, A., and Ganguli, S · 2017
Earlier work this paper cites.
‘a learning and masking approach to secure learning, 2017
Nguyen, L., and Sinha, A · 2017
Earlier work this paper cites.
Adversarial image perturbation for privacy protection a game theory perspective
Oh, S. J., Fritz, M., and Schiele, B · 2017
Earlier work this paper cites.
Practical black-box attacks against machine learning
Papernot, N., McDaniel, P., Goodfellow, I., Jha, S., Celik, Z. B., and Swami, A · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2017
Earlier work this paper cites.
Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients, 2017
Ross, A. S., and Doshi-Velez, F · 2017
Earlier work this paper cites.
Evaluating the visualization of what a deep neural network has learned
Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and Müller, K.-R · 2017
Earlier work this paper cites.
Upset and angri: Breaking high performance image classifiers
Sarkar, S., Bansal, A., Mahbub, U., and Chellappa, R · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Shrikumar, A., Greenside, P., and Kundaje, A · 2017
Earlier work this paper cites.
Ensemble methods as a defense to adversarial perturbations against deep neural networks
Strauss, T., Hanselmann, M., Junginger, A., and Ulmer, H · 2017
Earlier work this paper cites.
Hypernetworks with statistical filtering for defending adversarial examples
Sun, Z., Ozay, M., and Okatani, T · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Distant domain transfer learning
Tan, B., Zhang, Y., Pan, S., and Yang, Q · 2017
Earlier work this paper cites.
Adversarial discriminative domain adaptation
Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T · 2017
Earlier work this paper cites.
Toward robustness against label noise in training deep discriminative neural networks, 2017
Vahdat, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Select-additive learning: Improving generalization in multimodal sentiment analysis
Wang, H., Meghawat, A., Morency, L.-P., and Xing, E. P · 2017
Earlier work this paper cites.
Adversary resistant deep neural networks with an application to malware detection
Wang, Q., Guo, W., Zhang, K., Ororbia II, A. G., Xing, X., Liu, X., and Giles, C. L · 2017
Earlier work this paper cites.
Efficient defenses against adversarial attacks
Zantedeschi, V., Nicolae, M.-I., and Rawat, A · 2017
Earlier work this paper cites.
Generating natural adversarial examples
Zhao, Z., Dua, D., and Singh, S · 2017
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Zintgraf, L. M., Cohen, T. S., Adel, T., and Welling, M · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
Threat of adversarial attacks on deep learning in computer vision: A survey
Akhtar, N., and Mian, A · 2018
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Earlier work this paper cites.
Agnostic domain generalization
Carlucci, F. M., Russo, P., Tommasi, T., and Caputo, B · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Dhurandhar, A., Chen, P.-Y., Luss, R., Tu, C.-C., Ting, P., Shanmugam, K., and Das, P · 2018
Earlier work this paper cites.
Mma training: Direct input space margin maximization through adversarial training
Ding, G. W., Sharma, Y., Lui, K. Y. C., and Huang, R · 2018
Earlier work this paper cites.
Decoupled classifiers for group-fair and efficient machine learning
Dwork, C., Immorlica, N., Kalai, A. T., and Leiserson, M · 2018
Earlier work this paper cites.
Domain generalization with domain-specific aggregation modules
D’Innocente, A., and Caputo, B · 2018
Earlier work this paper cites.
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2018
Earlier work this paper cites.
Non-discriminatory machine learning through convex fairness criteria
Goel, N., Yaghini, M., and Faltings, B · 2018
Earlier work this paper cites.
Low frequency adversarial perturbation
Guo, C., Frank, J. S., and Weinberger, K. Q · 2018
Earlier work this paper cites.
Learning universal adversarial perturbations with generative models
Hayes, J., and Danezis, G · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Earlier work this paper cites.
Kannan, H., Kurakin, A., and Goodfellow, I · 2018
Earlier work this paper cites.
Art of singular vectors and universal adversarial perturbations
Khrulkov, V., and Oseledets, I · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems, 2018
Kiritchenko, S., and Mohammad, S. M · 2018
Earlier work this paper cites.
CausalGAN: Learning causal implicit generative models with adversarial training
Kocaoglu, M., Snyder, C., Dimakis, A. G., and Vishwanath, S · 2018
Earlier work this paper cites.
Adaptive sensitive reweighting to mitigate bias in fairness-aware classification
Krasanakis, E., Spyromitros-Xioufis, E., Papadopoulos, S., and Kompatsiaris, Y · 2018
Earlier work this paper cites.
Can neural machine translation be improved with user feedback?
Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S · 2018
Earlier work this paper cites.
Improving a neural semantic parser by counterfactual learning from human bandit feedback
Lawrence, C., and Riezler, S · 2018
Earlier work this paper cites.
Domain generalization with adversarial feature learning
Li, H., Pan, S. J., Wang, S., and Kot, A. C · 2018
Earlier work this paper cites.
Deep domain generalization via conditional invariant adversarial networks
Li, Y., Tian, X., Gong, M., Liu, Y., Liu, T., Zhang, K., and Tao, D · 2018
Earlier work this paper cites.
Learning noise-invariant representations for robust speech recognition, 2018
Liang, D., Huang, Z., and Lipton, Z. C · 2018
Earlier work this paper cites.
Defense against adversarial attacks using high-level representation guided denoiser
Liao, F., Liang, M., Dong, Y., Pang, T., Zhu, J., and Hu, X · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Lipton, Z. C · 2018
Earlier work this paper cites.
Detecting and correcting for label shift with black box predictors
Lipton, Z. C., Wang, Y.-X., and Smola, A · 2018
Earlier work this paper cites.
Learning adversarially fair and transferable representations
Madras, D., Creager, E., Pitassi, T., and Zemel, R · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2018
Earlier work this paper cites.
Methods for interpreting and understanding deep neural networks
Montavon, G., Samek, W., and Müller, K.-R · 2018
Earlier work this paper cites.
Learning with complex loss functions and constraints
Narasimhan, H · 2018
Earlier work this paper cites.
The book of why: the new science of cause and effect
Pearl, J., and Mackenzie, D · 2018
Earlier work this paper cites.
Zero-shot deep domain adaptation
Peng, K.-C., Wu, Z., and Ernst, J · 2018
Earlier work this paper cites.
Deep contextualized word representations, 2018
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Earlier work this paper cites.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk, V., Das, A., and Saenko, K · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K · 2018
Earlier work this paper cites.
Generalizing across domains via cross-gradient training
Shankar, S., Piratla, V., Chakrabarti, S., Chaudhuri, S., Jyothi, P., and Sarawagi, S · 2018
Earlier work this paper cites.
Achieving fairness through adversarial learning: an application to recidivism prediction
Wadsworth, C., Vera, F., and Piech, C · 2018
Earlier work this paper cites.
Deep visual domain adaptation: A survey
Wang, M., and Deng, W · 2018
Earlier work this paper cites.
Mitigating unwanted biases with adversarial learning
Zhang, B. H., Lemoine, B., and Mitchell, M · 2018
Earlier work this paper cites.
Regularizing neural machine translation by target-bidirectional agreement, 2018
Zhang, Z., Wu, S., Liu, S., Li, M., Zhou, M., and Xu, T · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W · 2018
Earlier work this paper cites.
One-network adversarial fairness
Adel, T., Valera, I., Ghahramani, Z., and Weller, A · 2019
Earlier work this paper cites.
Learning optimal and fair decision trees for non-discriminative decision-making
Aghaei, S., Azizi, M. J., and Vayanos, P · 2019
Earlier work this paper cites.
Adversarial invariant feature learning with accuracy constraint for domain generalization
Akuzawa, K., Iwasawa, Y., and Matsuo, Y · 2019
Earlier work this paper cites.
Uncovering and mitigating algorithmic bias through learned latent structure
Amini, A., Soleimany, A. P., Schwarting, W., Bhatia, S. N., and Rus, D · 2019
Earlier work this paper cites.
Arjovsky, M., Bottou, L., Gulrajani, I., and Lopez-Paz, D · 2019
Earlier work this paper cites.
Learning to understand goal specifications by modelling reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Kohli, P., and Grefenstette, E · 2019
Earlier work this paper cites.
Better rewards yield better summaries: Learning to summarise without references
Böhm, F., Gao, Y., Meyer, C. M., Shapira, O., Dagan, I., and Gurevych, I · 2019
Earlier work this paper cites.
Approximating cnns with bag-of-local-features models works surprisingly well on imagenet, 2019
Brendel, W., and Bethge, M · 2019
Earlier work this paper cites.
Unlabeled data improves adversarial robustness
Carmon, Y., Raghunathan, A., Schmidt, L., Duchi, J. C., and Liang, P. S · 2019
Earlier work this paper cites.
Classification with fairness constraints: A meta-algorithm with provable guarantees
Celis, L. E., Huang, L., Keswani, V., and Vishnoi, N. K · 2019
Earlier work this paper cites.
Improved adversarial learning for fair classification
Celis, L. E., and Keswani, V · 2019
Earlier work this paper cites.
Neural network attributions: A causal perspective
Chattopadhyay, A., Manupriya, P., Sarkar, A., and Balasubramanian, V. N · 2019
Earlier work this paper cites.
Sign-opt: A query-efficient hard-label adversarial attack
Cheng, M., Singh, S., Chen, P., Chen, P.-Y., Liu, S., and Hsieh, C.-J · 2019
Earlier work this paper cites.
Optimization with non-differentiable constraints with applications to fairness, recall, churn, and other goals
Cotter, A., Jiang, H., Gupta, M. R., Wang, S., Narayan, T., You, S., and Sridharan, K · 2019
Earlier work this paper cites.
Adversarial training methods for network embedding, 2019
Dai, Q., Shen, X., Zhang, L., Li, Q., and Wang, D · 2019
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
Davison, J., Feldman, J., and Rush, A · 2019
Earlier work this paper cites.
Explanations can be manipulated and geometry is to blame
Dombrowski, A.-K., Alber, M., Anders, C., Ackermann, M., Müller, K.-R., and Kessel, P · 2019
Earlier work this paper cites.
Adversarial robustness as a prior for learned representations, 2019
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Learning perceptually-aligned representations via adversarial robustness
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Learning fair representations via an adversarial framework
Feng, R., Yang, Y., Lyu, Y., Tan, C., Sun, Y., and Wang, C · 2019
Earlier work this paper cites.
Counterfactual fairness in text classification through robustness, 2019
Garg, S., Perot, V., Limtiaco, N., Taly, A., Chi, E. H., and Beutel, A · 2019
Earlier work this paper cites.
Dlow: Domain flow for adaptation and generalization
Gong, R., Li, W., Chen, Y., and Gool, L. V · 2019
Earlier work this paper cites.
Counterfactual visual explanations
Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
Visual attention consistency under image transforms for multi-label image classification
Guo, H., Zheng, K., Fan, X., Yu, H., and Wang, S · 2019
Earlier work this paper cites.
Differential privacy in deep learning: An overview
Ha, T., Dang, T. K., Dang, T. T., Truong, T. A., and Nguyen, M. T · 2019
Earlier work this paper cites.
Learning from dialogue after deployment: Feed yourself, chatbot!
Hancock, B., Bordes, A., Mazare, P.-E., and Weston, J · 2019
Earlier work this paper cites.
Fooling neural network interpretations via adversarial model manipulation
Heo, J., Joo, S., and Moon, T · 2019
Earlier work this paper cites.
A Benchmark for Interpretability Methods in Deep Neural Networks
Hooker, S., Erhan, D., Kindermans, P.-J., and Kim, B · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Stable and fair classification
Huang, L., and Vishnoi, N · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog, 2019
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R · 2019
Earlier work this paper cites.
Learning not to learn: Training deep neural networks with biased data
Kim, B., Kim, H., Kim, K., Kim, S., and Kim, J · 2019
Earlier work this paper cites.
Hallucinations in neural machine translation, 2019
Lee, K., Firat, O., Agarwal, A., Fannjiang, C., and Sussillo, D · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Taking advantage of multitask learning for fair classification
Oneto, L., Doninini, M., Elders, A., and Pontil, M · 2019
Cited alongside, same era.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Explain yourself! leveraging language models for commonsense reasoning
Rajani, N. F., McCann, B., Xiong, C., and Socher, R · 2019
Cited alongside, same era.
Toward learning human-aligned cross-domain robust models by countering misaligned features
Wang, H., Huang, Z., Zhang, H., and Xing, E · 2021
Later among the works it cites.
Assessing multilingual fairness in pre-trained multimodal representations
Wang, J., Liu, Y., and Wang, X. E · 2021
Later among the works it cites.
Dynamically disentangling social bias from task-oriented representations with adversarial attack
Wang, L., Yan, Y., He, K., Wu, Y., and Xu, W · 2021
Later among the works it cites.
Identifying and mitigating spurious correlations for improving robustness in nlp models
Wang, T., Sridhar, R., Yang, D., and Wang, X · 2021
Later among the works it cites.
Causal attention for unbiased visual recognition
Wang, T., Zhou, C., Sun, Q., and Zhang, H · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sagawa, S., Koh, P. W., Hashimoto, T. B., and Liang, P · 2019
Cited alongside, same era.
Cxplain: Causal explanations for model interpretation under uncertainty
Schwab, P., and Karlen, W · 2019
Cited alongside, same era.
Cycle-consistency for robust visual question answering, 2019
Shah, M., Chen, X., Rohrbach, M., and Parikh, D · 2019
Cited alongside, same era.
The woman worked as a babysitter: On biases in language generation
Sheng, E., Chang, K.-W., Natarajan, P., and Peng, N · 2019
Cited alongside, same era.
Estimating causal effects of tone in online debates, 2019
Sridhar, D., and Getoor, L · 2019
Cited alongside, same era.
One pixel attack for fooling deep neural networks
Su, J., Vargas, D. V., and Sakurai, K · 2019
Cited alongside, same era.
Robustly disentangled causal mechanisms: Validating deep representations for interventional robustness
Suter, R., Miladinovic, D., Schölkopf, B., and Bauer, S · 2019
Cited alongside, same era.
Weakly-supervised video object grounding via causal intervention, 2021
Wang, W., Gao, J., and Xu, C · 2021
Later among the works it cites.
Contrastive counterfactual visual explanations with overdetermination, 2021
White, A., Ngan, K. H., Phelan, J., Afgeh, S. S., Ryan, K., Reyes-Aldasoro, C. C., and d’Avila Garcez, A · 2021
Later among the works it cites.
Measuring association between labels and free-text rationales
Wiegreffe, S., Marasović, A., and Smith, N. A · 2021
Later among the works it cites.
Recursively summarizing books with human feedback, 2021
Wu, J., Ouyang, L., Ziegler, D. M., Stiennon, N., Lowe, R., Leike, J., and Christiano, P · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models, 2021
Wu, T., Ribeiro, M. T., Heer, J., and Weld, D. S · 2021
Later among the works it cites.
Generative counterfactuals for neural networks via attribute-informed perturbation, 2021
Yang, F., Liu, N., Du, M., and Hu, X · 2021
Later among the works it cites.
Causalvae: Disentangled representation learning via neural structural causal models
Yang, M., Liu, F., Chen, Z., Shen, X., Hao, J., and Wang, J · 2021
Later among the works it cites.
Causal attention for vision-language tasks
Yang, X., Zhang, H., Qi, G., and Cai, J · 2021
Later among the works it cites.
Improved ood generalization via adversarial training and pre-training, 2021
Yi, M., Hou, L., Sun, J., Shang, L., Jiang, X., Liu, Q., and Ma, Z.-M · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation, 2021
Yuan, W., Neubig, G., and Liu, P · 2021
Later among the works it cites.
Counterfactual zero-shot and open-set visual recognition, 2021
Yue, Z., Wang, T., Zhang, H., Sun, Q., and Hua, X.-S · 2021
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Ravfogel, S., and Goldberg, Y · 2021
Later among the works it cites.
What if we could not see? counterfactual analysis for egocentric action anticipation
Zhang, T., Min, W., Yang, J., Liu, T., Jiang, S., and Rui, Y · 2021
Later among the works it cites.
A survey on neural network interpretability
Zhang, Y., Tio, P., Leonardis, A., and Tang, K · 2021
Later among the works it cites.
Factual probing is [mask]: Learning vs. learning to recall, 2021
Zhong, Z., Friedman, D., and Chen, D · 2021
Later among the works it cites.
Examining and combating spurious features under distribution shift, 2021
Zhou, C., Ma, X., Michel, P., and Neubig, G · 2021
Later among the works it cites.
A closer look at how fine-tuning changes bert
Zhou, Y., and Srikumar, V · 2021
Later among the works it cites.
Deformable {detr}: Deformable transformers for end-to-end object detection
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J · 2021
Later among the works it cites.
Controllable generation from pre-trained language models via inverse prompting
Zou, X., Yin, D., Zhong, Q., Ding, M., Yang, H., Yang, Z., and Tang, J · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Later among the works it cites.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A. J., Lewkowycz, A., Misra, V., Ramasesh, V. V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Later among the works it cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Later among the works it cites.
Pada: Example-based prompt learning for on-the-fly adaptation to unseen domains, 2022
Ben-David, E., Oved, N., and Reichart, R · 2022
Later among the works it cites.
Let there be a clock on the beach: Reducing object hallucination in image captioning
Biten, A. F., Gomez, L., and Karatzas, D · 2022
Later among the works it cites.
Coin: Counterfactual image generation for vqa interpretation
Boukhers, Z., Hartmann, T., and Jürjens, J · 2022
Later among the works it cites.
Learning to perform complex tasks through compositional fine-tuning of language models
Bursztyn, V., Demeter, D., Downey, D., and Birnbaum, L · 2022
Later among the works it cites.
Docogen: Domain counterfactual generation for low resource domain adaptation
Calderon, N., Ben-David, E., Feder, A., and Reichart, R · 2022
Later among the works it cites.
DoCoGen: Domain counterfactual generation for low resource domain adaptation
Calderon, N., Ben-David, E., Feder, A., and Reichart, R · 2022
Later among the works it cites.
Fairlex: A multilingual benchmark for evaluating fairness in legal text processing
Chalkidis, I., Pasini, T., Zhang, S., Tomada, L., Schwemer, S. F., and Søgaard, A · 2022
Later among the works it cites.
Adapting pretrained vision-language foundational models to medical imaging domains
Chambon, P. J. M., Bluethgen, C., Langlotz, C., and Chaudhari, A · 2022
Later among the works it cites.
Can rationalization improve robustness?
Chen, H., He, J., Narasimhan, K., and Chen, D · 2022
Later among the works it cites.
Adversarial training for improving model robustness? look at both prediction and interpretation, 2022
Chen, H., and Ji, Y · 2022
Later among the works it cites.
Improving in-context few-shot learning via self-supervised training
Chen, M., Du, J., Pasunuru, R., Mihaylov, T., Iyer, S., Stoyanov, V., and Kozareva, Z · 2022
Later among the works it cites.
Causal intervention for subject-deconfounded facial action unit recognition, 2022
Chen, Y., Chen, D., Wang, T., Wang, Y., and Liang, Y · 2022
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Later among the works it cites.
Advcodemix: Adversarial attack on code-mixed data
Das, S. D., Basak, A., Mandal, S., and Das, D · 2022
Later among the works it cites.
Robustifying sentiment classification by maximally exploiting few counterfactuals
De Raedt, M., Godin, F., Develder, C., and Demeester, T · 2022
Later among the works it cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models
Delobelle, P., Tokpo, E. K., Calders, T., and Berendt, B · 2022
Later among the works it cites.
Align-deform-subtract: An interventional framework for explaining object differences, 2022
Eastwood, C., Nanbo, L., and Williams, C. K. I · 2022
Later among the works it cites.
Krona: Parameter efficient tuning with kronecker adapter
Edalati, A., Tahaei, M., Kobyzev, I., Nia, V. P., Clark, J. J., and Rezagholizadeh, M · 2022
Later among the works it cites.
Causal inference in natural language processing: Estimation, prediction, interpretation and beyond
Feder, A., Keith, K. A., Manzoor, E., Pryzant, R., Sridhar, D., Wood-Doughty, Z., Eisenstein, J., Grimmer, J., Reichart, R., Roberts, M. E., Stewart, B. M., Veitch, V., and Yang, D · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned, 2022
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Later among the works it cites.
Inducing causal structure for interpretable neural networks
Geiger, A., Wu, Z., Lu, H., Rozner, J., Kreiss, E., Icard, T., Goodman, N., and Potts, C · 2022
Later among the works it cites.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness, 2022
Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F. A., and Brendel, W · 2022
Later among the works it cites.
Translational lung imaging analysis through disentangled representations, 2022
Gordaliza, P. M., Vaquero, J. J., and Muñoz-Barrutia, A · 2022
Later among the works it cites.
Auto-debias: Debiasing masked language models with automated biased prompts
Guo, Y., Yang, Y., and Abbasi, A · 2022
Later among the works it cites.
Mitigating gender bias in distilled language models via counterfactual role reversal
Gupta, U., Dhamala, J., Kumar, V., Verma, A., Pruksachatkun, Y., Krishna, S., Gupta, R., Chang, K.-W., Steeg, G. V., and Galstyan, A · 2022
Later among the works it cites.
Trustworthy artificial intelligence in medical imaging
Hasani, N., Morris, M., Rhamim, A., Summers, R., Jones, E., Siegel, E., and Saboury, B · 2022
Later among the works it cites.
On distribution shift in learning-based bug detectors
He, J., Beurer-Kellner, L., and Vechev, M · 2022
Later among the works it cites.
MABEL: Attenuating gender bias using textual entailment data
He, J., Xia, M., Fellbaum, C., and Chen, D · 2022
Later among the works it cites.
CPL: Counterfactual prompt learning for vision and language models
He, X., Yang, D., Feng, W., Fu, T.-J., Akula, A., Jampani, V., Narayana, P., Basu, S., Wang, W. Y., and Wang, X · 2022
Later among the works it cites.
Attack-less adversarial training for a robust adversarial defense
Ho, J., Lee, B.-G., and Kang, D.-K · 2022
Later among the works it cites.
Causal information bottleneck boosts adversarial robustness of deep neural network
Hua, H., Yan, J., Fang, X., Huang, W., Yin, H., and Ge, W · 2022
Later among the works it cites.
The two dimensions of worst-case training and their integrated effect for out-of-domain generalization
Huang, Z., Wang, H., Huang, D., Lee, Y. J., and Xing, E. P · 2022
Later among the works it cites.
Atlas: Few-shot learning with retrieval augmented language models, 2022
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E · 2022
Later among the works it cites.
Training calibration-based counterfactual explainers for deep learning models in medical image analysis
J. Thiagarajan, J., Thopalli, K., Rajan, D., and Turaga, P · 2022
Later among the works it cites.
Prompt-based distribution alignment for domain generalization in text classification
Jia, C., and Zhang, Y · 2022
Later among the works it cites.
Embedding hallucination for few-shot language fine-tuning
Jian, Y., Gao, C., and Vosoughi, S · 2022
Later among the works it cites.
ROSE: Robust selective fine-tuning for pre-trained language models
Jiang, L., Zhou, H., Lin, Y., Li, P., Zhou, J., and Jiang, R · 2022
Later among the works it cites.
Are all spurious features in natural language alike? an analysis through a causal lens
Joshi, N., Pan, X., and He, H · 2022
Later among the works it cites.
De-biasing facial detection system using vae
Kandge, V. V., Kandge, S. V., Kumbharkar, K., Pattanshetti, P., et al · 2022
Later among the works it cites.
Unmasking the mask – evaluating social biases in masked language models
Kaneko, M., and Bollegala, D · 2022
Later among the works it cites.
Auto-encoding variational bayes, 2022
Kingma, D. P., and Welling, M · 2022
Later among the works it cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Later among the works it cites.
Efficient counterfactual debiasing for visual question answering
Kolling, C., More, M., Gavenski, N., Pooch, E., Parraga, O., and Barros, R. C · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2022
Later among the works it cites.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P · 2022
Later among the works it cites.
Language generation models can cause harm: So what can we do about it? an actionable survey
Kumar, S., Balachandran, V., Njoo, L., Anastasopoulos, A., and Tsvetkov, Y · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Later among the works it cites.
SPE: Symmetrical prompt enhancement for fact probing
Li, Y., Che, T., Wang, Y., Jiang, Z., Xiong, C., and Chaturvedi, S · 2022
Later among the works it cites.
Gendered mental health stigma in masked language models
Lin, I., Njoo, L., Field, A., Sharma, A., Reinecke, K., Althoff, T., and Tsvetkov, Y · 2022
Later among the works it cites.
What makes good in-context examples for GPT-3?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., and Chen, W · 2022
Later among the works it cites.
Domain confused contrastive learning for unsupervised domain adaptation
Long, Q., Luo, T., Wang, W., and Pan, S. J · 2022
Later among the works it cites.
Low-resource interactive active labeling for fine-tuning language models
Maekawa, S., Zhang, D., Kim, H., Rahman, S., and Hruschka, E · 2022
Later among the works it cites.
Cascading biases: Investigating the effect of heuristic annotation strategies on data and models
Malaviya, C., Bhatia, S., and Yatskar, M · 2022
Later among the works it cites.
Understanding stereotypes in language models: Towards robust measurement and zero-shot debiasing
Mattern, J., Jin, Z., Sachan, M., Mihalcea, R., and Schölkopf, B · 2022
Later among the works it cites.
Ganterfactual—counterfactual explanations for medical non-experts using generative adversarial learning
Mertes, S., Huber, T., Weitz, K., Heimerl, A., and Andr’e, E · 2022
Later among the works it cites.
Multi-modal understanding and generation for medical images and text via vision-language pre-training
Moon, J. H., Lee, H., Shin, W., Kim, Y.-H., and Choi, E · 2022
Later among the works it cites.
Semi-supervised domain adaptation with cyclegan guided by a downstream task loss
Mütze, A., Rottmann, M., and Gottschalk, H · 2022
Later among the works it cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2022
Later among the works it cites.
Counterfactual data augmentation via perspective transition for open-domain dialogues
Ou, J., Zhang, J., Feng, Y., and Zhou, J · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
Data augmentation for fairness-aware machine learning: Preventing algorithmic bias in law enforcement systems
Pastaltzidis, I., Dimitriou, N., Quezada-Tavarez, K., Aidinlis, S., Marquenie, T., Gurzawska, A., and Tzovaras, D · 2022
Later among the works it cites.
Identifiability of sparse causal effects using instrumental variables, 2022
Pfister, N., and Peters, J · 2022
Later among the works it cites.
Frameworks and results in distributionally robust optimization
Rahimian, H., and Mehrotra, S · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents, 2022
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
Multiple adversarial domains adaptation approach for mitigating adversarial attacks effects
Rasheed, B., Khan, A., Ahmad, M., Mazzara, M., Kazmi, S., et al · 2022
Later among the works it cites.
Causal scene bert: Improving object detection by searching for challenging groups of data, 02 2022
Resnick, C., Litany, O., Kar, A., Kreis, K., Lucas, J., Cho, K., and Fidler, S · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Learning to retrieve prompts for in-context learning
Rubin, O., Herzig, J., and Berant, J · 2022
Later among the works it cites.
Exploiting independent instruments: Identification and distribution generalization
Saengkyongam, S., Henckel, L., Pfister, N., and Peters, J · 2022
Later among the works it cites.
Just fine-tune twice: Selective differential privacy for large language models
Shi, W., Shea, R., Chen, S., Zhang, C., Jia, R., and Yu, Z · 2022
Later among the works it cites.
Salient imagenet: How to discover spurious features in deep learning?, 2022
Singla, S., and Feizi, S · 2022
Later among the works it cites.
An information-theoretic approach to prompt engineering without ground truth labels
Sorensen, T., Robinson, J., Rytting, C., Shaw, A., Rogers, K., Delorey, A., Khalil, M., Fulda, N., and Wingate, D · 2022
Later among the works it cites.
A comparison of strategies for source-free domain adaptation
Su, X., Zhao, Y., and Bethard, S · 2022
Later among the works it cites.
BERTScore is unfair: On social bias in language model-based metrics for text generation
Sun, T., He, J., Qiu, X., and Huang, X · 2022
Later among the works it cites.
Lst: Ladder side-tuning for parameter and memory efficient transfer learning
Sung, Y.-L., Cho, J., and Bansal, M · 2022
Later among the works it cites.
Masktune: Mitigating spurious correlations by forcing to explore, 2022
Taghanaki, S. A., Khani, A., Khani, F., Gholami, A., Tran, L., Mahdavi-Amiri, A., and Hamarneh, G · 2022
Later among the works it cites.
Identifying causal effects using instrumental time series: Nuisance iv and correcting for the past, 2022
Thams, N., Søndergaard, R., Weichwald, S., and Peters, J · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Later among the works it cites.
Neuron coverage-guided domain generalization
Tian, C. X., Li, H., Xie, X., Liu, Y., and Wang, S · 2022
Later among the works it cites.
Mitigating spurious correlation in natural language understanding with counterfactual inference
Udomcharoenchaikit, C., Ponwitayarat, W., Payoungkhamdee, P., Masuk, K., Buaphet, W., Chuangsuwanich, E., and Nutanong, S · 2022
Later among the works it cites.
Efficient fine-tuning of bert models on the edge
Vucetic, D., Tayaranian, M., Ziaeefard, M., Clark, J. J., Meyer, B. H., and Gross, W. J · 2022
Later among the works it cites.
Iteratively prompt pre-trained language models for chain of thought
Wang, B., Deng, X., and Sun, H · 2022
Later among the works it cites.
Toward learning robust and invariant representations with alignment regularization and data augmentation
Wang, H., Huang, Z., Wu, X., and Xing, E · 2022
Later among the works it cites.
Understanding gradual domain adaptation: Improved analysis, optimal path and beyond
Wang, H., Li, B., and Zhao, H · 2022
Later among the works it cites.
Generalizing to unseen domains: A survey on domain generalization
Wang, J., Lan, C., Liu, C., Ouyang, Y., Qin, T., Lu, W., Chen, Y., Zeng, W., and Yu, P · 2022
Later among the works it cites.
Learning robust representations for continual relation extraction via adversarial class augmentation
Wang, P., Song, Y., Liu, T., Lin, B., Cao, Y., Li, S., and Sui, Z · 2022
Later among the works it cites.
Causal intervention improves implicit sentiment analysis
Wang, S., Zhou, J., Sun, C., Ye, J., Gui, T., Zhang, Q., and Huang, X · 2022
Later among the works it cites.
On the convergence and robustness of adversarial training, 2022
Wang, Y., Ma, X., Bailey, J., Yi, J., Zhou, B., and Gu, Q · 2022
Later among the works it cites.
Adamix: Mixture-of-adapter for parameter-efficient tuning of large language models
Wang, Y., Mukherjee, S., Liu, X., Gao, J., Awadallah, A. H., and Gao, J · 2022
Later among the works it cites.
Fairness-aware adversarial perturbation towards bias mitigation for deployed deep models
Wang, Z., Dong, X., Xue, H., Zhang, Z., Chiu, W., Wei, T., and Ren, K · 2022
Later among the works it cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., brian ichter, Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Later among the works it cites.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Lopes, R. G., Hajishirzi, H., Farhadi, A., Namkoong, H., et al · 2022
Later among the works it cites.
Learning instrumental variable from data fusion for treatment effect estimation, 2022
Wu, A., Kuang, K., Xiong, R., Zhu, M., Liu, Y., Li, B., Liu, F., Wang, Z., and Wu, F · 2022
Later among the works it cites.
Adversarial soft prompt tuning for cross-domain sentiment analysis
Wu, H., and Shi, X · 2022
Later among the works it cites.
Generating data to mitigate spurious correlations in natural language inference datasets
Wu, Y., Gardner, M., Stenetorp, P., and Dasigi, P · 2022
Later among the works it cites.
Wu, Z., Xu, H., Fang, J., and Gao, K · 2022
Later among the works it cites.
Does your model classify entities reasonably? diagnosing and mitigating spurious correlations in entity typing
Xu, N., Wang, F., Li, B., Dong, M., and Chen, M · 2022
Later among the works it cites.
Improving stability of fine-tuning pretrained language models via component-wise gradient norm clipping
Yang, C., and Ma, X · 2022
Later among the works it cites.
Gram: Fast fine-tuning of pre-trained language models for content-based collaborative filtering
Yang, Y., Kim, K. S., Kim, M., and Park, J · 2022
Later among the works it cites.
Interventional training for out-of-distribution natural language understanding
Yu, S., Jiang, J., Zhang, H., Niu, Y., Sun, Q., and Bing, L · 2022
Later among the works it cites.
Actune: Uncertainty-based active self-training for active fine-tuning of pretrained language models
Yu, Y., Kong, L., Zhang, J., Zhang, R., and Zhang, C · 2022
Later among the works it cites.
COCO-DR: Combating the distribution shift in zero-shot dense retrieval with contrastive and distributionally robust learning
Yu, Y., Xiong, C., Sun, S., Zhang, C., and Overwijk, A · 2022
Later among the works it cites.
QA domain adaptation using hidden space augmentation and self-supervised contrastive adaptation
Yue, Z., Zeng, H., Kratzwald, B., Feuerriegel, S., and Wang, D · 2022
Later among the works it cites.
Zaharia, G.-E., Smădu, R.-A., Cercel, D.-C., and Dascalu, M · 2022
Later among the works it cites.
Pearl causal hierarchy on image data: Intricacies & challenges
Zečević, M., Willig, M., Dhami, D. S., and Kersting, K · 2022
Later among the works it cites.
Zhang, H., Liang, H., Zhang, Y., Zhan, L., Wu, X.-M., Lu, X., and Lam, A · 2022
Later among the works it cites.
RoChBert: Towards robust BERT fine-tuning for Chinese
Zhang, Z., Li, J., Shi, N., Yuan, B., Liu, X., Zhang, R., Xue, H., Sun, D., and Zhang, C · 2022
Later among the works it cites.
Fine-mixing: Mitigating backdoors in fine-tuned language models
Zhang, Z., Lyu, L., Ma, X., Wang, C., and Sun, X · 2022
Later among the works it cites.
Certified robustness against natural language attacks by causal intervention
Zhao, H., Ma, C., Dong, X., Luu, A. T., Deng, Z.-H., and Zhang, H · 2022
Later among the works it cites.
Gpt-3-driven pedagogical agents to train children’s curious question-asking skills
Abdelghani, R., Wang, Y.-H., Yuan, X., Wang, T., Lucas, P., Sauzéon, H., and Oudeyer, P.-Y · 2023
Closest in time.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Closest in time.
What learning algorithm is in-context learning? investigations with linear models, 2023
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Closest in time.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning, 2023
Allen-Zhu, Z., and Li, Y · 2023
Closest in time.
Fine-tuning deteriorates general textual out-of-distribution detection by distorting task-agnostic features
Chen, S., Yang, W., Bi, X., and Sun, X · 2023
Closest in time.
Large language models are few(1)-shot table reasoners, 2023
Chen, W · 2023
Closest in time.
How robust is gpt-3.5 to predecessors? a comprehensive study on language understanding tasks, 2023
Chen, X., Ye, J., Zu, C., Xu, N., Zheng, R., Peng, M., Zhou, J., Gui, T., Zhang, Q., and Huang, X · 2023
Closest in time.
On the relation between sensitivity and accuracy in in-context learning, 2023
Chen, Y., Zhao, C., Yu, Z., McKeown, K., and He, H · 2023
Closest in time.
Chatlaw: Open-source legal large language model with integrated external knowledge bases, 2023
Cui, J., Li, Z., Yan, Y., Chen, B., and Yuan, L · 2023
Closest in time.
A too-good-to-be-true prior to reduce shortcut reliance
Dagaev, N., Roads, B. D., Luo, X., Barry, D. N., Patil, K. R., and Love, B. C · 2023
Closest in time.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers, 2023
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F · 2023
Closest in time.
Evaluation of gpt-3.5 and gpt-4 for supporting real-world information needs in healthcare delivery, 2023
Dash, D., Thapa, R., Banda, J. M., Swaminathan, A., Cheatham, M., Kashyap, M., Kotecha, N., Chen, J. H., Gombar, S., Downing, L., Pedreira, R., Goh, E., Arnaout, A., Morris, G. K., Magon, H., Lungren, M. P., Horvitz, E., and Shah, N. H · 2023
Closest in time.
Cross-domain image captioning with discriminative finetuning, 2023
Dessì, R., Bevilacqua, M., Gualdoni, E., Rakotonirina, N. C., Franzon, F., and Baroni, M · 2023
Closest in time.
Raft: Reward ranked finetuning for generative foundation model alignment, 2023
Dong, H., Xiong, W., Goyal, D., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Closest in time.
A survey on in-context learning, 2023
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., Li, L., and Sui, Z · 2023
Closest in time.
Semi-supervised specific emitter identification method using metric-adversarial training
Fu, X., Peng, Y., Liu, Y., Lin, Y., Gui, G., Gacanin, H., and Adachi, F · 2023
Closest in time.
Improving zero-shot generalization and robustness of multi-modal models
Ge, Y., Ren, J., Gallagher, A., Wang, Y., Yang, M.-H., Adam, H., Itti, L., Lakshminarayanan, B., and Zhao, J · 2023
Closest in time.
Finetune like you pretrain: Improved finetuning of zero-shot vision models
Goyal, S., Kumar, A., Garg, S., Kolter, Z., and Raghunathan, A · 2023
Closest in time.
The false promise of imitating proprietary llms
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D · 2023
Closest in time.
Rationalization for explainable nlp: A survey
Gurrapu, S., Kulkarni, A., Huang, L., Lourentzou, I., Freeman, L., and Batarseh, F. A · 2023
Closest in time.
Preserving pre-trained features helps calibrate fine-tuned language models
He, G., Chen, J., and Zhu, J · 2023
Closest in time.
Towards compositional adversarial robustness: Generalizing adversarial training to composite semantic perturbations
Hsiung, L., Tsai, Y.-Y., Chen, P.-Y., and Ho, T.-Y · 2023
Closest in time.
Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Hu, Z., Lan, Y., Wang, L., Xu, W., Lim, E.-P., Lee, R. K.-W., Bing, L., and Poria, S · 2023
Closest in time.
Towards reasoning in large language models: A survey, 2023
Huang, J., and Chang, K. C.-C · 2023
Closest in time.
How to lie with statistics
Huff, D · 2023
Closest in time.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Closest in time.
Randomized adversarial training via taylor expansion
Jin, G., Yi, X., Wu, D., Mu, R., and Huang, X · 2023
Closest in time.
Chatgpt for good? on opportunities and challenges of large language models for education
Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., Stadler, M., Weller, J., Kuhn, J., and Kasneci, G · 2023
Closest in time.
Instrumental variable estimation of average partial causal effects
Kawakami, Y., Kuroki, M., and Tian, J · 2023
Closest in time.
Causal reasoning and large language models: Opening a new frontier for causality
Kıcıman, E., Ness, R., Sharma, A., and Tan, C · 2023
Closest in time.
Demystifying causal features on adversarial examples and causal inoculation for robust network by adversarial instrumental variable regression, 2023
Kim, J., Lee, B.-K., and Ro, Y. M · 2023
Closest in time.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Closest in time.
Meltr: Meta loss transformer for learning to fine-tune video foundation models
Ko, D., Choi, J., Choi, H. K., On, K.-W., Roh, B., and Kim, H. J · 2023
Closest in time.
Surgical fine-tuning improves adaptation to distribution shifts
Lee, Y., Chen, A. S., Tajwar, F., Kumar, A., Yao, H., Liang, P., and Finn, C · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression, 2023
Li, S., Song, Z., Xia, Y., Yu, T., and Zhou, T · 2023
Closest in time.
Sibling-attack: Rethinking transferable adversarial attacks against face recognition
Li, Z., Yin, B., Yao, T., Guo, J., Ding, S., Chen, S., and Liu, C · 2023
Closest in time.
Beyond one-model-fits-all: A survey of domain specialization for large language models
Ling, C., Zhao, X., Lu, J., Deng, C., Zheng, C., Wang, J., Chowdhury, T., Li, Y., Cui, H., Zhao, T., et al · 2023
Closest in time.
Webglm: Towards an efficient web-enhanced question answering system with human preferences, 2023
Liu, X., Lai, H., Yu, H., Xu, Y., Zeng, A., Du, Z., Zhang, P., Dong, Y., and Tang, J · 2023
Closest in time.
Twins: A fine-tuning framework for improved transferability of adversarial robustness and generalization, 2023
Liu, Z., Xu, Y., Ji, X., and Chan, A. B · 2023
Closest in time.
Doubly right object recognition: A why prompt for visual rationales
Mao, C., Teotia, R., Sundar, A., Menon, S., Yang, J., Wang, X., and Vondrick, C · 2023
Closest in time.
Segment anything model for medical image analysis: an experimental study, 2023
Mazurowski, M. A., Dong, H., Gu, H., Yang, J., Konz, N., and Zhang, Y · 2023
Closest in time.
Locating and editing factual associations in gpt, 2023
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2023
Closest in time.
Instrumental processes using integrated covariances, 2023
Mogensen, S. W · 2023
Closest in time.
Foundation models for generalist medical artificial intelligence
Moor, M., Banerjee, O., Shakeri, Z., Krumholz, H., Leskovec, J., Topol, E., and Rajpurkar, P · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Dinov2: Learning robust visual features without supervision, 2023
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P · 2023
Closest in time.
Medical image understanding with pretrained vision language models: A comprehensive study, 2023
Qin, Z., Yi, H., Lao, Q., and Li, K · 2023
Closest in time.
In-context retrieval-augmented language models, 2023
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Closest in time.
Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization
Ramamurthy, R., Ammanabrolu, P., Brantley, K., Hessel, J., Sifa, R., Bauckhage, C., Hajishirzi, H., and Choi, Y · 2023
Closest in time.
Crossing the gap: Domain generalization for image captioning
Ren, Y., Mao, Z., Fang, S., Lu, Y., He, T., Du, H., Zhang, Y., and Ouyang, W · 2023
Closest in time.
The programmer’s assistant: Conversational interaction with a large language model for software development
Ross, S. I., Martinez, F., Houde, S., Muller, M., and Weisz, J. D · 2023
Closest in time.
Adversarial training improves model interpretability in single-cell rna-seq analysis
Sadria, M., Layton, A., and Bader, G · 2023
Closest in time.
Chatgpt utility in healthcare education, research, and practice: systematic review on the promising perspectives and valid concerns
Sallam, M · 2023
Closest in time.
ChatGPT and other large language models are double-edged swords
Shen, Y., Heacock, L., Elias, J., Hentel, K. D., Reig, B., Shih, G., and Moy, L · 2023
Closest in time.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Babiker, A., Schärli, N., Chowdhery, A., Mansfield, P., Demner-Fushman, D., Agüera Y Arcas, B., Webster, D., Corrado, G. S., Matias, Y., Chou, K., Gottweis, J., Tomasev, N., Liu, Y., Rajkomar, A., Barral, J., Semturs, C., Karthikesalingam, A., and Natarajan, V · 2023
Closest in time.
Offline RL for natural language generation with implicit language q learning
Snell, C. V., Kostrikov, I., Su, Y., Yang, S., and Levine, S · 2023
Closest in time.
ChatGPT, bard, and large language models for biomedical research: Opportunities and pitfalls
Thapa, S., and Adhikari, S · 2023
Closest in time.
Trainable projected gradient method for robust fine-tuning, 2023
Tian, J., Dai, X., Ma, C.-Y., He, Z., Liu, Y.-C., and Kira, Z · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Closest in time.
When and how to fool explainable models (and humans) with adversarial examples, 2023
Vadillo, J., Santana, R., and Lozano, J. A · 2023
Closest in time.
Transformers learn in-context by gradient descent, 2023
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Closest in time.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
Wang, J., Hu, X., Hou, W., Chen, H., Zheng, R., Wang, Y., Yang, L., Huang, H., Ye, W., Geng, X., et al · 2023
Closest in time.
On the robustness of chatGPT: An adversarial and out-of-distribution perspective
Wang, J., HU, X., Hou, W., Chen, H., Zheng, R., Wang, Y., Yang, L., Ye, W., Huang, H., Geng, X., Jiao, B., Zhang, Y., and Xie, X · 2023
Closest in time.
What do neural networks learn in image classification? a frequency shortcut perspective, 2023
Wang, S., Veldhuis, R., Brune, C., and Strisciuglio, N · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q. V., Chi, E. H., Narang, S., Chowdhery, A., and Zhou, D · 2023
Closest in time.
Ad-aug: Adversarial data augmentation for counterfactual recommendation
Wang, Y., Qin, Y., Han, Y., Yin, M., Zhou, J., Yang, H., and Zhang, M · 2023
Closest in time.
Cfa: Class-wise calibrated fair adversarial training
Wei, Z., Wang, Y., Guo, Y., and Wang, Y · 2023
Closest in time.
Black-box sparse adversarial attack via multi-objective optimisation
Williams, P. N., and Li, K · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Wu, S., Irsoy, O., Lu, S., Dabravolski, V., Dredze, M., Gehrmann, S., Kambadur, P., Rosenberg, D., and Mann, G · 2023
Closest in time.
Masked images are counterfactual samples for robust fine-tuning, 2023
Xiao, Y., Tang, Z., Wei, P., Liu, C., and Lin, L · 2023
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation, 2023
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., and Dong, Y · 2023
Closest in time.
An adversarial training framework for mitigating algorithmic biases in clinical machine learning
Yang, J., Soltan, A. A., Eyre, D. W., Yang, Y., and Clifton, D. A · 2023
Closest in time.
Towards interpretable mental health analysis with chatgpt, 2023
Yang, K., Ji, S., Zhang, T., Xie, Q., Kuang, Z., and Ananiadou, S · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears, 2023
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Closest in time.
Transferable adversarial attacks on vision transformers with token gradient regularization
Zhang, J., Huang, Y., Wu, W., and Lyu, M. R · 2023
Closest in time.
Slic-hf: Sequence likelihood calibration with human feedback, 2023
Zhao, Y., Joshi, R., Liu, T., Khalman, M., Saleh, M., and Liu, P. J · 2023
Closest in time.
Lima: Less is more for alignment, 2023
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., and Levy, O · 2023
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q. V., and Chi, E. H · 2023
Closest in time.
Ligaa: Generative adversarial attack method based on low-frequency information
Zhu, H., Zhu, Y., Zheng, H., Ren, Y., and Jiang, W · 2023
Closest in time.
A pilot study of query-free adversarial attack against stable diffusion
Zhuang, H., Zhang, Y., and Liu, S · 2023
Closest in time.
Segment everything everywhere all at once, 2023
Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Gao, J., and Lee, Y. J · 2023
Closest in time.
Does distributionally robust supervised learning give robust classifiers?
Hu, W., Niu, G., Sato, I., and Sugiyama, M · 2037
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2057
Closest in time.
Toward learning human-aligned cross-domain robust models by countering misaligned features
Wang, H., Huang, Z., Zhang, H., Lee, Y. J., and Xing, E. P · 2084
Closest in time.