Vision-language pretrained models have substantially improved text-to-image retrieval by aligning visual and textual information in a shared semantic space. However, these systems may produce unfair outcomes across demographic groups while their vulnerability to adversarial manipulation remains largely unexplored. This work studies fairness in text-to-image retrieval under both natural and adversarial settings. We address three limitations of prior research: its focus on neutral queries that do not explicitly mention sensitive attributes; its limited attention to demographic imbalance in the retrieval database; and its assumption that unfairness arises naturally from data rather than through malicious intervention. We propose a unified framework for evaluating fairness in text-to-image retrieval. In the natural setting, we analyze how unfairness depends on both the pretrained model and the demographic composition of the retrieval database. In the adversarial setting, we define a realistic threat model in which an attacker injects a limited number of malicious images into the database. We introduce FairLess, a poisoning attack designed to amplify demographic bias while preserving semantic relevance, and evaluate a mitigation strategy to reduce its impact. We also extend fairness evaluation to attribute-labeled queries, where the target image is associated with a specific demographic group, and introduce metrics that jointly capture retrieval accuracy and fairness. Experiments on two demographically annotated versions of MSCOCO and using CLIP-based models, fairness-mitigated variants, and BLIP-2, show that unfairness stems from both model bias and database polarization. Adversarial manipulation further amplifies demographic disparities, highlighting the need for robustness-aware fairness evaluation and mitigation in text-to-image retrieval systems.

Fairness in vision-language text-to-image retrieval: Natural biases and adversarial threats

Roli, Fabio;Biggio, Battista;
2026-01-01

Abstract

Vision-language pretrained models have substantially improved text-to-image retrieval by aligning visual and textual information in a shared semantic space. However, these systems may produce unfair outcomes across demographic groups while their vulnerability to adversarial manipulation remains largely unexplored. This work studies fairness in text-to-image retrieval under both natural and adversarial settings. We address three limitations of prior research: its focus on neutral queries that do not explicitly mention sensitive attributes; its limited attention to demographic imbalance in the retrieval database; and its assumption that unfairness arises naturally from data rather than through malicious intervention. We propose a unified framework for evaluating fairness in text-to-image retrieval. In the natural setting, we analyze how unfairness depends on both the pretrained model and the demographic composition of the retrieval database. In the adversarial setting, we define a realistic threat model in which an attacker injects a limited number of malicious images into the database. We introduce FairLess, a poisoning attack designed to amplify demographic bias while preserving semantic relevance, and evaluate a mitigation strategy to reduce its impact. We also extend fairness evaluation to attribute-labeled queries, where the target image is associated with a specific demographic group, and introduce metrics that jointly capture retrieval accuracy and fairness. Experiments on two demographically annotated versions of MSCOCO and using CLIP-based models, fairness-mitigated variants, and BLIP-2, show that unfairness stems from both model bias and database polarization. Adversarial manipulation further amplifies demographic disparities, highlighting the need for robustness-aware fairness evaluation and mitigation in text-to-image retrieval systems.
2026
Vision-language pretrained models; Text-to-image retrieval systems; Fairness; Natural setting; Adversarial setting; Metrics; White-box attack; Mitigation
File in questo prodotto:
File Dimensione Formato  
1-s2.0-S0045790626004556-main.pdf

accesso aperto

Tipologia: versione editoriale (VoR)
Dimensione 3.59 MB
Formato Adobe PDF
3.59 MB Adobe PDF Visualizza/Apri

I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11584/488626
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact