We present IndicFairFace, a novel and balanced face dataset comprising 14,400 images representing geographical diversity of India.
Images were sourced ethically from Wikimedia Commons and open-license web repositories and uniformly balanced across states and gender.
Using IndicFairFace, we quantify intra-national geographical bias in prominent CLIP-based VLMs and reduce it using post-hoc Iterative Nullspace Projection debiasing approach.
We also show that the adopted debiasing approach does not adversely impact the existing embedding space as the average drop in retrieval accuracy on benchmark datasets is less than 1.5 percent.
Our work establishes IndicFairFace as the first benchmark to study geographical bias in VLMs for the Indian context.