It's How You Ask: Gender-Associated Linguistic Bias in LLMs
AuthorsKatherine Van Koevering, Anjalie Field
Resources
The study finds that everyday linguistic styles associated with women can cause LLMs to produce shorter and less formal workplace writing, revealing a subtle but consequential form of bias.
Key results
Professional-writing prompts sampled from WildChat-4.8M.
GPT-4, Gemma 2 27b, Mistral 7b, and Meta Llama 3.1 3B Instruct.
Rewritten prompt pairs judged to preserve the original task.
Peak Llama-3.2-3B-Instruct accuracy at transformer layer 5.
Peak accuracy for decoding explicit sign-off name gender.
Mean output-distribution shift when patching layer 0 activations.
What the paper found
This study tests whether LLMs respond differently to users based on gender-associated linguistic style rather than explicit identity. Using 427 professional-writing prompts from WildChat-4.8M, researchers created matched versions with hedges, tag questions, collective references, and expressive adjectives—women-associated linguistic features—or more direct, individual, and neutral wording. Across four models—OpenAI’s GPT-4, Gemma 2 27b, Mistral 7b, and Meta Llama 3.1 3B Instruct—the women-associated versions generally produced responses that were more readable but less lexically sophisticated, lower-grade-level, and less formal, especially for emails and job applications; politeness and clout did not significantly change. Human validation judged 95.0% of 60 rewritten pairs to preserve the original task. Regression and mediation analyses showed these effects were not explained by simple prompt-style mirroring or feature carry-over. Explicit gendered sign-off names produced no meaningful behavioral effect, while mechanistic analysis of Llama-3.2-3B-Instruct decoded linguistic condition at 0.988 accuracy, compared with 0.717 for name gender. The linguistic signal emerged by layer 1, peaked at layer 5, and activation patching found the largest causal shift at layer 0, with mean KL divergence of 6.574. Steering layer-5 activations could increase hedging, but stronger interventions degraded coherence, indicating entanglement with other generation features. The findings suggest that mitigation focused on names or pronouns misses implicit register bias embedded early in transformer computation.
Original abstract
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.