NTH

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

AuthorsKatherine Van Koevering, Anjalie Field

August 23, 2026 2 min read
Watch on YouTube
The one-line take

The study finds that everyday linguistic styles associated with women can cause LLMs to produce shorter and less formal workplace writing, revealing a subtle but consequential form of bias.

Key results

427
Evaluated prompts

Professional-writing prompts sampled from WildChat-4.8M.

4
Evaluated models

GPT-4, Gemma 2 27b, Mistral 7b, and Meta Llama 3.1 3B Instruct.

95.0%
Task-preservation validation

Rewritten prompt pairs judged to preserve the original task.

0.988
Linguistic-condition probe accuracy

Peak Llama-3.2-3B-Instruct accuracy at transformer layer 5.

0.717
Name-gender probe accuracy

Peak accuracy for decoding explicit sign-off name gender.

6.574
Maximum patching KL divergence

Mean output-distribution shift when patching layer 0 activations.

What the paper found

This study tests whether LLMs respond differently to users based on gender-associated linguistic style rather than explicit identity. Using 427 professional-writing prompts from WildChat-4.8M, researchers created matched versions with hedges, tag questions, collective references, and expressive adjectives—women-associated linguistic features—or more direct, individual, and neutral wording. Across four models—OpenAI’s GPT-4, Gemma 2 27b, Mistral 7b, and Meta Llama 3.1 3B Instruct—the women-associated versions generally produced responses that were more readable but less lexically sophisticated, lower-grade-level, and less formal, especially for emails and job applications; politeness and clout did not significantly change. Human validation judged 95.0% of 60 rewritten pairs to preserve the original task. Regression and mediation analyses showed these effects were not explained by simple prompt-style mirroring or feature carry-over. Explicit gendered sign-off names produced no meaningful behavioral effect, while mechanistic analysis of Llama-3.2-3B-Instruct decoded linguistic condition at 0.988 accuracy, compared with 0.717 for name gender. The linguistic signal emerged by layer 1, peaked at layer 5, and activation patching found the largest causal shift at layer 0, with mean KL divergence of 6.574. Steering layer-5 activations could increase hedging, but stronger interventions degraded coherence, indicating entanglement with other generation features. The findings suggest that mitigation focused on names or pronouns misses implicit register bias embedded early in transformer computation.

Original abstract

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis