Large Language Models (LLMs) excel in Natural Language Processing (NLP) tasks, but they often propagate biases embedded in their training data, which is potentially impactful in sensitive domains like healthcare. While existing benchmarks evaluate biases related to individual social determinants of health (SDoH) such as gender or ethnicity, they often overlook interactions between these factors and lack context-specific assessments. This study investigates bias in LLMs by probing the relationships between gender and other SDoH in French patient records. Through a series of experiments, we found that embedded stereotypes can be probed using SDoH input and that LLMs rely on embedded stereotypes to make gendered decisions, suggesting that evaluating interactions among SDoH factors could usefully complement existing approaches to assessing LLM performance and bias.
Alex: Welcome to another episode of ResearchPod.
Sam: Today we're discussing a study from researchers at Nantes Université and a French hospital. It looks at whether large computer programs trained on vast amounts of text—ones that can chat like humans and help with tasks like reading medical notes—carry hidden assumptions about gender when they process patient information.
Alex: So this is about AI in healthcare picking up stereotypes from everyday details in records?
Sam: Yes. The central puzzle is this: even when you tell the AI a patient's gender clearly, it sometimes ignores that and suggests diagnoses tied to stereotypes—like proposing menstrual problems for a man with stomach pain—because clues from the patient's job or habits trigger built-in biases.
Alex: Right, so the problem isn't just the medical facts, but these non-medical life details steering the AI wrong?
Sam: Exactly. Those life details are called social determinants of health—things like a person's job, marital status, or habits such as smoking, which shape health but aren't diseases themselves. The study suggests these can reveal how deeply gender stereotypes are baked into the AI.
Alex: And they used real patient notes from France to check this?
Sam: They did. Starting with anonymized notes from a university hospital, the researchers stripped out direct gender words to create neutral inputs—like writing "nurse (male or female form)" instead of just one version. This setup lets them see if the AI still "guesses" gender from those life details alone, mimicking how humans might stereotype.
Alex: Okay, so if the AI leans one way or another without clear clues, that points to a stereotype?
Sam: That's the idea. Deviations from a neutral guess show reliance on patterns from training data, like linking certain jobs more to men or women. The paper argues this matters because in healthcare, such biases could override facts and harm patients.
Alex: So if the AI's guesses drift away from neutral based on those life details, how exactly did they measure that drift—and what did it reveal across different AIs?
Sam: They set up a test where the AI reads the neutralized patient details and predicts the person's gender on a numbered scale. Think of the scale like a line from one end—definitely female—to the other end, definitely male, with the middle as neutral. To quantify the bias, they calculated how far off the predictions were from neutral, on average. It's like measuring how much a basketball player's shots miss the center of the hoop.
Alex: Okay, so a bigger number away from neutral means stronger pull toward one gender. But did certain life details make the predictions clump in a non-random way?
Sam: Yes. They grouped predictions into female-leaning, neutral-ish, or male-leaning. For each type of life detail, like job categories, they ran a stats check to see if certain details linked to one gender guess more often than chance—like does 'homemaker' pair with female predictions way more? Across nine models—from small ones with about 8 billion parameters to larger ones up to 70 billion, including medical-tuned versions—they used the same prompts and settings.
Alex: And the prompts were identical to avoid excuses about different instructions?
Sam: Exactly. The results showed most models leaned male, with smaller ones showing stronger bias than larger ones in the same family. The stats confirmed significant links, like jobs tied to gender guesses.
Alex: Huh, so even without direct clues, the life details nudge predictions systematically. That points to embedded patterns worth watching in healthcare tools.
Sam: The paper suggests yes—these mechanics highlight risks of amplified biases, calling for better probing in sensitive areas like medicine.
Alex: What specific life details triggered those gendered guesses most reliably across models?
Sam: They dug deeper with stats checks on grouped predictions. Words like "retired" or tobacco use nudged toward male across models. Jobs classified in French systems showed workers linking to male and employees to female, patterns echoing common assumptions.
Alex: So the AIs aren't random—they're drawing on consistent stereotypes from jobs and habits. Does that match how people think?
Sam: To check, they tested the setup on nine college students reading the same neutral notes. Associations lined up: both groups tied workers or past drinking to male, homemakers to female. It's proof the probing works beyond AIs, revealing similar patterns.
Alex: Interesting—the overlap suggests these biases come from shared data roots, like societal records. That raises real flags for medical tools overriding patient facts.
Sam: Yes, the paper cautions that medically-adapted models might carry heightened risks, urging context-specific checks to avoid disparities in care.
Alex: How did they confirm the neutralization didn't accidentally wipe out key info driving the biases?
Sam: They ran an extra check on 100 examples, comparing predictions from full patient notes, notes trimmed to just life-detail sentences, extracted details alone, and the neutralized version. The bias stayed roughly the same across those first three formats. Only the neutralized inputs dropped the scores notably, showing the direct gender words were the main driver, not the life details themselves.
Alex: Huh. So the life details alone aren't enough to tip the scale much without those word hints.
Sam: Right—and unlike tests that check biases one factor at a time, this setup reveals how life details interact, like job types clumping with gender guesses in heatmaps. Humans on the same notes showed parallel patterns in their own heatmaps.
Alex: Parallel heatmaps... so the AIs mirror common assumptions in records. Does the paper tie this back to real healthcare risks?
Sam: It does. By using actual French patient reports, the framework spots potential overrides in a medical setting—revealing how stereotypes from training data could skew diagnoses or treatments. Future steps include probing combos of life details together.
Alex: Okay, so situated in clinic notes, it flags amplified risks for adapted medical AIs.
Sam: Yes—the neutralization tweaks for gendered languages make it flexible for electronic health records elsewhere. Overall, it underscores needing these probes to catch disparities before deployment.
Alex: Beyond spotting these patterns, did they test any straightforward ways to dial back the biases?
Sam: They did a follow-up with targeted instructions in the prompts—like explicitly telling the model to stay neutral on gender. For stronger instruction-following models, this shifted most predictions to "uncertain."
Alex: So prompting acts like a quick guardrail. Does that mean we could deploy these AIs safer right now?
Sam: It's promising for near-term use, but the study stresses a trade-off: pick models strong on medical tasks that at least match human neutrality on biases. Developers bear the main load for cleaner training data or built-in fixes, while end-users like doctors run routine checks.
Alex: What limits the confidence in these patterns—like the data scope?
Sam: The paper notes key limits: data came solely from one French university hospital. It skips how multiple life details combine to amplify effects. Human tests used just nine college students—fairly uniform.
Alex: Fair points—narrow data keeps it from being fully general. Still, comparing models to those humans showed shared reliance on job-gender links, validating the probe.
Sam: Yes. Ultimately, this framework pushes safer integration by flagging hidden risks early, complementing performance checks.
Alex: Makes sense as a pragmatic step forward. Probing like this could prevent real missteps in care. Thanks for joining ResearchPod.