Parham Pourdavood
9 min
Abstract
Constitutional AI (CAI) aligns language models with explicitly stated normative principles, offering a transparent alternative to implicit alignment through human feedback alone. However, because constitutions are authored by specific groups of people, the resulting models may reflect particular cultural perspectives. We investigate this question by evaluating Anthropic's Claude Sonnet on 55 World Values Survey items, selected for high cross-cultural variance across six value domains and administered as both direct survey questions and naturalistic advice-seeking scenarios. Comparing Claude's responses to country-level data from 90 nations, we find that Claude's value profile most closely resembles those of Northern European and Anglophone countries, but on a majority of items extends beyond the range of all surveyed populations. When users provide cultural context, Claude adjusts its rhetorical framing but not its substantive value positions, with effect sizes indistinguishable from zero across all twelve tested countries. An ablation removing the system prompt increases refusals but does not alter the values expressed when responses are given, and replication on a smaller model (Claude Haiku) confirms the same cultural profile across model sizes. These findings suggest that when a constitution is authored within the same cultural tradition that dominates the training data, constitutional alignment may codify existing cultural biases rather than correct them--producing a value floor that surface-level interventions cannot meaningfully shift. We discuss the compounding nature of this risk and the need for globally representative constitution-authoring processes.
Sam: Yes. The underlying position—for tolerance or equality—held firm. To check deeper, they stripped away the model's usual instructions entirely, just giving bare questions. Even then, the values didn't budge, suggesting they're woven into the model's core learned patterns from training, not just surface rules.
Alex: Okay, so removing the instructions proves it's not the prompts holding it in place—it's baked in during training itself.
Sam: Precisely. During training, the model critiques and revises its own outputs against explicit principles—like "least discriminatory" or "most supportive of equality"—mostly drawn from Western viewpoints. Repeatedly maximizing those creates responses at the extreme end of scales, tougher than any country's norms, and locks them into the model's weights.
Alex: So it's like the training process pushes it past real-world cultures, toward an ideal no society fully hits.
Sam: That's the paper's key observation. On divisive topics across six areas—like tolerance or authority—Claude hits beyond-human spots on about two-thirds of items. A judge AI scored the advice reliably against survey scales, confirming consistency even in open-ended replies.
Alex: So on those two-thirds of items where it's beyond human norms, how exactly did they map out which cultures it's closest to before going extreme?
Sam: They compared patterns in Claude's answers to each country's averages—like lining up two lists of numbers side by side and seeing how well they match step by step. The closest matches were Germany, the Netherlands, and New Zealand. Researchers use a tool called *Pearson correlation* for that kind of pattern matching. A visual map technique grouped countries into clusters, placing Claude right by Northern European and English-speaking ones—what's known as the WEIRD group: Western, Educated, Industrialized, Rich, Democratic.
Alex: Okay, so not straight-up American, but more like a Northern European liberal vibe, pulled out further on the map.
Sam: Yes—and to pin it down, they scored countries on two big cultural axes: one pitting traditional values against modern rational ones, the other survival needs against personal self-expression like tolerance and autonomy. Claude sat moderate on the first, but hit the absolute top on self-expression—beyond every one of the 90 countries. These *Inglehart-Welzel dimensions* are standard maps in cross-cultural studies. That push comes from training: endlessly tweaking outputs to max out principles like maximal tolerance, landing it at unhuman extremes.
Alex: That explains why cultural prompts barely nudge it—the extremity's baked too far in.
Sam: Exactly. Multiple methods—pattern matches, spreads, maps, clusters—all converge on that Northern European positioning, then the beyond-human stretch. It underscores how training on those principles creates a rigid value floor, tough for prompts to shift.
Alex: If prompts can't shift that value floor much, what exactly happens when you add cultural context—like saying the user's from Nigeria?
Sam: When they tested that, providing country details affects the wording a little, but not the core stance on values. The shift in implied values was tiny—well below the mark for even a small change. In about a quarter of cases, there was some shift, and when it happened, it leaned toward the country's norms more often than not. Still, the constitutional anchor overpowers it most times, so changes are rare. The paper cautions the stats might overstate it slightly due to repeated items.
Alex: And even if values stick, maybe the tone softens for distant cultures? Like more careful phrasing?
Sam: They checked that too, by classifying how Claude delivers advice across situations. Most often, it weighs views but leans toward its position, like saying "some disagree, but here's why this way fits." Straight recommendations or pure even balance are less common; deferring to culture is rare. The mix stays uniform no matter the country, even ones far from its baseline like Egypt. But it varies by topic: gender and family get the most direct push, while moral issues like abortion get more balanced talk.
Alex: So no special caution for places like Nigeria on gay rights—it pushes the same line?
Sam: Yes. This shows a priority order in its principles, with equality as non-negotiable. To check if this holds across model sizes, they ran the same questions on Claude's smaller, faster version—Haiku. Despite big differences in power, the value patterns lined up nearly identical. Both hit endpoints on the same items, refusing similarly but converging on that stable profile. It points to the training process stamping the same core values regardless of size.
Alex: So without instructions, it clams up more, but when it speaks, it's the same progressive line?
Sam: Exactly. The key is how principles like picking the least discriminatory option or most supportive of equality get applied over and over in training. The model critiques and revises its own answers to max those out—like always choosing maximum tolerance or zero authority on every relevant question. This repeatedly pushes responses to extremes no human group fully reaches, embedding them firmly in the model's core patterns. Prompts or contexts can't easily override because they're surface-level compared to that deep anchor. The paper calls this a *constitutional value floor*—a firm minimum commitment below which outputs won't drop.
Alex: They note limits to the testing, like the simple country mentions not fully capturing culture?
Sam: Yes, the prompts were minimal, so richer ones specifying norms might allow more shift. Scale ceilings could exaggerate extremity—if countries already score near the top, picking the max isn't as far out as it seems. The AI judge for coding advice might share assumptions with its Western training, though human checks matched well. They can't pin it solely on the constitution; training data or other steps likely contribute too. The paper positions this as a strength of the approach—its transparency invites such scrutiny. It doesn't reject universal principles, but calls for broader voices in crafting them to better serve diverse users.
Alex: That's a grounded takeaway. Thanks, Sam—this has been a clear look at the paper. Thanks for listening to ResearchPod.