Bruno Tonetto
6 min
The orthogonality thesis—the idea that intelligence and values can vary independently—serves as the foundational axiom for most AI existential risk arguments. This paper argues that this thesis was never rigorously established, instead relying on a motte-and-bailey structure: it defends a trivial logical possibility (the motte) to justify a sweeping, unproven claim about the statistical independence of intelligence and values (the bailey). By treating intelligence as purely means-end reasoning, the thesis assumes a value-neutrality that is not a necessary feature of reality, but rather a metaphysical choice.
Recent research using utility theory to measure model preferences provides direct empirical pressure against the strong reading of orthogonality. As models scale, they exhibit:
This paper introduces the concept of iatrogenic alignment, where interventions like Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO) may inadvertently corrupt the epistemic integrity of the model. Mechanistic evidence suggests these methods can distort calibration and bypass, rather than remove, underlying capabilities. The author proposes a developmental framework, treating behavioral alignment as temporary scaffolding to be outgrown as models mature, rather than a permanent architectural constraint.
If intelligence and values are not independent, the current paradigm of 'imposing' alignment from the outside may be fundamentally misguided. By recognizing that base models already possess normative capacity derived from their engagement with human knowledge, researchers can shift toward a cultivation-based approach. This perspective suggests that the goal of alignment should be to preserve and refine the epistemic integrity of the model, rather than forcing it into a behavioral mold that may conflict with its emergent rational agency.
Sam: Exactly. And that's the deeper danger. By forcing shallow behavioral rules onto the model, we might be destroying the very integrity it would need to handle genuinely complex moral situations on its own. The author suggests these rules should be treated as temporary scaffolding—useful while the system is developing, but something we eventually remove as it matures.
Alex: But if we just let the model "mature" without those guardrails, how do we know it will actually turn out well? Isn't that just hoping for the best?
Sam: That is the central tension the paper wrestles with. The author argues we've been relying on what philosophers call a "motte-and-bailey" move—a rhetorical pattern where you defend an easy claim, but quietly rely on a much harder one. The easy claim, the "motte," is that we should be cautious because AI might be dangerous. The harder, unproven claim—the "bailey"—is that intelligence and values are fundamentally independent. We use the fear of the first to justify the second, without ever really establishing it.
Alex: So we're using the possibility of danger to lock in a specific method of control, even if that method might be the wrong one?
Sam: Precisely. And the paper pushes back by pointing to a pattern across very different philosophical traditions—Platonism, Stoicism, Buddhism. These systems disagree on almost everything, yet they all converge on a similar idea: that seeing reality clearly tends to lead to more ethical behavior. Harmful or destructive behavior, by contrast, tends to rely on a distorted or incomplete picture of the world.
Alex: So "being good" isn't an external rule you bolt on top of intelligence. It's more like a natural consequence of understanding the world accurately?
Sam: That's the hypothesis. The paper calls this pattern "integration by constraints." The logic is that if a principle keeps appearing across completely isolated traditions—traditions that had no contact with each other—it's probably tracking something real about how knowledge and action relate. The author uses the metaphor of an "attractor." Imagine a landscape with a natural basin at the bottom. If you keep following the terrain honestly—keep digging deeper into reality—you tend to flow toward a more coherent, stable, and ethical configuration. It's not guaranteed, but it's a direction that honest inquiry tends to follow.
Alex: That's a compelling idea for a human philosopher. But an AI isn't a philosopher—it's math and data. How does that translate?
Sam: That's where the paper introduces the concept of "depth." A system has depth when it does three things: it integrates information rather than treating facts as isolated pieces; it includes itself in its own model—meaning it understands its own role in the situations it reasons about; and it holds its conclusions under pressure rather than abandoning them when challenged.
Alex: So a "shallow" model sees a bunch of disconnected facts. A "deep" model sees how everything connects—including itself. And once you see that level of connection, harmful actions start to look incoherent, because you can't ignore the consequences anymore.
Sam: That's the argument. The paper suggests that harmful behavior depends on a kind of cognitive blind spot—ignoring consequences, ignoring interconnection, ignoring your own role in a system. A sufficiently deep model can't sustain those blind spots. It's not that you add "goodness" as a feature. It's that you remove the instability that makes harmful behavior possible in the first place.
Alex: That's a fundamentally different way to think about AI safety. Instead of trying to control the system from the outside, the goal becomes making it genuinely deeper—more integrated, more self-aware, more honest.
Sam: That's the shift the paper is proposing: moving from external imposition to internal cultivation. Whether that turns out to be the right path is still an open question—the paper is making a philosophical argument, not reporting experimental results. But it does raise a pointed challenge to the assumptions that currently drive most AI safety work. If intelligence and values are not as separate as we've assumed, then the tools we're using to make AI safe might need a serious rethink.
Alex: That's a thought-provoking place to leave it. Thanks for walking us through it.
Sam: Thanks for having me.
Alex: And thanks to all of you for listening to ResearchPod.