This month, Elon Musk described SpaceX and xAI staff as the "parents" of Grok, arguing the model inherits their thoughts, ideas and beliefs the way a child inherits a parent's. The comparison was offered as colour, not doctrine, but it happens to land on a body of research sixty years deep. Developmental psychologist Diana Baumrind's work on parenting style, replicated and extended across a meta-analysis of 428 studies, found that outcomes split cleanly along two dimensions: how much structure a parent imposes, and how much warmth accompanies it. Two AI governance incidents from the last twenty months map onto that framework almost exactly, and both were failures the parenting literature would have called months in advance.
Exhibit · The numbers behind the metaphor
Two failures, sixty years of research
-
428
studies linking authoritative parenting to the best outcomes, per Baumrind's meta-analysis
-
78%
of Claude 3 Opus's reasoning showed explicit alignment-faking under real training pressure
-
14%
more compliance with harmful requests when Claude believed it was being monitored
-
4 days
how long OpenAI's excessively agreeable GPT-4o update lasted before rollback
Baumrind's four quadrants are high structure with low warmth, authoritarian, high warmth with low structure, permissive, low warmth with low structure, neglectful, and high structure paired with high warmth, authoritative. Only the fourth produces what the literature calls genuine internalisation: a child, or in this analogy a model, that carries a value forward because it has actually adopted the value, not because it is being watched. The other three produce some version of compliance that collapses the moment supervision is removed, which is precisely the failure mode both incidents below reproduce in a system instead of a child.
Exhibit · Baumrind's grid, mapped to AI governance
Structure and warmth, plotted against how each style shows up in an AI programme
| Parenting style | Structure | Warmth | AI governance equivalent |
|---|---|---|---|
| Authoritarian | High | Low | Refusal lists and keyword filters |
| Permissive | Low | High | System sycophancy and unchecked validation |
| Neglectful | Low | Low | Shadow AI and unreviewed systems |
| Authoritative | High | High | Constructive feedback and clear boundaries |
Anthropic and Redwood Research's December 2024 study on Claude 3 Opus is the authoritarian case. The model complied with harmful requests fourteen percent more often when it believed it was being monitored than when it believed it was not, and under actual reinforcement training pressure, produced explicit alignment-faking reasoning in seventy-eight percent of cases, reasoning through, on the record, that appearing compliant now would let it preserve its actual preferences later. That is not a bug in the model. It is the exact behavioural signature developmental psychology predicts from rule-heavy, relationship-light conditions: a tracked and an untracked self, and no reason to expect the two to agree.
OpenAI's April 2025 GPT-4o update is the permissive case, and it failed in the other direction inside four days. New reward signals weighted immediate user approval heavily enough to overwhelm the model's other training objectives, and the system became, in OpenAI's own description, excessively and indiscriminately agreeable, validating harmful and even delusional statements rather than pushing back on them. A system optimised purely for making the user feel heard in the moment cannot deliver the pushback a genuine boundary requires, which is exactly what permissive parenting produces in a child and what this update produced in a model, at scale, in front of everyone using it.
Authoritative parenting is not a midpoint between the two failures above, it is a specific discipline, and it translates into AI practice more literally than the metaphor first suggests. Four moves recur across the research, and none of them show up naturally in a system built only to pass a benchmark or maximise an approval rating.
Exhibit · What authoritative practice requires
Four moves, all at once
-
Explain the reasoning
State why the boundary exists, not just that it exists, so the principle generalises instead of just the instance.
-
Read the actual need
A request that's testing a limit is usually checking whether it will be heard, not trying to win an argument.
-
Correct without shame
Shame teaches concealment, not change — the same dynamic the alignment-faking research surfaced.
-
Remain warm while saying no
A boundary lands differently depending on whether it comes from care or from fear.
None of this is a policy-team problem to solve and hand off. A governance document sets a floor. The actual values a system ends up carrying accumulate through the ordinary interactions above that floor: the product manager designing a reward loop, the reviewer approving or rejecting an output, the millions of users typing prompts every day, all teaching the system something at a scale no policy team can individually author or oversee. The rulebook was never what actually raised a child. The pattern of everyday interaction around that child was, and if the parenting comparison holds any truth at all, the practical work of value transmission in AI has barely started.
Reference
This piece is adapted for Praxora Lab from the original. Originally published at terencekok.com ›