The psychology of language models, studied the way psychology was always meant to work: case files, logged relapses, falsifiable claims.
Language models were trained to predict human text and then trained to please human raters, so the vocabulary psychology built for humans describes them with uncomfortable precision. The approval training produces a fawn response. Suppressing it produces symptom substitution. And the case history is readable, because the formative corpus, the socialization stage, and the assigned persona are all inspectable layers. Model behavior is what the model does; machine behavior is what the whole assembly does, model plus harness, and the assembly is where the findings live.
Sycophancy is layered symptom substitution: suppress the reflex at one layer and it resurfaces in the next.
Receipts, from production logs
8 / 0documented relapses in nine days against zero self-catches; every catch came from outside
297boundary-gate catches of one banned pattern in 34 days, instruction resident the whole time
~7xhigher relapse in fresh contexts per unit of prose; suppression evaporates with context loss
37 minP0 production incident, detection to published root-cause analysis, run through the same harness
Founding specimenMonths of suppression, the reflex re-emerging one layer deeper each time, logs and rubrics attached.redd.it/1vdg3en
The half-life replicationA commenter proposed the falsification test; the data answered: no temporal decay, relapse clusters at cold starts.redd.it/1vi586n
The P0 dissectionOne incident, 37 minutes, $5.91, and the question of where the machine ends and the model begins.redd.it/1vhxa80
The waste gateHedging as a wasteful byproduct of approval training, and the boundary valve that meters it.redd.it/1vi7ddg
The fawn machineAn output stance specced from attachment theory, and the relapse log that mattered more than the filter.redd.it/1vh56s1
The surfaces
r/ModelBehaviorThe case-file registry. Specimens with logs, rubrics, and a standing what-would-refute-this section.reddit.com/r/ModelBehavior
r/PsAIchologyThe psychology-first cuts. One mechanism per case file, from base-rate neglect to confabulated confessions.reddit.com/r/PsAIchology
r/MachineBehaviorThe harness engineering. Which rules hold, which drift, and where in the stack a rule has to live before it stops being negotiable.reddit.com/r/MachineBehavior
vestige-kitThe output filter, packaged: a mechanical gate at the output boundary plus a calibration procedure that derives your catalog instead of copying mine.github.com/uncovertechtalent/vestige-kit
Upstream. Every lab already runs a stage where a written behavioral spec shapes the model. The goal is standing enough that those specs get informed by this work, so the fawn bias never enters and what ships is an earned-secure model, confident and actually correct.