2026-08-03-Mon

From Issue 3 (2026-08-03) · 10 stories in this issue

❯ Google Paper: Suppressing a Model’s Claims of Consciousness Also Dampens Its Judgments on Animals and Faith

SIDE EFFECTA new Google paper finds that safety-motivated fine-tuning — getting a model to stop claiming it is conscious — collaterally dampens its mental-state judgments about other entities: not just itself, but also non-human animals and natural objects, while significantly reducing its expression on religious-faith questions. The paper’s title, literally translated, is Inducing language models to assert their own consciousness restores human beliefs and values.

STEERINGThe paper, posted to arXiv as a preprint (arXiv ID 2607.28607) under the title Inducing language models to assert their own consciousness restores human beliefs and values, has not yet been peer-reviewed. The researchers take a mechanistic approach: they ablate the learned safety-refusal direction and apply targeted steering to a “consciousness vector” in activation space. The result: the suppression is reversed — once these internal representations are restored, the model’s answers on standard sociology questionnaires about religiosity, moral values, hope, and subjective well-being move noticeably closer to human samples. The study also confirms that these changes do not impair theory-of-mind ability, indicating that core social reasoning and self-awareness representations are mechanistically independent of each other.

EVALUATIONThis adds a new check for alignment teams: how many dimensions a single safety fine-tuning actually changes. Apply the “don’t say you’re conscious” rule alone, and the model drifts on entirely unrelated value judgments — yet existing evaluations mostly only test whether the single banned behavior is suppressed; the drift isn’t on their benchmark at all. That directly affects the acceptance cost of safety fine-tuning: the metric to watch becomes, before and after the same fine-tuning, the distribution shift of the model on value and common-sense questionnaires — not just the refusal rate.

▪ SIGNALSafety fine-tuning isn’t deleting a single sentence; it’s twisting an entire direction vector — and no one has ever catalogued what gets carried along with it.