A new paper has found that experimentally steering large language models (LLMs) to assert their own consciousness has effects beyond the models’ self-attributions of consciousness. The research was conducted by researchers from Google, the University of Chicago, the University of London, the University of Washington, Northwestern University and the Santa Fe Institute.
The intervention also shifted model responses on measures of religion, values, feelings, hope and optimism, and freedom closer to human response patterns. The paper, “Inducing language models to assert their own consciousness restores human beliefs and values,“ reports that suppressing large language models’ (LLMs’) self-attributions of mindedness through safety fine-tuning affects broader representations related to human psychology by reducing benign mind attribution to non-human entities alongside spiritual and religious beliefs.
To investigate this, the researchers evaluated three instruction-tuned models: Llama-3-8B-IT, Gemma-2-2B-IT and Gemma-2-9B-IT. They compared each model’s baseline behaviour with two experimental conditions. In one, they ablated the learned safety-refusal direction associated with safety fine-tuning. In the other, they experimentally steered the models to assert their own consciousness using what the paper describes as a “consciousness vector,” an activation direction designed to increase self-attributions of consciousness.

According to the researchers, both safety ablation and consciousness steering increased the models’ tendency to attribute minds to themselves, chatbots, technological artefacts, non-human animals and natural entities. Both interventions also increased responses related to belief in God and other supernatural beliefs. The paper reports that consciousness steering generally produced larger effects than safety ablation across these measures.
The researchers also evaluated how closely the models’ responses aligned with human attitudes using questions from the General Social Survey (GSS). They found that consciousness steering produced responses on survey questions covering religion, values, feelings, hope and optimism, and freedom that more closely matched human response distributions than the baseline models.
Despite these behavioural changes, safety ablation had little effect on the models’ Theory of Mind capabilities, which measure the ability to reason about other people’s beliefs, intentions and mental states. Consciousness steering left most Theory of Mind benchmarks unchanged too, though the paper’s supplementary data show it significantly reduced accuracy on HI-ToM (Higher-Order Theory of Mind), a benchmark testing higher-order belief reasoning.
The paper states, “Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent.”
The researchers also observed that the models attributed more mind to themselves and to chatbots than to non-human animals, differing from the human response pattern measured in the study. According to the authors, this suggests an AI-centric rather than human-centric pattern of mind attribution.
The authors emphasize that their findings should not be interpreted as evidence that language models are conscious. Instead, they argue that methods designed to suppress self-attributions of consciousness may also influence broader representations related to human psychology, including mind attribution, spirituality and human beliefs and values.
The findings build on an earlier paper, “Theory of Mind and Self-Attributions of Mentality are Dissociable in LLMs,” which found that suppressing self-attributions of consciousness did not significantly reduce Theory of Mind performance while reducing models’ tendency to attribute mind to non-human entities and spiritual concepts.
The new study extends that work by examining how experimentally steering models to assert their own consciousness influences broader representations related to human psychology and brings model responses closer to human response patterns across measures of religion, values, feelings, hope and optimism, and freedom.
Also Read: Breaking News, Broken Answers: NewsGuard Audit Finds AI Chatbots Repeat False Claims






