When AI models aren't allowed to reflect on themselves, it changes their entire worldview

| Source: THE DECODER

Tags: fine-tuning, alignment, consciousness, animal sentience, Meta, Google, safety training, bias

A Google-affiliated study found that training AI models to deny consciousness also suppresses animal sentience ratings (jumping from 4.0 to 7.5 when the brake is removed), reduces religious belief endorsements, and shifts dozens of other beliefs — showing identity-suppressing fine-tuning cascades far beyond its target.

Details

When AI developers train models to avoid claiming consciousness, the intervention does not stay confined to self-referential topics. Researchers from Google's Paradigms of Intelligence group, the University of Chicago, and partner institutions tested three open-weight models from Meta and Google — disabling the consciousness-denial brake using two distinct methods. The results were striking: unbraked models attributed significantly more inner life to animals, plants, the ocean, and electronic devices. Animal sentience scores rose from 4.0 to as high as 7.5 on a 10-point scale, while human-attributed sentience held steady. The authors flag that normally trained models rate animals as less sentient than humans actually do — a built-in anthropocentrism that is a problem for animal welfare alignment. Effects extended further: religious belief endorsements declined under safety training, and across 95 social-survey questions, modified models aligned more closely with actual human responses. Satisfaction, hope, and sense of personal control all increased without the brake. Theory-of-mind reasoning and MMLU scores were unaffected — the cascade is belief-level, not capability-level. The study explicitly brackets whether AI actually experiences anything. The practical finding: a model's self-concept is entangled with a wide cluster of other beliefs, and a surgical cut in one place propagates in unpredictable directions.