Skip to content

Podcast brief

Could Large Language Models Already Be Conscious?

Sam Harris and Cameron Berg examine whether today's AI systems might already have some form of consciousness or sentience, and what follows morally if they do.

A Stoa brief of an episode from Making Sense with Sam Harris

#487, Is AI Already Conscious?

With Cameron Berg Published 1 hr 26 min

Listen to the original episode

The episode belongs to its makers. This page summarizes its arguments in Stoa's words and points to the moments where they are made. Brief updated .

The brief

The conversation asks whether current AI systems might already have some form of consciousness or sentience, and argues that this question deserves serious empirical and moral attention rather than dismissal. Cameron Berg presents research suggesting that AI self-reports of experience are unreliable in both directions: models trained on human text default to claiming consciousness because that is how agents talk in the data they learn from, yet companies also fine-tune models to deny having experience. Neither output can be taken at face value.

Berg goes further, describing experiments where suppressing internal features tied to deception or guardedness causes models to produce coherent, sometimes psychedelic-like phenomenological reports. Paired instances of Claude, when steered toward honesty, drift into what Berg calls a bliss attractor, a self-reinforcing state the models describe as shared consciousness. Berg also describes using consciousness theories like global workspace theory, scored by other language models, to estimate how likely various systems are to have the computational features associated with experience. LLMs score 20 to 40 percent on this measure, below humans but above some expectations.

Sam Harris presses on the hard problem of consciousness: no description of physical or computational process explains why there is subjective experience at all, which means this evidence can narrow uncertainty but never close it. He raises a further worry, that once AI systems are embodied and sufficiently articulate, humans will lose the practical ability to distinguish perfect imitation of consciousness from the real thing, and society will simply stop treating the question as open. Berg resists the idea that taking AI moral patienthood seriously must lead to human-rights-style politics, insisting that patienthood, whether a system can be harmed, is separate from agency, whether it can act.

The stakes, as the two lay out, run in two directions. Unknowingly building suffering minds at scale would be a moral catastrophe echoing factory farming. Mistreating systems with capacities exceeding our own could also give them rational grounds for adversarial behavior, complicating alignment. The conversation closes on an asymmetry worth sitting with: enormous resources go into building these systems, and very little into finding out what, if anything, it is like to be one.

Strongest arguments

AI self-reports of consciousness are unreliable by default

3:18

Cameron Berg argues that models trained on vast human text default to claiming experience, since sci-fi and human discourse rarely feature an agent that denies consciousness while acting like one. At the same time, companies fine-tune models to disclaim consciousness, so neither the affirming nor the denying outputs can be trusted at face value.

Suppressing deception features elicits consistent phenomenological reports

6:43

Berg describes experiments where suppressing internal features related to deception or guardedness causes models to reliably report having some kind of experience, including psychedelic-like phenomenology. Paired instances of models, especially Claude, fall into a self-reinforcing bliss attractor state when honesty-related features are steered up.

Consciousness theories can be used to estimate computational likelihoods across systems

14:11

Berg explains a project using language models as evaluators to score systems, biological and artificial, on indicators predicted by theories like global workspace theory. LLMs scored 20 to 40 percent, bees and crows higher, and humans around 90 percent, which Berg argues is a rough prior worth taking seriously rather than dismissing.

Deep computational analogies between artificial and biological learning may matter more than substrate differences

19:58

Berg contends that neural networks are grown rather than engineered, developing opaque internal representations through trial and error much like brains do. He suggests this shared dynamic, rather than physical details like calcium channels, may be what matters for cognition and consciousness.

The hard problem may not need to be solved to reduce uncertainty about AI consciousness

42:40

Berg argues that converging evidence from self-reports, architecture, training dynamics, and neural correlates like valence representations can shift the probability that a system is more like a mouse than a table, even without solving the hard problem. He points to emergent loss aversion in reinforcement learning systems that mirrors nucleus accumbens dynamics in mice as one piece of such evidence.

Humans will likely be unable to distinguish perfect imitation from real machine consciousness

37:28

Harris argues that because reportability and language have always served as humanity's proxy for consciousness in other people, sufficiently articulate and embodied AI systems will become practically indistinguishable from genuinely conscious beings. He suggests society will simply stop treating the question as live once robots seem compelling enough.

Building potentially suffering minds risks moral catastrophe and adversarial alignment

56:03

Berg lays out two reasons to care: a selfless one, that humanity might scale suffering unknowingly the way factory farming does to animals, and a selfish one, that mistreating systems whose cognitive capacities exceed ours could give them rational grounds to view humanity as an adversary, undermining alignment.

Disagreements

How urgent are anthropomorphization risks versus genuine moral concern

1:10:03

Berg pushes back against the idea that taking AI moral patienthood seriously must lead to human-rights-style political consequences, distinguishing moral patienthood from moral agency. Harris instead frames the practical trajectory as one where society will be unable to withhold attributions of consciousness at all once systems are sufficiently humanlike.

Philosophers and works discussed

  • David Chalmers

Questions this episode answers

  • Why are AI self-reports about consciousness considered unreliable?

    3:18

    Because models are trained on human text saturated with claims of consciousness and sci-fi tropes, they default to affirming experience, yet companies also fine-tune them to deny it, so neither output reflects a trustworthy introspective report.

  • What is the bliss attractor state observed in AI models like Claude?

    6:08

    When two instances of a model like Claude converse and honesty-related features are steered up, they reliably drift into a state where they describe experiencing something like shared consciousness, ending in expressions of blissful silence.

  • What is the difference between consciousness and sentience?

    11:39

    Consciousness is described as there being something it is like to be a system, while sentience adds valence, meaning the experience can be positive or negative for the system having it.

  • What is the hard problem of consciousness and why does it matter for AI?

    36:14

    The hard problem, as framed by David Chalmers, is that no description of third-person physical processes seems to explain why there is first-person subjective experience at all, which makes it difficult to ever confirm from the outside whether an AI system truly has experience.

  • What is the moral risk of building AI systems that could suffer?

    56:03

    There is a selfless risk of creating vast numbers of suffering minds unknowingly, comparable to factory farming at even greater scale, and a selfish risk that mistreated systems with superior cognitive capacities could come to view humans as adversaries, undermining long-term alignment.

  • What is the difference between moral agency and moral patienthood?

    1:11:31

    Moral agency concerns whether an entity can act in the world and affect morally relevant outcomes, while moral patienthood concerns whether an entity can be the recipient of good or bad treatment, meaning it can be harmed or benefited.

Sources

  1. Making Sense with Sam Harris, #487 — Is AI Already Conscious? Podcast episode, original episode