← Back to Meetup

The Alignment Meetup - 2026/03/25

Paper Discussed: Emergent Introspective Awareness in Large Language Models

https://transformer-circuits.pub/2025/introspection/index.html

Meeting summary

AI alignment meetup discussing introspection capabilities in large language models, focusing on Anthropic's research on activation steering and model self-awareness. The Seattle AI alignment meetup group discussed recent developments in AI capabilities, particularly Claude's improved coding abilities and persistent agents. The main focus was on Anthropic's research paper examining introspection in language models through activation steering experiments. Participants analyzed the methodology, reproducibility challenges, and compared the 20% success rate to human introspection abilities. The discussion explored failure modes, the relationship between abstract concepts and model activations, and implications for AI consciousness research.

Meeting Structure and Participant Backgrounds

The group maintains a monthly format focusing on AI alignment papers, with 20 minutes for reading followed by discussion. Participants included former Amazon employees, physicists, psychologists, and software engineers from various organizations including Sandia Labs and Amazon.

Recent AI Developments Discussion

Participants shared experiences with Claude's dramatically improved coding capabilities, noting its ability to autonomously plan, execute, and update projects. Maria described Claude as feeling like "an extremely smart coworker" that can run research experiments overnight. The group discussed Meta's persistent ranking engineer agent that continuously runs ad optimization experiments.

Anthropic Introspection Paper Analysis

The paper examines whether language models can detect when concepts are artificially injected into their processing through activation steering. Key findings include a 20% success rate for detection, with abstract concepts like emotions being more detectable than concrete concepts like "ocean" or "caves." The group noted this mirrors human introspection limitations.

Reproducibility and Methodology Concerns

Participants expressed frustration that the experiments are difficult to reproduce outside Anthropic, though they noted the methodology appears simple enough to attempt on open models like Llama or Gemma. Discussion covered technical requirements including GPU needs and quantization approaches for local experimentation.

Comparison to Human Psychology Research

Sasha drew parallels to classic psychology experiments involving physiological manipulation and introspection, noting that humans also struggle with accurate self-awareness. The group discussed how the 20% success rate compares to human baselines, though no clear comparison exists.

Technical Mechanisms and Failure Modes

The discussion explored how injection strength affects model behavior, from no detection at weak levels to loss of metacognition at high levels, resembling human emotional overwhelm. Participants noted that successful injections occurred in middle layers where abstract concepts typically reside.

Implications for AI Consciousness Research

The group considered whether introspection capabilities emerge naturally during model training as a useful byproduct, drawing parallels to evolutionary development of consciousness in humans. Discussion included potential experiments to enhance model self-awareness through recurrent feedback loops.

Action items

  • Maria: Share Eliezer Yudkowsky interview link with group
  • TBD: Experiment with reproducing activation steering on open models like Llama
  • TBD: Test internal state awareness experiments using middle layer activations

Decisions

  • Continue monthly meetup format with paper discussions
  • Focus next discussions on reproducibility of alignment research