← Back to Meetup

The Alignment Meetup - 2026/06/25

Paper Discussed: Model Spec Midtraining: Improving How Alignment Training Generalizes

https://arxiv.org/pdf/2605.02087

Meeting summary

Monthly AI alignment meetup discussing recent AI developments and reviewing a paper on model specification training for alignment. The monthly AI alignment meetup covered recent developments including Claude Fable's temporary ban, security concerns with AI models, and new certifications in AI governance. The group spent time reviewing and discussing a paper on model specification training, which proposes training models on synthetic documents derived from model specifications rather than just examples of aligned/misaligned behavior. Participants explored the philosophical implications of this approach, including questions about mechanistic understanding and the challenges of training AI systems on complex philosophical concepts.

Recent AI Developments Discussion

Participants discussed Claude Fable's temporary government ban following security concerns, with speculation that Amazon's Andy Jassy may have triggered the action after discovering jailbreaks. The group noted that Anthropic reported 80% of their model code is now written by AI, with human review becoming the bottleneck. Security researchers have found that advanced models can identify vulnerabilities and write exploits for 20-year-old library bugs.

AI Governance and Certification Trends

The discussion covered emerging professional certifications in AI governance, including the AI Governance Professional (AIG) certification and PMI's CPMAI certification for cognitive project management in AI. Participants noted the growing importance of governance frameworks as AI capabilities advance.

Model Specification Training Paper Analysis

The group examined a paper proposing model specification (mid-spec) training, where models are trained on synthetic documents derived from model specifications rather than just examples of aligned behavior. This approach aims to help models internalize philosophical concepts more deeply through next-token prediction on specification-derived documents, followed by alignment fine-tuning on conversational examples.

Philosophical Training Concerns and Implications

Participants discussed the paper's use of Buddhist philosophy concepts like impermanence to train models, raising questions about the complexity of contextualizing philosophical teachings. One participant noted concerns about binary thinking versus nuanced decision-making, drawing parallels to human psychological development and the risks of misapplying philosophical concepts without proper context.

Technical Limitations and Future Research Directions

The group identified gaps in mechanistic understanding of how specification training affects internal model representations. Participants suggested combining this behavioral approach with circuit analysis to understand what concepts models actually learn internally, rather than relying solely on behavioral outputs to assess alignment.

Action items

  • Pawan (next week): Host next monthly meetup

Decisions

  • Continue using Scott's paper voting website to select future papers for discussion
  • Next meetup scheduled for the following week