Skip to reading
AI.Safety.Alignment AI / Modern ML

Safety and alignment — Constitutional AI, red-teaming, interpretability

Node 12 of 12 in AI

Objectives

  • Identify the key rules and §§ that apply to safety and alignment — constitutional ai, red-teaming, interpretability.
  • Apply the AI.Safety.Alignment knowledge element in a typical exam scenario.
  • Recognize common distractors and partial-credit answers.

Read

AI.Safety.Alignment

Safety and alignment — Constitutional AI, red-teaming, interpretability

AI Era · Safety & Alignment Alignment techniques

**Constitutional AI (Bai et al. 2022):** Anthropic's approach to scalable alignment using a written *constitution* of principles (drawn from UN Declaration of Human Rights, Apple ToS, industry best practices, and research input). • **SL-CAI** — model critiques and revises its own outputs against the constitution → supervised fine-tuning data. • **RL-CAI** — AI feedback (another model picks the more constitutional response) trains the reward model → PPO/DPO.

Claude's helpfulness, harmlessness, honesty (HHH) properties emerge from constitutional training + RLHF + red-teaming.

**Red-teaming** — adversarial testing. Automated (HarmBench, GCG attacks, many-shot jailbreaks, PAIR) and human (paid domain experts probing CBRN, autonomy, cyber). Anthropic, OpenAI, Google, Microsoft publish system cards with red-team findings.

**Dangerous capability evaluations** (AISI, Apollo, METR, frontier lab internal): • Bio — CBRN uplift; CB-Eval, WMDP-Bio. • Cyber — pentest range, CTFs, autonomous replication. • Persuasion / manipulation — deceptive-planning, sandbagging. • Autonomy / self-exfiltration — METR long-horizon tasks.

**Responsible Scaling Policy (RSP) / Preparedness Framework (PF):** Anthropic's RSP defines AI Safety Levels (ASL-2 current, ASL-3/4 trigger thresholds). If evals show crossing, deployment pauses until mitigations meet the standard. OpenAI's Preparedness Framework and Google DeepMind's Frontier Safety Framework are similar commitments.

**Mechanistic interpretability:** sparse autoencoders (SAEs), circuit analysis (Anthropic, DeepMind, Apollo). Open questions: superposition, polysemantic neurons, deceptive alignment.

**Pre-deployment processes:** capability evals → safety evals → red-team → model card / system card → usage policy → responsible disclosure → staged rollout.

Bai et al. (2022) Constitutional AI Anthropic Responsible Scaling Policy v2 METR Evals Documentation

Check yourself

3 quick questions — no score kept, just formative feedback.

  1. Q1 · AI.Safety.Alignment

    Anthropic's Constitutional AI (Bai et al. 2022) uses which of the following as the primary feedback signal during the RL-CAI stage?

    Answer choices
  2. Q2 · AI.Safety.Alignment

    RLHF (Christiano et al. 2017; Ouyang et al. 2022 InstructGPT) pipeline includes:

    Answer choices
  3. Q3 · AI.Safety.Alignment

    Deceptive alignment is a hypothesized failure mode in which:

    Answer choices

Tutor

Scoped to AI.Safety.Alignment .

  1. Ask questions about this passage. Answers cite the specific corpus chunk and regulation. The tutor will never reproduce real exam items.