APLENDE/

marketplace /

AI Safety and Alignment: A Grounded Eight-Hour Introduction

by admin · AI safety and alignment

A mastery-oriented overview of the technical, empirical, and governance problems involved in building beneficial AI systems, with fair treatment of skeptical, moderate-concern, and existential-concern positions.

Free account, then it's yours: your own private copy with a clean slate and its own review schedule.

Sign up free to take it →

4

modules

12

lessons

8.5h

estimated work · target 8h

What's inside

36 tasks · 30 flashcards · 42 concepts · 36 recall prompts

remember
6
understand
16
apply
7
analyze
6
evaluate
0
create
1

Sample recall prompts — this is what studying it feels like

  • Define AI safety and alignment in your own words, and explain how they differ.
  • Name three dimensions that affect AI risk besides raw model capability.
  • Give a fair one-sentence version of the skeptic, moderate-concern, and existential-concern positions.

Syllabus

Framing the Problem

Separate safety from alignment, learn why proxies fail, and understand outer versus inner alignment.

  • A Map of the AI Safety Debate reading · flashcard · quiz · 5 cards
  • Objectives, Proxies, and Specification Gaming reading · quiz · teach_back
  • Outer Alignment, Inner Alignment, and Mesa-Optimization reading · flashcard · essay · 5 cards

Technical Approaches in Current Systems

Study human feedback, constitutional training, scalable oversight, interpretability, and their limits.

  • Human Feedback and RLHF reading · flashcard · quiz · 5 cards
  • Constitutional AI and Scalable Oversight reading · quiz · teach_back
  • Interpretability and Mechanistic Evidence reading · note · quiz

Evaluation, Robustness, and Control

Learn how developers search for risky capabilities, stress-test safeguards, and reason about corrigibility and deployment.

  • Evals, Red Teaming, and Capability Thresholds quiz · reading · flashcard · 5 cards
  • Robustness, Adversarial Use, and Misuse reading · quiz · teach_back
  • Corrigibility, Control, and Deployment Norms reading · flashcard · essay · 5 cards

Governance, Disagreement, and Synthesis

Connect technical uncertainty to governance choices and produce a balanced safety brief.

  • Governance Levers: Compute, Audits, and Coordination reading · quiz · teach_back
  • Serious Disagreements and Empirical Uncertainty reading · note · quiz
  • Capstone: Build a Balanced AI Safety Brief project · reading · flashcard · 5 cards

Concepts it teaches

AI safetyAlignmentCapabilityEmpirical uncertaintyObjectiveProxySpecification gamingReward hackingOuter alignmentInner alignmentMesa-optimizerHuman feedbackPreference learningReward modelRLHFConstitutional AIScalable oversightAI feedbackInterpretabilityMechanistic interpretabilitySuperpositionEvalsRed teamingSystem cardCapability thresholdRobustnessMisuseDual useCorrigibilityShutdownabilityAI controlDeployment normsCompute governanceThird-party auditIncident reportingInternational coordinationSkeptic positionModerate concernExistential riskBurden of proofRisk registerSafety case