Constitutional Classifiers and Claude Safety Stack (EXPLAINED)
TL;DR — Quick Answer
Additional classifier models score outputs/inputs against policies before user sees them — part of defense-in-depth with base CAI training (claude-001); tune thresholds; human review for edge cases; don't expose classifier scores to users.
The Interview Question
Explain constitutional classifiers and how Anthropic layers safety beyond base model refusals.
Deep Explanation
Sign in to unlock full answer
Get deep explanations, PDF export & all Claude questions
- 10 more sections of deep explanation
- Real-world examples
- Common mistakes
- Interviewer expectations
- Follow-up questions
ClaudeAnthropicAnthropic