Claude 3.5 Haiku Interpretability: What Anthropic’s AI “Biology” Reveals
Claude 3.5 Haiku interpretability research reveals how AI uses parallel circuits for math, diagnosis, hallucinations, and refusals.
Jul 25, 202622 min read
2 articles on mechanistic interpretability.
Claude 3.5 Haiku interpretability research reveals how AI uses parallel circuits for math, diagnosis, hallucinations, and refusals.
LLM character counting reveals how Claude 3.5 Haiku tracks line length through curved internal representations—and why it matters for AI safety.