AI Engineering / Anthropic 与 Claude
Anthropic 可解释性 + Claude 行为研究(Natural Language Autoencoders + 用户使用洞察)
Anthropic research 板块 5 篇 hub:Natural Language Autoencoders(把 Claude 内部 activations 解码成自然语言,2026-05 重大突破)+ How people ask Claude for personal guidance(100 万对话采样 / 6% 是寻求人生建议)+ BioMysteryBench(生物信息学评测)+ Why do AI models hallucinate? + What is sycophancy? + 4 个研究 team 页(Alignment / Economic Research / Interpretability / Societal Impacts)。覆盖"AI 内部怎么想 + 用户怎么用 + 容易出什么 bug"3 块研究