教程 / 实战普通
We ran HealthBench on our health AI's safety layer. It scored lower than the bare model.
内容摘要
A 150-conversation HealthBench subset, a safety layer that costs points, one real dose leak we found while measuring, and what moved the score.
综合分类、实体标签与信息来源匹配