TutorialsOrdinary
We ran HealthBench on our health AI's safety layer. It scored lower than the bare model.
Summary
A 150-conversation HealthBench subset, a safety layer that costs points, one real dose leak we found while measuring, and what moved the score.
CategoryAI Tutorials & Practice
TierOrdinary
Published
Indexed by AIQB
SourceDEV Community
AIQB record IDintel-1ef7afa4039d04c8c084b4e2