← Back to feed
research
Unit 42
Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
New research reveals that AI safety refusal lives in a thin neural layer, highlighting the critical need for external, multi-layered security. The post Pertu...
Read the full story
Unit 42 →