Request, Aggregate, Bypass: How Attackers Can Evade LLM Safety Classifiers

CrowdStrike’s research team shows that decomposing a harmful task into individually benign subtasks can systematically bypass the safety classifiers guarding frontier AI models, producing working offensive code across 9 of 10 tested attack categories. Microsoft Research independently discovered the same technique, which it calls “Capability Laundering.”

Source: www.crowdstrike.com

Curated Permalink