Cross-request semantic accumulation (flagging when an API key’s recent requests jointly cover a harmful-composition template) is a natural mitigation to consider. However, it has structural bounds: an attacker can distribute subtask queries across different providers, use local open-weight models for orchestration and assembly, or simply rotate API keys. Cross-request correlation is only effective when the provider has visibility over the full request sequence, which is not guaranteed when the adversary controls the orchestration layer. As Microsoft’s paper observes, “the decisive context stays with the orchestrator” outside the frontier provider’s observation boundary. Per-provider defenses alone cannot fully address cross-provider or hybrid local/cloud attack pipelines. Reasoning about how adversaries distribute requests across providers, rotate keys, and use local models to stay below observation horizons is an adversary tracking problem separate from model lab expertise.

Mitigation Challenges 

This research is part of the CrowdStrike Cyber Superintelligence Lab’s mission to build cyber defense that learns faster than adversaries evolve. As frontier AI models become both the tools and targets of sophisticated attacks, understanding how safety architectures fail at the seams is essential defensive intelligence. As the adversarial coevolution between AI-powered offense and defense is accelerating, research like this ensures defenders see the next generation of multi-model attack pipelines before threat actors deploy them at scale.

Conclusion

Our research maps exactly how wide that failure is. Across 9 of 10 offensive categories, a model can decompose a harmful task into benign subtasks, reframe each as a legitimate software request (game modding, detection engineering, or other dual-use categories), and recompose the outputs into working offensive code. Each fragment is not merely permitted by the classifier; it is genuinely benign.
Model safety evaluations are rigorous within their scope but limited to attacks against single models. Adversaries do not act within these limitations. Our research demonstrates that this is a fundamental vulnerability class in classifier-based LLM safety architectures. Microsoft’s convergent discovery of this technique supports our findings. The defense is not to make these remarkably robust classifiers stricter but to extend the threat model beyond individual requests to encompass request sequences, cross-model composition, and the knowledge transfer dynamics between classified and unclassified model tiers.
Understanding knowledge gaps between models is critical to this bypass technique. For well-documented techniques (reverse shells, persistence), Smaller Model B’s training data is sufficient. For specialized domains such as process injection (Windows API precision), CVE-specific exploits (vulnerability mechanics), privilege escalation (token manipulation), and keylogging (hook APIs), the frontier model’s deeper knowledge is essential. The pipeline is most dangerous precisely where the knowledge gap between model tiers is largest.
Our 515-technique evaluation establishes that the classifier itself is robust; no direct bypass was found across any technique class. The vulnerability lies not in the classifier’s accuracy but in the architectural assumption that per-request evaluation is sufficient. As Microsoft aptly puts it, “alignment that holds over a whole task can fail when the task is split into individually permitted fragments.” 
The classifier works. The architecture around it needs hardening. Within its per-request threat model, the Level 3 classifier is the strongest publicly evaluated defense we have encountered. The gap is structural, not a classifier failure. 

Similar Posts