Cisco Systems has released details regarding its AI Defense platform, highlighting its performance in an independent benchmark conducted by ML6. The evaluation focused on the system's ability to secure artificial intelligence applications that operate through complex, multilingual, and multi-turn conversations. According to the company, traditional guardrails that analyze only single-turn English prompts are insufficient for modern enterprise AI use cases.
In the ML6 test, which analyzed 80,000 Dutch-language prompts, Cisco AI Defense recorded the highest F1 score among the providers evaluated. The benchmark included scenarios such as prompt injection, policy bypass attempts, ambiguous instructions, and realistic enterprise interactions. Cisco reported an F1 score of 0.845 for this specific cohort.
The company noted that this Dutch-specific dataset was assembled independently and is not directly comparable to other evaluations.
Cisco attributes its performance to a technical approach that distinguishes between user intent and content. The company states that its system uses constitutional definitions to provide precise operational specifications for classification. This method, Cisco claims, reduces inter-model disagreement by up to 57 times compared to paragraph-level definitions. The system is designed to apply these specifications with equal precision across languages including French, Japanese, and Arabic.
The company also shared results from an evaluation using an augmented version of the LMSYS Chat-1M and WildChat datasets. This evaluation covered eight additional languages and consisted of approximately 5,800 to 5,900 conversations per language. The dataset included predominantly benign general-purpose conversations alongside an adversarial subset representing roughly 14 percent of the labeled examples.
Cisco stated that the false positive rate measured in this specific mix would be significantly lower in a real-world distribution.
Cisco research across 15 frontier models found that every model tested showed meaningful vulnerability to multi-turn attacks. The company argues that real adversaries often iterate and reframe refusals across multiple turns, meaning attack success rates do not consistently correlate with single-turn benchmarks.
