ARMOR++: Multi-Primitive Transfer Attacks Against Deepfake Detectors via Multi-Agent Orchestration
Deepfake detectors fail under black-box adversarial transfer due to their reliance on fragile architectural cues. We present ARMOR++, a multi-agent framework for high-transferability deepfake evasion. The framework leverages Qwen2.5-VL for spatial-semantic priors and Qwen3 LLM for orchestrating primitive selection, hyperparameter-adaptive reparameterization, and entropy-regularized perturbation mixing. By synthesizing five complementary primitives—dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structure modifications—ARMOR++ targets the diverse inductive biases of heterogeneous detectors. Rigorous evaluation on the AADD-2025 benchmark demonstrates that ARMOR++ significantly outperforms existing proxy and proxy-free baselines in both low-quality and high-quality image regimes. Statistical analysis confirms substantially higher blind-target attack success rates (ASR) compared to state-of-the-art proxy baselines, along with advantages on non-proxy benchmarks and robust defense configurations. These findings highlight a significant reliability gap in current deepfake detector deployments and validate the effectiveness of agent orchestration for uncovering latent vulnerabilities.
Background and Context
The security and reliability of deepfake detection technologies remain central concerns within the artificial intelligence safety domain, yet existing detection models exhibit significant vulnerabilities when subjected to black-box adversarial transfer attacks. This fragility primarily stems from the detectors' heavy reliance on specific, brittle forensic cues embedded in their architectural design. Consequently, adversarial perturbations generated by surrogate models, such as Convolutional Neural Networks (CNNs), often fail to transfer effectively to target models based on Transformer architectures. This architectural mismatch creates a substantial gap in attack efficacy, limiting the ability to evaluate the true robustness of deployed systems. Furthermore, current methodologies generally lack semantic awareness, struggling to maintain attack effectiveness under strict query-free constraints. The absence of high-level semantic understanding means that many existing attacks operate purely at the pixel level, ignoring the structural and contextual integrity of the image, which is increasingly critical as detectors evolve to incorporate more sophisticated visual-language reasoning.
To address these critical limitations, the ARMOR++ framework has been introduced as a novel multi-agent orchestration system designed to achieve high-transferability deepfake evasion. This research represents a paradigm shift from static, single-modality attacks to dynamic, semantically aware adversarial generation. The core contribution lies in the construction of an automated system capable of comprehending image spatial semantics and dynamically adjusting attack strategies. By integrating the reasoning capabilities of Large Language Models (LLMs), ARMOR++ enables the intelligent selection and combination of attack primitives. This approach not only resolves the failure of traditional methods in cross-architecture transfer scenarios but also significantly enhances the robustness and generalizability of the attacks. The framework exposes deep-seated defects in current detection systems, demonstrating that reliance on isolated feature detectors is insufficient against coordinated, multi-domain adversarial strategies.
Deep Analysis
ARMOR++ employs an innovative multi-agent orchestration architecture that deeply integrates the strengths of Vision-Language Models (VLMs) and Large Language Models. Specifically, the framework leverages the Qwen2.5-VL VLM to extract spatial-semantic priors from the input images. This step is crucial as it allows the attack process to understand the semantic structure of the image content rather than focusing solely on pixel-level noise. By grounding the adversarial perturbations in semantic context, the framework ensures that the modifications remain imperceptible to human observers while effectively disrupting the detector's logic. The Qwen3 LLM then acts as the central orchestrator, coordinating the activities of various agents. It is responsible for selecting the most appropriate attack primitives, performing hyperparameter-adaptive reparameterization, and executing entropy-regularized perturbation mixing. This hierarchical structure allows for a sophisticated balance between attack strength and stealth, adapting in real-time to the feedback received during the optimization process.
To cover diverse detection blind spots, ARMOR++ synthesizes five complementary attack primitives: dense optimization, saliency-based methods, spatial transformations, frequency-domain perturbations, and block-structure modifications. This multi-primitive design enables the framework to target the diverse inductive biases of heterogeneous detectors. For instance, frequency-domain perturbations are specifically designed to disrupt frequency-based feature extractors, while spatial transformations challenge the spatial invariance assumptions of other models. By combining these diverse techniques, ARMOR++ can effectively bypass detectors that rely on different underlying mechanisms. The integration of these primitives is not arbitrary; it is guided by the LLM's strategic planning, which evaluates the effectiveness of each primitive in the current context. This dynamic selection process ensures that the attack remains effective even as the detector's defense mechanisms evolve, highlighting the limitations of static defense strategies in the face of adaptive adversarial threats.
Industry Impact
The rigorous evaluation of ARMOR++ on the AADD-2025 benchmark demonstrates its significant superiority over existing proxy and proxy-free baselines in both low-quality and high-quality image regimes. The results indicate that ARMOR++ achieves substantially higher Attack Success Rates (ASR) in blind-target attack settings, where the attacker has no knowledge of the target detector's architecture or parameters. This finding is particularly alarming for the industry, as it suggests that current deployment models are highly susceptible to sophisticated attacks that do not require prior knowledge of the system. Statistical analysis confirms these advantages, showing that ARMOR++ outperforms state-of-the-art proxy baselines across various scenarios. The framework's ability to maintain high ASR under robust defense configurations further underscores the fragility of existing security measures. These results highlight a significant reliability gap in current deepfake detector deployments, warning the industry that reliance on single-architecture or single-feature detection models is no longer sufficient to ensure content integrity.
Moreover, the findings from ARMOR++ have profound implications for the future development of deepfake detection technologies. The research reveals that current mainstream detectors have significant vulnerabilities when facing advanced adversarial attacks, necessitating a rethinking of detection architecture. The multi-agent orchestration framework proposed by ARMOR++ offers a new research paradigm for adversarial machine learning, proving that combining the semantic understanding capabilities of LLMs with the perceptual abilities of vision models can greatly enhance the intelligence of attacks. This, in turn, drives the development of more robust detection algorithms, forcing researchers to consider semantic and structural aspects in detector design. The potential for open-sourcing and the scalability of the ARMOR++ framework provide a solid foundation for subsequent research, helping the community explore safer AI-generated content detection mechanisms. Ultimately, this work serves as a critical wake-up call for the industry, emphasizing the need for comprehensive, multi-layered security strategies in the era of generative AI.
Outlook
Looking forward, the implications of ARMOR++ extend beyond mere academic curiosity, serving as a stress test for the entire digital media ecosystem. As generative AI tools become more accessible, the volume of synthetic media will continue to grow exponentially, making reliable detection not just a technical challenge but a societal imperative. The success of ARMOR++ in bypassing current detectors suggests that future defense mechanisms must be inherently more robust, potentially incorporating real-time adversarial training and multi-modal verification. The integration of LLMs in both attack and defense roles indicates a future where AI systems are constantly evolving in an adversarial dance, requiring continuous updates and rigorous testing. Organizations deploying deepfake detection solutions must therefore adopt a proactive stance, regularly testing their systems against advanced frameworks like ARMOR++ to identify and patch vulnerabilities before they can be exploited maliciously.
Additionally, the research highlights the importance of standardizing benchmarks for adversarial robustness. The AADD-2025 benchmark, while comprehensive, needs to be expanded to include a wider variety of detector architectures and attack scenarios to ensure that new detection models are truly robust. The community must also focus on developing explainable AI techniques for detectors, allowing for better understanding of why certain images are classified as fake or real. This transparency is crucial for building trust in automated content verification systems. As the field progresses, collaboration between researchers, industry practitioners, and policymakers will be essential to establish ethical guidelines and regulatory frameworks that address the risks posed by sophisticated adversarial attacks. The ARMOR++ framework, with its emphasis on semantic-aware, multi-agent orchestration, sets a new standard for what is possible in adversarial machine learning, challenging the industry to rise to the occasion and develop more resilient and trustworthy AI systems.