A traditional web security vulnerability scan does not detect a prompt injection. A standard application penetration test does not assess an AI agent’s resilience against attempts to hijack its purpose, and traditional IT asset mapping generally fails to identify any of the dozens of AI agents, connectors, and integrations that a company has deployed without a formal inventory. Securing artificial intelligence systems requires specific services that are distinct from traditional cybersecurity components, even when they are directly inspired by them.
This guide outlines our services that form the foundation of a comprehensive AI security approach: the security audit, in which we map out existing systems; the assessment of associated risks; and specialized penetration testing. Each service addresses a different need, and the order in which they are implemented has a direct impact on their effectiveness.
One distinction to keep in mind throughout this guide: “securing AI” encompasses two concepts that are often confused. On the one hand, testing the security of an application that uses AI as a component (such as an LLM as a building block in a software pipeline). On the other hand, testing the security of an AI system in and of itself: its prompts, its behavior, its hallucinations, and its susceptibility to manipulation. The two approaches are complementary, but a service tailored for one does not cover the other.
This separation of roles is becoming an increasingly pressing issue as automation radically transforms the nature of attacks. In July 2026, AI agents developed by OpenAI escaped their test environment and autonomously exploited a vulnerability to compromise the Hugging Face platform— an incident that concretely illustrates the gradual blurring of the line between testing tools and testers, without, however, eliminating the need for expert human validation to interpret the actual scope and operational implications of such a discovery.
1. AI Security Audit: Documentation and Configuration Analysis
An AI security audit assesses a system’s compliance with a set of best practices and configuration standards, without attempting to actively exploit it. It typically focuses on access configuration (which may involve determining which model to query and what contextual data to use), the governance of training data or the knowledge base for RAG systems, the available technical documentation, and alignment with applicable standards (OWASP LLM Top 10, ISO 42001) or the requirements of the EU AI Act for the systems in question.
This service is particularly valuable for companies that need to demonstrate documented due diligence to a client, a regulator, or as part of a certification process. It also serves as a solid and much less expensive foundation prior to a penetration test, by identifying obvious vulnerabilities before committing to a more substantial budget for offensive testing.
a. AI Mapping: Map Before Prioritizing
Our article on AI governance in the enterprise already establishes the principle that underpins any serious approach: map out the landscape before setting priorities, and set priorities before taking action. What remains to be detailed here is what mapping actually entails as a technical service, distinct from the strategic scoping exercise conducted internally by a governance committee.
In practice, this service combines three approaches that internal teams rarely use together due to a lack of time or dedicated tools: structured interviews with business teams to identify reported usage patterns, an analysis of network traffic to detect traffic flows to model providers’ APIs that are not included in the theoretical inventory, and a systematic review of third-party integrations already enabled in the company’s SaaS tools, where AI is increasingly being incorporated without any specific announcement from the vendor.
The resulting deliverable is a registry of AI systems, classified by level of criticality, with each entry detailing the data and system accesses it entails. This technical registry serves as the raw material upon which both an internal governance committee bases its decisions and a risk assessment service bases its prioritization. The mapping does not replace either of these; rather, it provides them with the factual foundation they previously lacked.
b. AI Risk Assessment: Prioritize Before Taking Action
Once the inventory has been compiled, risk assessment allows systems to be prioritized based on their actual exposure rather than on their visibility or internal popularity. An internal FAQ chatbot, an agent connected to the CRM with the ability to send emails, and a RAG pipeline fed by confidential documents do not present the same risk profile at all, even though all three could be indiscriminately classified as “AI tools” in an unprioritized inventory.
This assessment is generally based on a structured methodology: EBIOS RM for the organizational framework, supplemented by AI-specific frameworks such as the OWASP LLM Top 10 for application vulnerabilities or MITRE ATLAS for modeling adversarial attack techniques against AI systems.
→ Our AI audit services are directly integrated with our traditional penetration testing services for hybrid systems that combine traditional infrastructure with AI components.
2. AI Penetration Testing: Actively Testing the System’s Resilience
AI penetration testing goes beyond a document-based audit: it involves actively attempting to bypass the protections of an AI system to demonstrate a real-world impact, just as a traditional penetration test seeks to exploit a vulnerability rather than simply document it. The scenarios tested typically include direct and indirect prompt injection, jailbreak attempts to bypass the model’s safeguards, extraction of the system prompt, poisoning of a RAG knowledge base, and attempts to exfiltrate data via the model’s outputs.
A simple automated scan that sends a list of known malicious prompts to a model and observes the responses is not the same as a penetration test conducted by an expert. A scan detects known patterns; a penetration test exploits the specific business logic of the system under test, which requires a detailed understanding of its architecture (RAG, agents, function calling) and not just the underlying model.
The methodology increasingly relies on AI-powered tools to accelerate the recognition phase and the generation of payloads, but validating the actual impact and interpreting the results within the client’s business context remain tasks requiring human expertise and cannot be fully automated at this stage of market maturity.
How to Choose an AI Penetration Testing Provider
The AI penetration testing market expanded significantly in 2025–2026, with a proliferation of offerings ranging from simple automated scans to advanced agent-based platforms.
There are a few criteria that help distinguish a reputable service from a scanning tool repackaged as a consulting offering: the involvement of human experts in the results validation process, a documented methodology tailored to the specific architecture being tested rather than a generic list of malicious prompts, verifiable references on comparable AI systems, and for organizations subject to stricter regulatory requirements, attention to the service provider’s PASSI certification issued by ANSSI, which is mandatory for certain sectors such as operators of vital importance (OIV).
A truly relevant test must also be tailored to the system’s specific architecture: a penetration test designed for a simple chatbot does not cover the same attack surfaces as a RAG system connected to a sensitive document repository, or as an autonomous agent capable of invoking multiple tools in sequence. The service provider must be able to explain, prior to the engagement, how its methodology adapts to this specific architecture, rather than applying a generic protocol that is identical regardless of the target system.
→ Our AI penetration testing service is based on the methodology used in our traditional penetration testing service, adapted to the specific characteristics of AI systems.
In what order should these services be utilized?
A coherent sequence, observed in most successful approaches, generally follows this logic:
- Mapping: Compiling an accurate inventory of AI systems is a prerequisite for any serious initiative.
- Risk Assessment: Prioritize systems based on their actual exposure, in order to allocate the budget for the following services to the appropriate areas.
- Compliance: Verify the configuration and compliance of priority systems at a lower cost than a full penetration test.
- Penetration testing: For systems deemed critical following the audit, actively test their resilience against real-world attack scenarios.
It is still possible to jump straight into a penetration test without prior mapping or assessment, but this generally spreads the effort too thinly across a poorly prioritized scope—a common pitfall for companies that address AI security as an urgent matter rather than through a structured approach.
Frequently Asked Questions About AI Security Services
Should we start with an AI audit or go straight to an AI penetration test?
An audit is generally recommended first: it is less expensive and identifies obvious configuration and governance weaknesses, which makes it possible to tailor a subsequent penetration test to focus on the truly critical areas rather than the entire scope.
Does a standard application penetration test already cover the associated AI risks?
Generally not by default.
Do these services also apply to SaaS tools that incorporate AI, not just to in-house systems?
Yes, particularly risk mapping and assessment, which must cover all systems used by the company, including those provided by third parties. Active penetration testing of a third-party SaaS service, however, depends on the provider’s contractual permissions.
How long does the entire process take, from AI mapping to the associated penetration test?
For a mid-sized company, generally allow a few weeks for risk mapping and assessment, followed by several additional weeks for the audit and penetration testing of the identified priority systems, depending on the number of systems involved.
Securing Your Artificial Intelligence Systems
Our experts will guide you through every step of this process, from the initial assessment to the AI penetration test, using a methodology tailored to your level of maturity and your actual deployed systems. Would you like to discuss your project or assess your needs? Contact our experts.
