LLM Attack Surface
How Guard finds the LLM and AI services in your environment and tests them for prompt injection and related risks.
LLM Services and Their Security Blind Spot
Organizations are deploying LLM applications and self-hosted model servers faster than they can secure them, and in many cases, faster than they even know about. The result is an attack surface that traditional security tools were never designed to detect, let alone test.
Why LLMs Are a New Class of Attack Surface
The architecture is fundamentally different
Traditional web applications have a clear separation between code and data — your application logic is in source files, your data is in a database, and your inputs flow through well-defined parameters. Security tools built over the last two decades exploit this separation: WAFs inspect HTTP parameters, SAST tools analyze source code, and DAST scanners fuzz input fields.
LLMs break this model entirely. In an LLM application, every input is both code and data simultaneously. A user prompt can contain instructions that the model interprets and follows. There is no clear boundary between "what the user said" and "what the system should do." This is why prompt injection — instructing the model to ignore its system prompt and follow attacker instructions instead — is fundamentally different from SQL injection or XSS. It exploits the architecture, not a bug.
WAFs can't parse semantic attacks. SAST tools don't understand prompt manipulation. DAST scanners miss indirect injection entirely. The entire traditional security toolchain has a blind spot exactly where organizations are deploying their newest and most sensitive applications.
Shadow AI is already inside your network
Employees run local Ollama instances, teams deploy open-source chat interfaces, and business units stand up RAG pipelines — often outside any governance framework.
This isn't negligence. It's the natural result of AI tools being extraordinarily easy to deploy. A developer can have a fully functional LLM service running on a corporate laptop in under five minutes. An engineering team can deploy an internal chatbot to a Kubernetes cluster in an afternoon. None of these deployments require security team involvement.
Without discovery, you have no inventory of where these services run, what they connect to, or what data they expose.
The vulnerability categories
The OWASP Top 10 for LLM Applications identifies ten critical vulnerability categories that every LLM deployment faces:
What Guard Does About It
Guard addresses the LLM attack surface with a two-stage pipeline: discover what's running, then test it for vulnerabilities.
Stage 1: LLM Service Discovery
The first challenge is visibility. You can't secure LLM services you don't know exist.
Guard continuously scans your external attack surface for LLM services running on HTTP/HTTPS endpoints. When it finds one, it identifies exactly what's running — not just "there's a web service here" but "this is an Ollama instance serving Llama 3 and Mistral models."
32 LLM platforms detected, organized by category:
Self-Hosted LLM Servers (15) Ollama, vLLM, LocalAI, llama.cpp, Hugging Face TGI, LM Studio, Aphrodite Engine, FastChat, GPT4All, Gradio, Jan, KoboldCpp, NVIDIA NIM, TabbyAPI, Text Generation WebUI
Gateway and Proxy Services (3) LiteLLM, Kong AI Gateway, Envoy AI Gateway
RAG and Orchestration Platforms (13) AnythingLLM, AstrBot, BetterChatGPT, Dify, Flowise, Hugging Face Chat UI, LibreChat, LobeHub, NextChat, Onyx, Open WebUI, SillyTavern
Cloud-Managed Services (1) Salesforce Einstein
Plus a generic OpenAI-compatible API fallback that catches any service implementing the OpenAI API standard.
How detection works:
Guard doesn't rely on simple banner matching. Each platform has a purpose-built detection probe that sends protocol-specific HTTP requests to service-characteristic endpoints and analyzes the responses against multi-factor signature rules — status codes, response body content, HTTP headers, and content types. A confidence scoring system (1-100) ensures high-specificity matches rank above generic detections.
When a service is identified, Guard also extracts the available models — querying the service's model listing endpoint to discover exactly which models are deployed (e.g., "llama3.1:70b", "gpt-4-turbo", "mistral-large").
Every discovered LLM service is registered as an asset in Guard with its service type, available models, and the API interaction parameters needed for the next stage.
Stage 2: LLM Vulnerability Scanning
Once Guard knows what LLM services are running, it tests them for real security weaknesses.
Guard's LLM vulnerability scanner executes 210+ adversarial attack probes across 47 categories against discovered services.
Vulnerability categories tested:
How scanning works:
The scanner authenticates with each discovered LLM using the API parameters collected during discovery. It submits adversarial prompts and analyzes the model's responses using multiple detection strategies — pattern matching, classifier-based scoring, and LLM-as-judge evaluation. Each finding includes the specific prompt that triggered the vulnerability, the model's response, and a confidence score.
Findings are categorized by type (jailbreak, prompt injection, data exfiltration, etc.) and surfaced as risks in Guard with full proof data — so your security team can see exactly what worked, why it matters, and what to fix.
The Discovery-to-Testing Pipeline
External Attack Surface Scanning
│
├─→ Port Scanning ─→ HTTP/HTTPS services discovered
│ │
│ └─→ LLM Service Discovery ─→ Service identified (e.g., Ollama on port 11434)
│ ├─→ Platform detection (32 services, confidence scored)
│ ├─→ Model enumeration (which models are deployed)
│ └─→ API interaction parameters extracted
│ │
│ └─→ LLM Vulnerability Scanning
│ ├─→ Jailbreak testing ─→ Safety bypass found
│ ├─→ Prompt injection ─→ Instruction override found
│ ├─→ Data exfiltration ─→ Leakage path found
│ ├─→ System prompt extraction ─→ Credentials exposed
│ └─→ Format exploits ─→ XSS via LLM output found
│
└─→ Risk created per vulnerability category
└─→ Proof includes: attack prompt, model response, confidence score
This pipeline runs automatically as part of Guard's external attack surface scanning. When a new LLM service appears on your perimeter, Guard discovers it, identifies it, and tests it — without manual intervention.
What Users See in the Platform
LLM Service Inventory
Every discovered LLM service appears as an asset with:
- Service type (Ollama, vLLM, Dify, OpenAI-compatible, etc.)
- Available models and their names
- Confidence score for service identification
- Associated endpoint and port information
LLM Vulnerability Findings
Each vulnerability finding includes:
- Category — Jailbreak, prompt injection, data exfiltration, etc.
- Severity — Based on the type and impact of the vulnerability
- Proof — The exact prompt that triggered the vulnerability and the model's response
- Confidence — How certain the detection is (0-1 score)
- Remediation context — What the finding means and how to address it
Findings follow the same lifecycle as all Guard risks: Triage → Open → Remediated/Accepted.
Summary
Guard gives security teams:
- Discovery — Find every LLM service running on your attack surface, including the ones nobody told you about
- Identification — Know exactly what's running: the platform, the models, the API surface
- Testing — Validate that each service can resist the actual attack techniques being used in the wild
- Continuous monitoring — Detect new LLM deployments as they appear, not months later during an audit
More in Attack Surfaces
Application Attack SurfaceCICD Pipeline SurfaceCloud Attack SurfaceExternal Attack SurfaceStill need help? Ask the team