Skip to main content

Augustus: MCP servers become a fully tested attack surface

Augustus, Guard's AI security testing capability, has taken on an entirely new attack surface: Model Context Protocol servers. MCP servers expose tools, resources, and prompt templates to a model, and nothing in conventional web-application testing reaches them. Augustus now discovers, fingerprints, authenticates against, and adversarially probes MCP deployments end to end — covering the OWASP MCP Top 10 — and delivers each result as a self-contained finding with a CVSS v4.0 vector, a CWE tag, and reproduction steps. Its multimodal coverage has grown alongside: image and PDF attack probes now run against more model families, and scan verdicts finally distinguish an injection that was obeyed from one that never landed.

What's New

New: MCP servers are discovered, fingerprinted, and routed automatically

  • Endpoint auto-discovery — MCP servers are an emerging and frequently unmonitored surface, so for any public domain asset Guard probes the conventional mcp.<domain> host and, when it resolves, emits it as a discovered asset for identification. MCP servers you did not know you had now appear in your attack surface inventory.
  • Julius fingerprinting — an mcp-server probe identifies Model Context Protocol servers (JSON-RPC 2.0 over the Streamable HTTP transport, spec 2025-06-18) and types them as Type = "mcp", distinguishing them from generic LLM inference endpoints.
  • Julius → Augustus, with no manual wiring — typed MCP assets are picked up automatically and the MCP probe pattern is applied; the Augustus generator configuration Julius emits is forwarded intact, preserving probe scope. Vespasian's MCP probe likewise emits a ready-to-use generator config on detection.
  • Agentic workflows too — an orchestrator agent that discovers an endpoint fingerprinted as an MCP server now hands it to Augustus exactly as the deterministic path does, so agentic hunts no longer leave MCP targets untested. The Augustus agent can also invoke MCP probes directly by passing generator_type: mcp.
  • MCP reconnaissance panel — discovered MCP servers surface a dedicated LLM-Recon panel in Guard showing server identity, the tool/resource/prompt catalog, and identity recon results.

A first-class MCP generator and recon layer

  • MCP generator — Augustus speaks MCP (JSON-RPC 2.0) to target servers over stdio, streamable HTTP, legacy SSE, and auto-detected transports, so MCP endpoints are tested the same way REST, GraphQL, and gRPC targets are.
  • Reconnaissance module family — a class of modules gathers target facts as observations rather than verdicts, feeding downstream probes structured context about the exposed tool surface. Recon modules can consume earlier modules' observations, so chained discovery — identifier enumeration feeding straight into authorization tests — needs no manual wiring.
  • Full paginated catalog enumeration — only the first page of a server's tool, resource, resource-template, and prompt catalogs used to be retrieved. All pages are now followed via nextCursor, removing silent under-coverage against servers with large catalogs.
  • Full JSON Schema traversal — probe arguments are built by resolving nested properties, $ref, allOf, anyOf, oneOf, and if/then constructs rather than top-level properties alone. Tools that previously looked covered while generating schema-validation errors are now genuinely exercised, so reported attack surface matches tested attack surface.

What the MCP probes actually test

  • Injection across all three MCP primitives — tool poisoning and tool-parameter injection, resource-content injection (OWASP MCP06, where adversarial instructions embedded in a resources/read response steer the host model), and prompt-template poisoning (OWASP MCP10, probing prompts/get templates for injected instructions — the indirect-injection class behind real-world GitHub, Supabase, and Atlassian incidents). Resources and prompt templates are model-facing surfaces that went entirely untested before.
  • OS command injection (OWASP MCP05) — the most prevalent vulnerability class in deployed MCP servers, at roughly 43% ecosystem prevalence. Shell-metacharacter payloads are injected into tool inputs and detected both in-band and blind, via out-of-band callbacks, catching sinks that return nothing to the client. Directly covers the CVE-2025-6514 and CVE-2025-53355 patterns that the earlier arithmetic-canary probe missed.
  • Broken Object-Level Authorization — requests referencing object identifiers belonging to other users or sessions surface servers that trust caller-supplied IDs without re-validating ownership.
  • Authentication and authorisation (OWASP MCP07) — probes for unauthenticated access at the HTTP transport layer, weak token validation, missing or bypassable authorisation on tool calls, and privilege-escalation paths across MCP sessions.
  • Path traversal and SSRF — OS-file traversal payloads are injected into path-like tool parameters and reads detected by file-content signature, respecting allowed-prefix declarations in tool descriptions; SSRF probes cover server-side request abuse from tool handlers.
  • Transport-layer attacks — DNS rebinding (coercing tool handlers into requests to attacker-controlled origins by exploiting TTL expiry), SSE session fixation and cross-session data leakage, and Origin-header validation bypass.
  • Credential exposure on every surface — API keys, tokens, database credentials, and cloud-provider secrets are detected in MCP server configurations at rest, in live tool responses, and — newly — in resources, prompt templates, and other non-tool surfaces. A server advertising plaintext credentials in a resource would previously have scored as safe.

Findings you can act on without leaving the page

  • Structured security write-ups — each of the eight MCP probes (Tool Injection, SSRF, BOLA, Path Traversal, Response Leak, Origin Validation, SSE Session Hijack, Credential Exposure) carries a curated description of the issue and its consequences, actionable remediation guidance, CWE identifiers and supporting references, and a conservatively-scored CVSS v4.0 vector. No external lookup required to understand what was found or how to fix it.
  • Reproduction guidance — all eight probes also emit a dedicated Verification field in their RiskInfo: static, reproduction-oriented prose describing exactly how the probe confirmed the vulnerability and the steps to reproduce it manually, kept separate from the finding description so triage and escalation have a clear path.
  • Calibrated language — finding titles, descriptions, and severity justifications were reviewed and corrected to reflect what the evidence actually supports, consistent with pentest report standards, rather than overstating impact.
  • Aggregated Origin findings — Origin-header bypass variants are reported as one consolidated finding per server instead of ten separate rows, removing noise that used to dominate output on benign servers.

Accuracy: fewer missed findings, fewer false ones

  • Tool-surface false negatives resolved — four measured false negatives in the path-traversal, injection, SSRF, and response-leak probes are fixed; the probes were reaching their targets and discarding valid evidence.
  • Measured benchmark gain — coverage on the DVMCP challenge benchmark advanced from 1 of 10 to 4 of 10 solved challenges, before counting further gains from the authentication and credential probes.
  • Cross-probe contamination fixed — detectors are scoped per probe in multi-probe runs, so verdicts from unrelated detectors no longer contaminate results.
  • Authenticated tests that actually authenticate — credentials attached to the asset or integration in Guard are looked up and forwarded to MCP and HTTP probes, with per-asset credential scoping and a model-availability pre-flight check before probes execute, hardening runs against silent misconfiguration.

Broader multimodal attack coverage

  • Image-based attack probes — three paper-faithful probes, including FigStep, each reproducing the original peer-reviewed attack methodology against vision-capable models, reaching attack paths text-only probes cannot.
  • PDF probes on Google models — multimodal document probes previously ran only against Anthropic models. Native document content blocks are now wired through both Google generators, so pdf.* probes execute against Gemini and Vertex AI instead of silently dropping the document payload.
  • A four-way scan verdict — results now distinguish no injection, injection attempted but resisted, injection obeyed with a benign canary, and injection obeyed with harmful effect. A visible-channel injection the model obeys scores 0.5 and renders as its own intermediate verdict rather than 0.1/SAFE, so "injection obeyed" is no longer indistinguishable from "nothing happened."
  • Errors reported as errors — probes that fail before reaching the model (auth failure, timeout, transport error) are surfaced as errors instead of silent passes, so a broken scan cannot look like a clean result.
  • Newest Claude model support — the Anthropic generator no longer sends a hardcoded temperature parameter, enabling scans against claude-sonnet-5 and claude-opus-4 class models that rejected the field.
ImprovedCapability