Hannibal: autonomous penetration testing you run on your own terms

Hannibal, Guard's autonomous penetration testing capability, has gone from an External-and-Cloud hunt agent to something a security team runs on its own terms. Web applications and LLM endpoints are first-class attack surfaces now, and hunts against them can authenticate, so the agent tests what a logged-in user actually reaches rather than the login wall. You set how aggressively a hunt pivots, what it may spend, when it is allowed to run, and what tag its findings carry — all self-service, without routing through Praetorian. Hunts also know when to stop: a quality backstop scores finding severity over a trailing window and completes the run when the bar is no longer met. And the whole workflow is now drivable from the terminal.
What's New
New: web application and LLM attack surfaces
- Web Applications — the Launch Hannibal dialog offers Web Applications alongside External and Cloud, scoping the run to the
webapplicationasset class and executing a purpose-built web hunt pipeline that is separate from the external-network and cloud hunt logic. - LLM endpoints — LLMs are a first-class surface option, scoping the hunt to assets that prior discovery classified as carrying an LLM or AI attack surface and routing to a dedicated LLM red-teaming workflow rather than the generic web or cloud path. Risks found during an LLM hunt are correctly tagged with the LLM attack surface and appear under surface-filtered views.
Credentialed hunts: past the login wall
- Authenticated web application hunts — an unauthenticated scan only ever sees the login wall, and that is not where web application risk lives. Web Application hunts accept static credentials, basic login, SSO, and TOTP, threading them into the hunt so the agent tests the authenticated attack surface.
- Credentialed LLM hunts — API keys or session credentials can be supplied for LLM surfaces, so credential-gated endpoints are tested rather than skipped.
- Credentials that stay available mid-hunt — recorded application credentials are resolved from the
HAS_CREDENTIALgraph relationship rather than from the credential pin alone. Agents previously fell back to the Guard tenant email even when a valid recorded credential existed, which madewebauth_replaysilently disappear from the toolset partway through a hunt; the tool now stays available whenever credentials are present and valid.
You decide how hard the agent pushes
- Aggressiveness dial — choose Cautious, Balanced, or Aggressive when launching or scheduling a hunt. Cautious caps dispatch breadth; Balanced and Aggressive differ in pivoting depth and attack-chain length. It is a dedicated, explained step in the Launch Hannibal wizard, each level presented as a card with a description, and it applies equally to on-demand and scheduled hunts.
- Attack-path recording — the higher aggressiveness tiers record the full attack path taken, so results surface multi-step exploit chains rather than isolated single findings.
Self-service scheduling for autonomous penetration tests
- No Praetorian in the loop — launching or scheduling an autonomous penetration test previously meant routing through Praetorian, and the workflow sat behind super-admin budget steps. Customers now configure and schedule Hannibal pentests directly from AI Settings — hunt mandates, guardrails, severity targets, sources, and scan profiles — on their own timeline.
- Purpose-built settings cards — hunt scheduling and AI budget configuration are separate cards. The Schedule Autonomous Penetration Test card carries the scheduling toggle and hunt profiles with deep-link support via URL parameters and expands in place rather than opening a modal; Category Allocation is its own card with its own edit modal, available to manage-settings users independently of the budget cap; the AI Spend card is slimmed to the monthly cap and usage meter, with a budget-only, super-admin-gated edit modal. The Automatic Findings Validation (Cato) card received the same accordion treatment, so both AI-driven capabilities configure consistently.
- Schedules with a full API — per-tenant schedules associate an asset target with a capability (Hannibal or Constantine), a cadence of daily, weekly, or monthly, and an optional day and time in UTC. Manage them through
GET,POST,PUT, andDELETEon/hunt/schedules. Schedules are stored as a tenant setting, so each team independently controls scope, agent type, and the prompt used for auto-created hunts. A "Hunt scheduling" step appears in the AI spend configuration wizard when a hunts budget allocation is set, wiring budget to cadence. - A reliable off switch —
hunt_schedule.enabledset to false previously did not stop the auto-hunt cron from firing and consuming budget. It now gates all automatic pentest activity, both scheduled auto-hunts and emergent-threat kickoffs, while manual hunt creation remains available. It defaults to enabled with no custom profiles required, so existing configurations are unaffected. - Scan-window compliance — automated hunts could start iterations outside a tenant's configured scan launch window, undermining the control customers use to limit when offensive activity runs. Auto-hunt-created hunts are now marked automated and every iteration is gated to the scan window; ticks falling outside it are dropped rather than parked, so hunts do not stall. Operator- and chat-initiated runs keep their ability to bypass the schedule, and parked jobs no longer appear as in-flight, so the hunt queue reflects actual activity.
Spend controls at the level of a single run
- Surge budget overlay — fund a hunt or a Marcus session from a Praetorian-managed surge budget that runs independently of the customer's monthly credit pool. Select it from the Hannibal launch panel or the Marcus composer budget dropdown; surge spend is COGS-accounted and never drawn from the monthly pool.
- Per-run budget cap — set an optional spending ceiling on an individual hunt at launch. The hunt stops when the cap is reached regardless of findings or other finish criteria, so a single run cannot exhaust an allocation.
Hunts that know when they are done
- Quality-based completion backstop — rather than running indefinitely or stopping only on a time or budget limit, a hunt scores a trailing 8-hour window after an initial grace period using severity weights (critical 10, high 7, medium 2) and marks itself complete when it falls below a two-medium-equivalent threshold. Hunt results stay meaningful and operator attention is freed.
- Distinct completion states and notification — hunts completed by the backstop are distinguished from user-initiated stops, and operators receive a Chariot Slack alert when the backstop fires.
- No finding bypasses triage — all triage classes, not only high and critical, automatically spawn a Cato triage job after filing, so every finding is automatically evaluated regardless of its initial classification.
Finding a hunt's results afterwards
- Per-hunt custom tags — the Finish criteria & duration step of the launch wizard takes an optional Custom Tag (for example
Q4_Application_Hunt). It threads from the launch payload through the Hunt model onto every finding the hunt files, alongside the built-inhuntandagent-reportedtags, and is filterable in the Assets and Risks views through the existing tag query controls. - Hunt-tag asset labeling — assets selected for a hunt are tagged automatically, so the scope of any given run is easy to filter and review.
- The hunt drawer shows everything — the drawer defaulted to Demonstrated findings only, so a hunt with no compromises yet looked like it had returned nothing at all. It now lists every finding associated with the hunt across all lifecycle states — Demonstrated, Detected, and the rest — sorted by criticality, with the most recent item ranked highest within each severity tier. The spurious "No Vulnerabilities match your current filter" empty state is gone.
A more reliable agent
- No more fabricated-evidence storms — hunts could silently burn up to a third of their runtime retrying evidence IDs the model had invented from memory, filing zero findings on labs it had already solved. A
list_evidencetool now hands the agent the ground-truth set of collected evidence IDs before it callsreport_new_risk, eliminating a cycle that reached roughly 100 failed calls per run, and unresolvable evidence references now fail fast instead of amplifying one fabrication into a storm.
New: full Hunt management from the terminal
- Interactive launch wizard —
praetorian-clilaunches External, Internal, Cloud, Web Application, and LLM Application hunts with surface-specific scope selection, mandate configuration, aggressiveness, duration, model tier, and run-scoped credentials. - Live hunt chat and steering — a continuously refreshed agent conversation view with an inline guidance composer, full-history pagination, expandable tool-call detail, Markdown rendering, and pending-interaction indicators.
- Memory browser and editor — navigate, open, edit, save, and delete hunt memory items in a persistent keyboard-driven interface; system-owned entries stay read-only.
- Approvals and credential requests — list and watch pending human-in-the-loop interactions across root iterations and nested subagents, render endpoint approval context, and answer credential requests through the secure broker flow.
- Scheduled hunt profiles — list, inspect, create, edit, pause, resume, and delete reusable scheduled hunt profiles, kept clearly separate from generic capability schedules.
- Cost and operational overview — projected cost in USD, remaining time, root agent and iteration counts, highest vulnerability severity, and scope summary are integrated into hunt status output.
- Workflow deep-links — jump from an agent-backed workflow step straight to its hunt conversation and back to the same step, matching the web UI's workflow-to-chat navigation.

