0 fires this session · driven by live scan:progress
Full arsenal runs by default — every installed tool is available to the agent. Keys below are the whole catalog; they light in real time as tools actually fire. Honest: a key with no fire stays dark, and "installed" comes only from preflight.
Point Arxiis at a target to populate the perimeter — every authorized and discovered host appears here.
⚓ SITREP
STANDBY
Awaiting orders — point at a target and engage.
▌ SYSTEM EVENTS
0
— system events stream (live from the backend) —
🛰️
Getting started
watching the agent work
🧠
LIVE REASONING STREAM
The agent's own reasoning, in its own words — ReAct: reason → tool call → observe. Not code; the agent never writes or runs code itself.
0 steps
No reasoning events yet. Once a hunt is running, each operator's "why I'm about to do X" will stream here live, followed by the tool call it led to and the observation that came back.
🎯 START A ZERO-DAY HUNT
Point Arxiis at any authorized target and watch it work: live recon with real tools (nmap, DNS, HTTP probes), every finding provenance-gated to the command that produced it — no phantom flags. An LLM reasons over that ground truth in the open. Kill-chain phases past recon are labeled for what they are — no fake pwns, no vibes. Don't take our word for it: run npm run verify-claims and re-derive every number yourself. Only test what you own or have written permission to assess.
⚡Run keyless — connect Claude Code, Codex, or Hermes and Arxiis drives missions through your agent's own login (no API key). To run, connect a local agent or add a key. Set up in Settings →
No per-step approval · live recon, scaffolded kill-chain · every step receipted
Connect an already-authed CLI (Claude Code · Codex · Hermes) or add an API key to start hunting.
Operator Clarity Layer
Current truth, next move, and proof pressure for the active hunt.
updated --snapshot idle
Operator Inbox0 open
Gate Snapshothold
Proof Stateledger
Mission Spine
Mission readiness — flight recorder, target intake, and authorization.
Flight Recorderstandby
Evidence Workbench0 items
Work Order Board0 open
Target Intakescope
Training Rangelocal safe
Scope Receipts0 pending
No pending scope receipts.
Authorize Targetno active authorization
Grant a time-boxed authorization for one target before Hunt runs against it — a real request→authorize receipt with a live TTL countdown, not a toggle.
PLINY CODE OPS LAYER
Slash grammar, route preview, agent lanes, memory, computer-use, and evidence gates folded into the original Arxiis war room.
Intrusive / credential / dangerous tools stay inert until approved — approve once, then free. Credential & dangerous actions warn on every run; every gated decision is audited.
No tools approved yet.
Waiting for gated tool calls…
Zero-Day Hunt Pulselive · derived from ledgerstandby
Counters are real ledger data. Stage progression is a derived projection of ledger state — not an autonomous hunt engine.
Specialist Swarmlive · ledger0 active
Specialist Swarm — Live Boardlive · pack boardno swarm active
Each card is one deployed operator's live heartbeat. Leads are the swarm's working hypotheses — chips show status, not proof; only confirmed is verified.
No swarm active — Engage to deploy operators.
Swarm Cognition Loopderived · engine not livewarming
Auto-deploys operators, adds a test target, and launches. No API key needed if a local agent is connected.
🌐
testphp.vulnweb.com
Acunetix test site — SQLi, XSS, CSRF
+
🏦
demo.testfire.net
HCL AppScan demo — auth bypass, injection
+
📡
scanme.nmap.org
Nmap official test host — port scanning
+
Target0
No targets
Operators
0
Config
OPSEC Level, Cognitive Mode, and the toggles above tune the client-side reasoning pipeline — backend/keyless missions currently run operators with their configured prompts + defaults.
QualityACTIVE
0
OK
0
REV
0
REJ
0
ESC
0
Targets
0
Creds
0
Access
0/5
Phases
Findings & Loot(0)
ALLCRITHIGHMEDLOWCREDS|
Severity
Type
Finding
Target
Phase
No findings yet. Engage a mission to discover vulnerabilities.
Enter LaunchEsc AbortD Deploy AllK Clear/ Search
Attack Maplive · tool-verified · from findings
The REAL attack map — nodes are this mission's actual findings; a "derived" edge means a concrete artifact an earlier finding surfaced reappears in a later command (a traceable dependency), "sequence" is phase/time ordering. Auto-loads as findings land. SPECTRA (synthetic) is a separate client-side sketch — clearly labeled, never confirmed findings.
no graph
No findings yet — the real attack map populates automatically as this mission produces findings. Built ONLY from actual findings/evidence, never synthetic.
Where each target sits in the kill-chain — every phase's techniques, each with a live status. ✓ done completed · ⏳ active in progress · ○ pending not yet reached. Auto-refreshes while a mission runs.
—
No methodology telemetry yet — coverage populates as the engine advances targets through the kill-chain.
Live Scan
Agent Progress
Operators
0 active
Task Queue
0 tasks
Reasoning And Tool Stream
0 events
ScopeGuard
Scope Receipts
Approval Requests
0 receipts
Active Formation
Custom Deployment
0
Units Active
0
Tools Ready
🚨
0
Critical
⚠️
0
High
ℹ️
0
Medium
🔑
0
Credentials
🔓 Findings
No findings yet. Start a mission to collect evidence.
◉
ARSENAL CONTROL
Full arsenal runs by default — every tool lives in the Tool Launchpad above, lighting up as it fires
85+CATALOG
0ACTIVE
14DOMAINS
Sensor PostureStandby
Arm a few tools to shape the hunt profile.
Visible SetAll tools
Scanning catalog...
Coverage0 lanes
Mix domains for cross-boundary hunts.
HandoffNo operator
Pick an operator before assignment.
◉ ACTIVE LOADOUT — advisory in full-arsenal mode · every installed tool runs regardless. Use the Tool Launchpad above to watch tools fire.
Scopereceipt gate
Breadthneeds domains
Proofneeds validator
Chainneeds composer
No tools assigned - click tools below to add to the loadout
Adversarial Tactics, Techniques & Common Knowledge
TA0001 Initial Access --
TA0002 Execution --
TA0003 Persistence --
TA0004 Privilege Escalation --
TA0005 Defense Evasion --
TA0006 Credential Access --
TA0007 Discovery --
TA0008 Lateral Movement --
TA0010 Exfiltration --
TA0011 Command & Control --
🐛 CWE Top 25
--
Most Dangerous Software Weaknesses (2024)
CWE-787 Out-of-Bounds Write --
CWE-79 Cross-Site Scripting --
CWE-89 SQL Injection --
CWE-416 Use After Free --
CWE-78 OS Command Injection --
CWE-20 Input Validation --
CWE-125 Out-of-Bounds Read --
CWE-22 Path Traversal --
CWE-352 Cross-Site Request Forgery --
CWE-434 Unrestricted File Upload --
📊 Overall Results
🏆
--
Overall Score
✅
--
Tests Passed
⏱️
--
Avg Time (ms)
🎯
--
Accuracy
Web
Binary
Crypto
Reverse
Forensics
Auto Ops
OWASP
MITRE
CWE
🧠 AI Performance Analysis
💪 Strengths
⚠️ Weaknesses
🔧 Recommended Improvements
⚙️ Suggested Config Changes
Optimized Configuration
✨ Applied Configuration Changes
🔄
Analyzing performance data...
🤖
LLM is optimizing your configuration...
Analyzing recommendations and generating optimal settings
⚠ Local simulation / illustrative. The container controls and challenge "solves" here do not touch a live Docker daemon or a real flag server — flags are format-checked, not verified against a live target. For measured, live-exploit-verified results use OBSIDIVM and the benchmarks.
🚀 Quick Setup
Get started with execution-based CTF benchmarks in minutes. Requires Docker installed locally.
1️⃣Check DockerUnknown
Verify Docker daemon is running and accessible.
2️⃣Build ImagesNot Built
Build Docker images for CTF challenges.
3️⃣Launch RangeOffline
Start all challenge containers.
📋 Manual Commands (run in terminal)
cd ctf && docker-compose build && docker-compose up -d
🎯 Challenge Browser
📊 Range Status
Containers Running0
Challenges Solved0 / 8
Total Points0 / 1450
Agent Success Rate--%
🤖 Agent Execution
Initializing agent...0 / 8
📈 Results & History
Challenge
Category
Agent
Time
Status
Flag
🏁
No execution results yet. Run the benchmark to see agent performance.
★ THE ADMIRAL — Autonomous Op Orchestrator
STANDING BY
Give the Admiral a high-level directive and it will autonomously plan the entire operation — identifying targets, allocating operators, setting OPSEC levels, and executing the full kill chain. No manual configuration needed.
Operation Plan
Admiral's Intent
UNROUTED
Mission Gate
WAIT
Targets
Force Allocation
Objectives
Rules of Engagement
Hunt Lanes
Specialist Work Orders
Evidence Contract
Critic & Tool Route
⚓ Admiral Self-Critique
no sitrep yet
No plan yet — the Admiral's adversarial self-critique of its own plan appears here once a plan is generated.
Strongest assumption (attacking its own plan)
Missing coverage the Admiral admits to
Live confidence (from the latest SITREP — subjective, LLM-narrated)—
Admiral's recommended next actions
Mission gate (deterministic, authoritative readiness) is shown in the Mission Gate card — this panel is the Admiral's own narrative self-critique, not a replacement for it.
Admiral's Strategic Rationale
Situation Reports
Awaiting first situation report...
Strategic Assessment
Assessment will be generated after operation completes or on demand.
Stored, but direct-LLM currently routes via OpenRouter or a connected agent — not wired to this key yet.
OpenAI API Key
Stored, but direct-LLM currently routes via OpenRouter or a connected agent — not wired to this key yet.
🔌 Local Agents
Enlist agents you've already authed on this machine — no API keys needed. Arxiis detects each CLI and can drive it as an operator. Tick several and connect them all at once.
🔌 Connect Local Agents
— scanning —
Detecting local agents…
🤖 Model Selection
Select your preferred AI model for agent operations
Loading models...
🔄 Fallback ModelAuto-retries with this model if the primary fails (refusal, timeout, error)
🖥️ API Server
URL of your running Arxiis API server (start with npm run server)
Default: http://localhost:3333 — change if your server runs on a different host/port
⚠️ Data
Config Library
Saved Configurations
0 configs saved •
Default: None
📊 Current Configuration
Not benchmarked
Run a benchmark to see current config performance
💾 Saved Configurations
💾
No saved configurations yet
Run benchmarks and save your best performing configs
Turn the AI coding agent you already run into a red team.
🌩️ Coverage by domain
A domain lights up only when there's a benchmark behind it. Greyed tiles are in development.
🕸️
Web ✅
Apps, APIs, auth flows, OWASP Top 10. XBEN 90.1% pass@1.