White Rabbit: server security audits in Claude Code
I periodically review my production server and development environments: SSH settings, listening ports, login history, and package updates. The information comes from several places. I then have to connect it: which service owns a port, whether the firewall covers it, and what the logs show.
I moved some of that work into Claude Code and packaged it as White Rabbit. The plugin contains data collectors, reference material, and instructions for reviewing the results. I wanted a report that helps me decide what to investigate and fix first.
White Rabbit suggests commands for repairs, which I review and run separately. During the audit, a hook checks Bash commands and blocks those that don’t pass its read-only policy.
Running a server security audit
After setting up the plugin, start a full audit in Claude Code:
/wr all user@host
White Rabbit collects server configuration, SSH logs, web access logs, and an inventory of OS packages. It connects findings across those sources and compares them with the previous run when one is available. The results go into a combined report.
You can also run an individual check:
| Command | What it covers |
|---|---|
/wr server user@host | SSH, firewall rules, listening ports, privileged accounts, persistence mechanisms, and Docker |
/wr logs user@host | SSH/auth logs: password guessing, username enumeration, and successful logins from suspicious addresses |
/wr web user@host | nginx and Caddy access logs: sensitive file requests, scanners, and suspicious request patterns |
/wr cve user@host | Known vulnerabilities in installed OS packages |
Run /wr without arguments to see the available checks and guard status. Use a dedicated audit session with the hook enabled: its restrictions apply to Bash commands throughout the session, including ordinary development commands. The local machine needs jq; the hook blocks execution if it’s missing.
Collecting data over SSH
The target doesn’t need a resident agent. Collectors use the utilities already available, such as ss, journalctl, dpkg or rpm, and the firewall’s inspection commands. The script is passed to Bash over SSH. From the plugin directory, collecting the server snapshot looks like this:
ssh user@host 'bash -s' < scripts/collect/server_snapshot.sh
The output has named sections for ports, services, users, SSH settings, and other checks. Local analyzers and Claude work with that format. Missing utilities or insufficient permissions can leave gaps in the audit.
A large access log forced me to change the collector. On a server with 1.3 million requests, it accumulated too much data in shell variables and the SSH session ended with exit code 255. The web collector now processes a stream and keeps the most recent 200,000 lines by default. If it reaches the limit, the report notes that older entries were not analyzed.
Restricting commands
I wanted to keep repairs out of the audit session. White Rabbit’s PreToolUse hook inspects Bash commands before execution, checking forbidden patterns, output redirection, and the commands used in pipelines.
Commands such as rm, chmod, apt install, and firewall changes are blocked. The hook also blocks execution if it can’t parse its input or read the policy files. This is the dependency check in hooks/guard.sh:
if ! command -v jq >/dev/null 2>&1; then
echo "White Rabbit guard: jq not found — failing closed (read-only)." >&2
exit 2
fi
The local analysis scripts need their own exceptions. A filename such as correlate.sh isn’t enough to identify a trusted script: a file with that name could be anywhere. The hook checks canonical paths against the scripts inside the plugin. A different file with the same name doesn’t get that exception.
These are command checks based on rules and patterns, with the limits that implies. The hook isn’t a complete sandbox. It also needs to be active when the session starts; writing “read-only” in a prompt doesn’t enable it.
Reviewing the findings
The useful part for me is connecting evidence that would otherwise sit in separate lists. For example, ufw status can show an active firewall while Docker publishes a container port externally. The report compares firewall settings with published ports so the exposure is easier to spot.
Logs benefit from the same treatment. If an IP address appears in both SSH password guessing and web scanning, White Rabbit connects those events. A successful SSH login from an address associated with attacks gets a high priority. I still need to check the account, time, and session: a matching IP alone doesn’t establish a compromise.
For OS packages, the analyzer looks up CVEs through OSV.dev. Prioritization uses CISA KEV, the catalog of vulnerabilities known to be exploited, and EPSS, an estimate of exploitation probability. This report currently includes only vulnerabilities with an available fix. It therefore doesn’t cover vulnerabilities that have no patch yet. An irrelevant finding can be suppressed with a VEX justification.
Subsequent audits mark findings as [NEW], [UNCHANGED], or [RESOLVED]. That makes it easier to follow up on repairs and spot new issues.
Adding a check
Audit procedures live in skills/, reference material in knowledge/, and collectors in scripts/collect/. Scripts in scripts/analyze/ process the collected data locally.
If the collector already returns the evidence a check needs, I can describe the check in markdown: what to look for, which values to compare, how to assess the risk, and what repair to suggest. The Docker port check in knowledge/checks/docker.md is an example. A check that needs more data also needs a collector change.
Limits
Behavioral tests exercise the collectors, analyzers, and hook against test fixtures. Claude’s interpretation varies with the model. It can miss a connection or overstate a finding, so I need the underlying evidence in the report to review its conclusions.
Logs and configuration are untrusted input too. An HTTP request can contain text that looks like an instruction to the model. The playbooks tell Claude to treat it as data, which doesn’t guarantee protection against prompt injection. Audit data sent to Claude is processed by the model service; that matters when working with a client’s server.
White Rabbit currently covers individual audits of server settings, logs, and OS package CVEs. Source code and application dependency checks are still planned. It doesn’t monitor the server between runs, and I review and apply suggested repairs separately.
Repository: github.com/IvanShishkin/white-rabbit.