Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
15 changes: 10 additions & 5 deletions .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -30,7 +30,7 @@ jobs:
- name: linux
host: ubuntu-latest
- name: windows
host: ubuntu-latest
host: windows-latest
runs-on: ${{ matrix.settings.host }}
defaults:
run:
Expand Down Expand Up @@ -74,10 +74,15 @@ jobs:
working-directory: packages/client
run: bun run check:generated

- name: Run HttpApi exerciser gates
- name: Build the V2 CLI with the web app
if: runner.os == 'Linux'
working-directory: packages/cli
run: OPENCODE_CHANNEL=local bun run script/build.ts --single --skip-install

- name: Smoke-test the compiled V2 service
if: runner.os == 'Linux'
working-directory: packages/opencode
run: bun run test:httpapi
working-directory: packages/cli
run: bun run script/service-smoke.ts

e2e:
name: e2e (${{ matrix.settings.name }})
Expand All @@ -88,7 +93,7 @@ jobs:
- name: linux
host: ubuntu-latest
- name: windows
host: ubuntu-latest
host: windows-latest
runs-on: ${{ matrix.settings.host }}
env:
PLAYWRIGHT_BROWSERS_PATH: ${{ github.workspace }}/.playwright-browsers
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,7 @@ UPCOMING_CHANGELOG.md
logs/
*.bun-build
tsconfig.tsbuildinfo
/packages/cli/.cache/

# Internal docs (not for public repo)
docs/architecture-rfc.md
57 changes: 14 additions & 43 deletions .oxlintrc.json
Original file line number Diff line number Diff line change
@@ -1,51 +1,22 @@
{
"$schema": "https://raw.githubusercontent.com/nicolo-ribaudo/oxc-project.github.io/refs/heads/json-schema/src/public/.oxlintrc.schema.json",
"options": {
"typeAware": true
},
"categories": {
"suspicious": "warn"
"correctness": "off",
"suspicious": "off",
"pedantic": "off",
"perf": "off",
"style": "off",
"restriction": "off",
"nursery": "off"
},
"rules": {
"typescript/no-base-to-string": "warn",
// Effect uses `function*` with Effect.gen/Effect.fnUntraced that don't always yield
"require-yield": "off",
// SolidJS uses `let ref: T | undefined` for JSX ref bindings assigned at runtime
"no-unassigned-vars": "off",
// SolidJS tracks reactive deps by reading properties inside createEffect
"no-unused-expressions": "off",
// Intentional control char matching (ANSI escapes, null byte sanitization)
"no-control-regex": "off",
// SST and plugin tools require triple-slash references
"triple-slash-reference": "off",

// Suspicious category: suppress noisy rules
// Effect's nested function* closures inherently shadow outer scope
"no-shadow": "off",
// Namespace-heavy codebase makes this too noisy
"unicorn/consistent-function-scoping": "off",
// Opinionated — .sort()/.reverse() mutation is fine in this codebase
"unicorn/no-array-sort": "off",
"unicorn/no-array-reverse": "off",
// Not relevant — this isn't a DOM event handler codebase
"unicorn/prefer-add-event-listener": "off",
// Bundler handles module resolution
"unicorn/require-module-specifiers": "off",
// postMessage target origin not relevant for this codebase
"unicorn/require-post-message-target-origin": "off",
// Side-effectful constructors are intentional in some places
"no-new": "off",

// Type-aware: catch unhandled promises
"typescript/no-floating-promises": "warn",
// Warn when spreading non-plain objects (Headers, class instances, etc.)
"typescript/no-misused-spread": "warn"
},
"options": {
"typeAware": true
},
"options": {
"typeAware": true
"no-restricted-globals": [
"error",
{
"name": "Reflect",
"message": "Use typed property access or direct invocation. Suppress this rule only for genuine reflection."
}
]
},
"ignorePatterns": ["**/node_modules", "**/dist", "**/.build", "**/.sst", "**/*.d.ts", "**/sdk.gen.ts"]
}
149 changes: 100 additions & 49 deletions CLAUDE.md

Large diffs are not rendered by default.

122 changes: 68 additions & 54 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,8 @@

> **Beta** — it holds up on real engagements and CTFs, but expect rough edges. [File an issue](https://github.com/s0ld13rr/pentestcode/issues) when something breaks; that's what makes it better.

This branch migrates PentestCode to the **OpenCode V2 2.0.18 runtime** (local product version `0.3.0-v2.1`). It retains PentestCode's tools, specialist agents, engagement files, and knowledge. See the [V2 migration guide](specs/v2/pentestcode-migration.md) for compatibility and local validation. The public installation commands below install the published release; build this branch from source to try V2.

## What it does

One instruction in, a full attack chain out:
Expand All @@ -29,15 +31,15 @@ One instruction in, a full attack chain out:
you: "pentest 10.10.10.5, goal is domain admin"
```

| Stage | What the agent does |
|-------|---------------------|
| **Scan** | `nmap -sS -p-` finds 7 open ports and parses the XML straight into engagement state |
| Stage | What the agent does |
| ------------- | -------------------------------------------------------------------------------------------- |
| **Scan** | `nmap -sS -p-` finds 7 open ports and parses the XML straight into engagement state |
| **Recognize** | Ports 88 + 389 → Domain Controller. Fans out three enumerators in parallel (SMB, LDAP, HTTP) |
| **Enumerate** | Null SMB session → writable share. LDAP → user list. Gobuster → web dirs |
| **Attack** | AS-REP roast → crackable hash → first valid credential |
| **Spray** | That credential sprayed across SMB, WinRM, LDAP and RDP on every known host |
| **Exploit** | WinRM foothold → post-exploit agent dumps SAM / LSA / DPAPI |
| **Result** | Domain admin hash in hand — every step recorded with its evidence chain |
| **Enumerate** | Null SMB session → writable share. LDAP → user list. Gobuster → web dirs |
| **Attack** | AS-REP roast → crackable hash → first valid credential |
| **Spray** | That credential sprayed across SMB, WinRM, LDAP and RDP on every known host |
| **Exploit** | WinRM foothold → post-exploit agent dumps SAM / LSA / DPAPI |
| **Result** | Domain admin hash in hand — every step recorded with its evidence chain |

It's methodical where people get lazy: it sprays every credential against every service on every host, and it doesn't forget to check things. Everything it learns lands in a structured state you can query mid-run with `/status`, `/vulns`, or `/creds`.

Expand All @@ -57,20 +59,25 @@ A single self-contained binary — no Bun, Node, or runtime to install. Linux an
<summary>Other options</summary>

**Pin version:**

```bash
PENTESTCODE_VERSION=0.1.7 curl -fsSL https://raw.githubusercontent.com/s0ld13rr/pentestcode/main/install.sh | bash
```

**Custom directory:**

```bash
PENTESTCODE_INSTALL=/usr/local/bin curl -fsSL https://raw.githubusercontent.com/s0ld13rr/pentestcode/main/install.sh | bash
```

**From source:**

```bash
# Bun 1.4.2
bun install
bun run build --single --skip-embed-web-ui
# binary at packages/opencode/dist/pentestcode-<os>-<arch>/bin/pentestcode
cd packages/cli
OPENCODE_VERSION=0.3.0-v2.1 OPENCODE_CHANNEL=local bun run script/build.ts --single --skip-install
# binary at dist/cli-<os>-<arch>/bin/pentestcode
```

</details>
Expand All @@ -80,10 +87,10 @@ bun run build --single --skip-embed-web-ui
```bash
pentestcode auth login # connect your LLM provider
pentestcode # interactive session
pentestcode --prompt "scan 10.10.10.0/24 and enumerate all services" # one-shot
pentestcode run "scan 10.10.10.0/24 and enumerate all services" # one-shot
```

Works with 20+ providers through [ai-sdk](https://github.com/vercel/ai) — Anthropic, OpenAI, Google, Azure, AWS Bedrock, Ollama, and more.
Native V2 provider integrations support Anthropic, OpenAI, Google, Azure, AWS Bedrock, Ollama, and more.

## How it works

Expand Down Expand Up @@ -125,52 +132,59 @@ State survives the session: close the terminal, come back tomorrow, and the agen

## Tools

18 built-in pentest tools beyond bash. Parser tools are mandatory: after running nmap the agent must pipe the output through `nmap_parse` rather than grep the XML by hand, so every finding reaches the engagement state.

| Tool | What it does |
|------|-------------|
| `nmap_parse` | Parse nmap XML → auto-populate hosts/services |
| `nuclei_parse` | Parse Nuclei JSON → create vulns with severity |
| `cme_parse` | Parse NetExec output → update creds/access/hosts |
| `gobuster_parse` | Parse dir brute output → classify findings |
| `bloodhound_parse` | Parse SharpHound JSON → populate AD model |
| `sqlmap_parse` | Parse sqlmap output → extract injection points |
| `xss_detect` | Analyze responses for reflected/stored XSS |
| `jwt_analyze` | Decode JWT, check alg:none/weak HMAC/expiry |
| `cred_spray` | Plan credential spray across all discovered services |
| `scope_check` | CIDR/wildcard scope validation |
| `attack_path_suggest` | Cost-based path finding through the relationship graph |
| `tunnel_manage` | Plan SSH/chisel/ligolo tunnels, track live sessions |
| `phase_control` | Phase management with quality gates |
| `report_gen` | Generate markdown/JSON pentest reports |
| `state_update` | Record findings (30+ mutation types, batch mode) |
| `state_query` | Query engagement state (20+ query types) |
21 built-in pentest tools alongside V2's `shell` and `subagent` tools. Parser tools are mandatory: after running nmap the agent must pipe the output through `nmap_parse` rather than grep the XML by hand, so every finding reaches the engagement state.

| Tool | What it does |
| --------------------- | -------------------------------------------------------------------------- |
| `nmap_parse` | Parse nmap XML → auto-populate hosts/services |
| `nuclei_parse` | Parse Nuclei JSON → create vulns with severity |
| `cme_parse` | Parse NetExec output → update creds/access/hosts |
| `gobuster_parse` | Parse dir brute output → classify findings |
| `bloodhound_parse` | Parse SharpHound JSON → populate AD model |
| `sqlmap_parse` | Parse sqlmap output → extract injection points |
| `xss_detect` | Analyze responses for reflected/stored XSS |
| `jwt_analyze` | Decode JWT, check alg:none/weak HMAC/expiry |
| `cred_spray` | Plan credential spray across all discovered services |
| `scope_check` | CIDR/wildcard scope validation |
| `attack_path_suggest` | Cost-based path finding through the relationship graph |
| `tunnel_manage` | Plan SSH/chisel/ligolo tunnels, track live sessions |
| `phase_control` | Phase management with quality gates |
| `report_gen` | Generate markdown/JSON pentest reports |
| `state_update` | Record findings (30+ mutation types, batch mode) |
| `state_query` | Query engagement state (20+ query types) |
| `task_graph` | Plan dependent tasks, dispatch V2 sessions, and cancel individual children |
| `ensure_tools` | Check availability of required local tooling |
| `os_hook` | Produce platform-specific command guidance |
| `inject_probe` | Build bounded injection probes |
| `knowledge` | Retain effectiveness, false-positive patterns, and engagement reflections |

## Skills

19 curated knowledge packs, loaded on demand so they cost context only when relevant:
Bundled knowledge packs load on demand so they cost context only when relevant:

- **Phase checklists** (6) — what to do in each pentest phase
- **Service knowledge** (9) — SMB, SSH, FTP, DNS, databases, web servers, mail, Docker/K8s, CI/CD
- **Playbooks** (4) — infrastructure, Active Directory, web application, cloud
- **Phase guidance** — reporting and post-exploitation
- **Service knowledge** — SMB, SSH, FTP, databases, web servers, Docker/K8s, CI/CD, mobile and pivoting
- **Web testing** — injection, SSRF, XXE, file upload, traversal, deserialization and authorization
- **Playbooks** — Active Directory, web applications and cloud

Skills are plain markdown. Add your own by dropping a `SKILL.md` into the skills directory — no code changes needed.

## Commands & modes

Drive a live session with slash commands:

| Command | What it does |
|---------|-------------|
| `/status` | Engagement dashboard — hosts, vulns, creds, phase |
| `/targets` | Host & service table |
| `/vulns` | Findings by severity |
| `/creds` | Discovered credentials |
| `/scope` | View/edit target scope |
| `/phase` | Phase management |
| `/mode` | Switch auto / free / guided |
| `/pause` | Pause on findings (never / always / checkpoint) |
| `/report` | Generate a pentest report |
| Command | What it does |
| ------------- | ------------------------------------------------- |
| `/engagement` | Create or select the shared engagement |
| `/status` | Engagement dashboard — hosts, vulns, creds, phase |
| `/targets` | Host & service table |
| `/vulns` | Findings by severity |
| `/creds` | Discovered credentials |
| `/scope` | View/edit target scope |
| `/phase` | Phase management |
| `/mode` | Switch auto / free / guided |
| `/pause` | Pause on findings (never / always / checkpoint) |
| `/report` | Generate a pentest report |

And set how much rope the agent gets:

Expand All @@ -192,19 +206,19 @@ One toolkit across offensive security:

## Configuration

Config lives at `.pentestcode/pentestcode.jsonc`:
Project config can live at `.pentestcode/pentestcode.jsonc`; global config is in `~/.config/pentestcode/`. Existing PentestCode and OpenCode config names are also discovered. A minimal native V2 configuration is:

```jsonc
{
"provider": {
"anthropic": {
"model": "claude-sonnet-4-20250514"
}
}
"$schema": "https://opencode.ai/config.json",
"model": "anthropic/claude-sonnet-4-20250514",
"default_agent": "pentest",
}
```

Providers: Anthropic, OpenAI, Google, Azure, AWS Bedrock, Ollama, Together, Groq, Fireworks, DeepSeek, Mistral, and more via [ai-sdk](https://github.com/vercel/ai).
Terminal preferences use the separate global `cli.json`. Engagements and shared knowledge remain under `~/.pentestcode/`; `PENTESTCODE_HOME` overrides that domain-data root.

Providers include Anthropic, OpenAI, Google, Azure, AWS Bedrock, Ollama, Together, Groq, Fireworks, DeepSeek and Mistral.

## Contributing

Expand Down
Loading