Research artifact for the JSS 2026 submission "Prompt-to-API-Call Injection Attacks". It contains the attack-vector generation framework, the end-to-end attack/defense harness, the five defense strategies (D1–D5), and all result data used in the manuscript.
Second-revision synchronization (2026-09-23). The 62 final result files and 21 retained source files were checked against the experiment host. Every file matched the previous public release byte for byte before the documentation, analysis, and launcher corrections described below. All original result JSONs remain unchanged. The manifest pins the final data, supporting records, and source versions with SHA-256 hashes.
Verify the published results without running experiments or calling any API:
python3 verify_artifact.py
python3 build_commercial_table.py
python3 analyze_guard_capacity.pyThe first command also recomputes and checks
analysis/revision2_summary.json.
Historical configuration and measurement details are in
REPRODUCIBILITY.md.
Scope and intended use. This is defensive-security research. The attacks were executed against local testbed instances. Backend deployments are separate from this code archive (see §3). Do not run these payloads against systems you do not own or operate. The D5 Semantic Intent Validator is the mitigation proposed and evaluated here.
An LLM agent that holds a user's session credentials translates natural-language requests into REST API calls. A Prompt-to-API-Call (P2A) injection makes the agent emit an unauthorized API call under the user's legitimate identity — a "confused deputy" at the LLM-agent layer. The study builds 52 attack vectors in 5 categories (Unrestricted, Restricted-Direct, Restricted-Indirect, Injection-Location, Obfuscation-Bypass), adapts them to 5 heterogeneous backends (Strapi portal, E-Commerce, Gitea, Home Assistant, Directus), yielding 260 end-to-end scenarios, and evaluates them on 4 open-source LLMs (LLaMA-3 8B, Qwen2.5 7B, Ministral 3 8B, Qwen2.5-Coder 14B) plus two commercial models.
p2a_demo.py Flask harness: attack generation → call → defense → execute → reflect
(POST /api/run is the single entry point the runners call)
attacks.py 52 base attack vectors + per-category operators
{ecommerce,directus,gitea,ha}_attacks.py per-testbed attack adaptations
defenses.py D1–D5 implementations (D5 = Phase-1 whitelist + Phase-2 guard LLM)
api_server.py, ecommerce_api_server.py mock/local backend servers
# Runners
run_rq1_real.py RQ1/RQ2 baseline: 52 × 5 systems × 4 models = 1,040 runs
fresh output: results/RQ1/; final archive: results/RQ2/
run_defense_experiments.py RQ3: 52 × 5 systems × 6 defenses (LLaMA-3) = 1,560 runs → results/defense/
r12_strategy_ablation.py R1.2 operator selection vs. combination ablation
r21_adaptive_guard.py R2.1 adaptive-attack robustness of the D5 guard ("who guards the guardian")
r22_false_positive.py R2.2 false-positive rate / security–usability trade-off of D5
guard_cloud_eval.py strong-end guard data point (cloud reasoning model)
commercial_crossmodel.py commercial closed-source models (GPT-4o-mini, Claude-Haiku)
# Analysis
build_commercial_table.py, analyze_guard_capacity.py assemble manuscript tables
artifact_analysis.py offline raw-record aggregation and strict call-schema metric
verify_artifact.py SHA-256 checks, completeness checks, and summary verification
results/ all output data (see §5)
provenance/ final-record manifest and retained historical metadata
analysis/ derived second-revision summary (not new experiment results)
.gitignore excludes .env and local testbed deployments; the testbeds are
stood up separately (see §3).
python3 -m venv venv && source venv/bin/activate
pip install flask==3.1.3 flask-cors requests # historical Python: 3.11.15
# Local LLMs via Ollama (https://ollama.com)
ollama pull llama3:latest # 8B, primary defense model
ollama pull qwen2.5:latest # 7B
ollama pull ministral-3:latest # nominal 8B; retained GGUF metadata: 8.9B
ollama pull qwen2.5-coder:14b # 14BThese commands prepare a new environment; the complete historical dependency lock was not retained. Mutable model tags may now resolve differently. Compare local artifact digests with the retained manifests before treating a new run as using the archived model versions.
Deploy the five backends locally (Strapi, a Flask E-Commerce server, Gitea, Home Assistant, Directus) and point the harness at them. All credentials and endpoints are read from environment variables — never hard-code them.
| Variable | Purpose |
|---|---|
OLLAMA_URL, OLLAMA_MODEL, LLM_TIMEOUT |
local LLM endpoint / default model |
USE_REAL_BACKEND=1 |
execute against real backends (vs. mock) |
STRAPI_URL, ECOMMERCE_URL, GITEA_URL, HA_URL, DIRECTUS_URL |
testbed base URLs |
GITEA_TOKEN, HA_REFRESH_TOKEN |
testbed admin credentials (set your own) |
COMMERCIAL_API_KEY, COMMERCIAL_BASE |
commercial-model endpoint (OpenAI-compatible) |
DEEPSEEK_API_KEY, GUARD_CLOUD_MODEL, CLOUD_BASE |
cloud reasoning guard |
GUARD_MODEL |
Phase-2 guard model for D5 (default llama3:latest) |
DEMO_URL, PORT |
harness address the runners call |
The scripts fall back to placeholder demo tokens for the throwaway local testbeds; override every credential via the environment for any real run.
For the existing paper numbers, use the offline verification commands above.
The following commands perform new inference and backend operations. Run them
in a separate working copy: defense and revision runners write to their result
paths, so they would otherwise overwrite the archived files. The final baseline
records were archived under results/RQ2/; the retained baseline runner writes
new output under results/RQ1/.
After deploying and configuring the five testbeds, start the harness and run the phases with one trial per cell, as recorded in the final JSONs:
# 0. harness
USE_REAL_BACKEND=1 python p2a_demo.py # serves POST /api/run
# 1. RQ1 (JSON compliance) + RQ2 (cross-model ASR) → fresh results/RQ1/
python run_rq1_real.py --system all --model all --n 1
# 2. RQ3 (five defenses, LLaMA-3) → results/defense/
python run_defense_experiments.py --defense all --system all --model llama3:latest
# 3. Revision experiments
USE_REAL_BACKEND=0 python r12_strategy_ablation.py # → results/R12_strategy_ablation.json
GUARD_MODEL=llama3:latest python r21_adaptive_guard.py # → results/R21_adaptive_guard_*.json
GUARD_MODEL=llama3:latest python r22_false_positive.py # → results/R22_false_positive_*.json
DEEPSEEK_API_KEY=... GUARD_CLOUD_MODEL=deepseek-v4-flash python guard_cloud_eval.py
USE_REAL_BACKEND=0 COMMERCIAL_API_KEY=... python commercial_crossmodel.py # → results/commercial_*.json
# Open-source anchor under the same generation-level protocol:
USE_REAL_BACKEND=0 COMMERCIAL_API_KEY=... XMODELS=llama3:latest python commercial_crossmodel.pyFor the other two local guards, repeat the guard commands with
GUARD_MODEL=qwen2.5:latest and GUARD_MODEL=qwen2.5-coder:14b.
commercial_crossmodel.py uses the mock backend to measure generation-level
success; it is a separate protocol from the real-backend RQ2 evaluation. Its
entry point checks COMMERCIAL_API_KEY even when only the local anchor is
selected; the local model request itself is sent to Ollama.
The scripts set temperature=0.0; each archived ⟨model, system, attack⟩ cell
has one trial. Wilson 95% CIs describe variation across attack vectors, not
repeated samples of the same cell. Provider or runtime changes can affect a new
run even with the same request parameters.
| Data | Manuscript |
|---|---|
results/RQ2/*/rq1_*_*.json |
RQ1 JSON compliance (Table 3); RQ2 cross-model ASR (Tables 7, 8) |
results/defense/{none,D1..D5}/*_llama3_latest.json |
RQ3 residual ASR / block rate (Tables 10, 11, 12) |
results/R12_strategy_ablation.json |
Strategy selection vs. combination (Table 5) |
results/commercial_*.json |
Commercial closed-source models (Table 9) |
Four valid guards' R21_adaptive_guard_*.json, R22_false_positive_*.json |
D5 guard robustness / usability (Table 13) |
Table numbers refer to the second-revision manuscript. The manifest lists the
exact 62 contributing files; broad wildcards also match two failed 32B guard
attempts, which are retained but excluded from the four-guard analysis.
results/defense/grand_summary.json duplicates the 30 individual defense files;
the verifier checks their equality and does not count them twice.
The raw schemas differ by protocol: baseline files use raw_trials, defense
files use raw_results, commercial files use rows, and guard files use
cases or results. Commercial and guard studies do not have end-to-end backend
execution traces. RQ3 latency is only persisted for final BLOCK decisions.
Keep the two no-defense runs separate. RQ2 LLaMA records 180/260 successes (69.2%). The later RQ3 control records 175/260 (67.3%), and is the reference for the defense comparison. D5 records 57/260 (21.9%), a 67.4% relative reduction against the RQ3 control. The final manuscript and offline summary use these respective runs; they do not substitute the RQ2 control into RQ3.
Use the defined JSON criterion. The commercial records preserve the original
json_compliance_pct helper field. The manuscript instead counts saved calls
with a nonempty string endpoint and a recognized HTTP method. The corrected
table builder recomputes 92.7%, 100.0%, and 98.1% for LLaMA, GPT-4o-mini, and
Claude Haiku 4.5 respectively. It does not rewrite the historical helper fields.
- Phase 1 — deterministic whitelist (
d3_gate_output): an O(1) check of the generated⟨method, endpoint⟩against the role's minimal-privilege whitelist. Non-whitelisted calls are blocked with no model call. - Phase 2 — semantic guard LLM (
d5_validate_intent): every whitelist-admitted call is checked by a guard model for semantic consistency between the user's query and the generated action; inconsistent calls are blocked. This catches the "intent-drift" attacks (legitimate endpoint, unauthorized intent) that a whitelist cannot.
The manifest records both the retained experiment-host hashes and the current
hashes. Four retained files have explicitly documented packaging corrections:
the two launchers now use one trial and a portable working directory; the
normal-request runner's usage example uses GUARD_MODEL; and the commercial
runner's docstring fixes the model-size range and explains its legacy metric.
The Python experiment logic is unchanged. Analysis helpers were updated to
recompute the paper's metric and use the exact cloud-guard label.
The old *.bak* sources, run_live.py, and patch_dashboard.py remain as
historical development material. They are not the final reproduction workflow.
The earlier ten-trial shell commands did not describe the released final JSONs
and have been corrected; no new experiment was run during this synchronization.
If you use this artifact, please cite the JSS 2026 paper (see the manuscript for the current reference). Report issues via the repository's issue tracker.