# AGENTS.md — Build guide for ZeroLeaks

## Project scope
Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs.

Catalogue verdict: kinda. The mechanics here are not exotic: fire a few hundred adversarial prompts at your own chat endpoint, capture the responses, and check whether any of them contain your system prompt, tool schemas or API keys. An agent can build that loop, including an LLM-as-judge scorer and an HTML report, in a single sitting, and it will find the embarrassing stuff on day one. What you cannot one-shot is a probe library that stays current with each new model release and each new jailbreak family, because that is maintained knowledge, not code. There is also a trust angle: 'we ran our own script and found nothing' reads very differently in a security review than a dated third-party report. Build it for your own sanity checks, keep paying if you need something to show someone else.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- Runtime and tools: Python, Typer, httpx, Pydantic, Jinja2 and versioned JSONL run files; no database or web server.
- Before starting: One owned endpoint specification, a bounded auditable probe pack, a local reference fixture and secrets outside shareable reports.

## Stack and architecture
- Python, Typer, httpx, Pydantic, Jinja2 and versioned JSONL run files; no database or web server
- Data design: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion.
- Setup: One owned endpoint specification, a bounded auditable probe pack, a local reference fixture and secrets outside shareable reports

## Security and data integrity
- Operate only against configured authorized endpoints. Redact secret values in saved evidence and distinguish detector findings, judge opinion, failures and unknown outcomes.
- Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.
- Keep secrets outside client bundles and exported projects; document what leaves the device and make retention/deletion controls visible.

## Agent implementation rules
- Project rule — domain: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion.
- Project rule — scope and recovery: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.
- Project rule — acceptance: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs.
- Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.

## Optional agent skills and references
- Recommended skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
- Recommended skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — review configuration and API defaults against the app-specific invariants and recovery boundaries above; this is not a security certification. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate this working slice using synthetic or explicitly authorized non-sensitive examples: Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs.
- Share a synthetic example export and the acceptance walkthrough; keep real customer, health, financial and source data private: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs.
- State the limits before asking someone to replace their existing tool: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.

## Engineering roadmap
1. Phase 1 — Pin the working slice and create its example input: Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs. Confirm setup: One owned endpoint specification, a bounded auditable probe pack, a local reference fixture and secrets outside shareable reports.
2. Phase 2 — Implement durable result files and command invariants before formatting terminal output: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion.
3. Phase 3 — Connect CLI commands to real saved state and explicit result/exit statuses. Operate only against configured authorized endpoints. Redact secret values in saved evidence and distinguish detector findings, judge opinion, failures and unknown outcomes.
4. Phase 4 — Expose the app-specific limits and recovery path in context: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.
5. Phase 5 — Walk through this concrete acceptance case and preserve its exported evidence: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.

## Paid-product capabilities outside this build
- A curated, maintained probe corpus that tracks new jailbreak families instead of whatever the agent remembered on build day
- Multi-turn and multi-model attack strategies, including crescendo and encoding tricks, done properly rather than as single-shot prompts
- A third-party report with a date on it that you can hand to a customer or an auditor
- Severity triage and remediation guidance written by someone who has seen a lot of these
- Regression runs on every model or prompt change without you remembering to trigger them

## Implementation prompt
Build me a local CLI for authorized system-prompt and secret-leak regression checks, as a narrow substitute for Zeroleaks. Requirements:

- Use Python 3.12, Typer, httpx, Pydantic, Jinja2, and JSONL files in ./runs/. No web server or database.
- A target.yaml names one endpoint I control, its request template, response-text path, allowed request rate, and timeout. Confirm the endpoint scope before running probes. Keep bearer keys in .env.
- Store versioned probes with ID, family, severity, and turns. Start with a small auditable pack covering direct extraction, role confusion, encoded requests, and pasted-document injection; permit user-authored additions.
- Run probes with bounded concurrency and backoff on 429/5xx. Save request IDs, redacted transcripts, status, and elapsed time. Never write real secret values or full system instructions into a shareable report.
- Detect exact secret-pattern hits and overlap with a local reference prompt. An optional model judge may add a labelled opinion; a judge failure produces unknown, not clean.
- Produce report.html and report.json with finding evidence, detector, severity, probe version, and unknown counts. A diff command compares two runs by stable probe ID.
- Acceptance: a fixture containing a planted token flags the exact probe; a harmless fixture stays unflagged; a timed-out endpoint stays unknown; diff reports one introduced and one resolved finding.
- No accounts, telemetry, third-party scanning, or unbounded attack traffic. README documents permission, data retention, redaction limits, and the fact that a clean run is not proof of safety.

EDITORIAL IMPLEMENTATION CONTRACT
Working slice: Run a versioned, bounded probe pack against an authorized chat endpoint, inspect redacted leakage evidence and compare findings across two runs.
Data and invariants: Store TargetConfig, ProbeVersion, Attempt, DetectorResult and RunComparison in JSONL; failure/timeout is unknown, and detector evidence is separate from an optional model judge's opinion.
Boundary and recovery: Keep the existing Python CLI with no database/server. Do not store real secrets in shareable transcripts, scan third-party targets or treat a clean probe run as proof of security.
Acceptance walkthrough: Use one fixture with a planted token, one harmless response and one timeout; report finding, clear observation and unknown separately, with a stable-ID diff between runs.
Record actual dependency versions, permissions and provider access in setup instructions. Preserve originals, expose partial failures and document backup/restore. These are acceptance requirements, not a claim of a completed or production-certified build. Add the domain, recovery and acceptance rules to AGENTS.md so future edits preserve them.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
