# AGENTS.md — Build guide for AdaL

## Project scope
Build a terminal planner that assigns bounded repository tasks to workers with separate contexts. Offer read/search first, inspectable patches second, and explicitly approved shell commands; resume a saved run with a configurable model adapter and budget.

Catalogue verdict: kinda. The core loop is genuinely reachable: open-source agents already route across model families, run tools, and edit repos, and a weekend of wiring gets you a personal version. What does not fall out of a weekend is the harness around the model, which is where AdaL claims its numbers come from, plus browser-based verification, clustered code review, and the team controls. You end up with an agent that works and a noticeably worse recovery rate on long tasks.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- Node, Git, a selected repository, a provider API key and a per-run budget. Start with a fixture repository and a read-only tool set; browser workers are outside this release.
- Implementation components: Node.js and TypeScript with a terminal interface and provider SDK adapters. SQLite run metadata, append-only JSONL tool traces and isolated Git worktrees for explicitly assigned workers.
- Scope boundary: Enterprise identity, provider billing consolidation and autonomous browser verification are separate projects.

## Stack and architecture
- Node.js and TypeScript with a terminal interface and provider SDK adapters.
- SQLite run metadata, append-only JSONL tool traces and isolated Git worktrees for explicitly assigned workers.
- Domain model: repository roots, branch/HEAD snapshots, agent runs, assigned worker scopes, tool calls and usage receipts

## Security and data integrity
- Resolve and enforce repository/worktree root containment, including symlinks, before every file operation. Respect filesystem permissions, require explicit tool/shell approval and bound command duration. Keep provider secrets and confidential file contents out of terminal diagnostics and run traces; model output cannot expand tool access.
- Correctness boundary: Workers cannot write outside their assigned root or overwrite another worker's change; reported usage distinguishes provider counts from estimates.
- Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
- Save run state and tool receipts without secrets. Resume against the recorded branch/HEAD or stop for reconciliation; retain patches and worktrees until the owner accepts or discards them.

## Agent implementation rules
- Project rule — data model: repository roots, branch/HEAD snapshots, agent runs, assigned worker scopes, tool calls and usage receipts
- Project rule — preserve this invariant: Workers cannot write outside their assigned root or overwrite another worker's change; reported usage distinguishes provider counts from estimates.
- Project rule — acceptance evidence: Two workers touching the same file produce a conflict for review; an interrupted tool call resumes without rerunning an unknown external action.

## Optional agent skills and references
- Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate the actual AdaL-inspired workflow with owned or clearly labeled sample data: Build a terminal planner that assigns bounded repository tasks to workers with separate contexts. Offer read/search first, inspectable patches second, and explicitly approved shell commands; resume a saved run with a configurable model adapter and budget.
- Publish a reproducible walkthrough with this observable result: Two workers touching the same file produce a conflict for review; an interrupted tool call resumes without rerunning an unknown external action.
- Explain who can operate this scoped tool, its setup and ongoing costs, and these remaining product gaps: Enterprise identity, provider billing consolidation and autonomous browser verification are separate projects. Avoid guaranteed savings, performance scores or implied endorsement.

## Engineering roadmap
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Build a terminal planner that assigns bounded repository tasks to workers with separate contexts. Offer read/search first, inspectable patches second, and explicitly approved shell commands; resume a saved run with a configurable model adapter and budget. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Enterprise identity, provider billing consolidation and autonomous browser verification are separate projects.
2. Phase 2 — Durable model. Model repository roots, branch/HEAD snapshots, agent runs, assigned worker scopes, tool calls and usage receipts Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Workers cannot write outside their assigned root or overwrite another worker's change; reported usage distinguishes provider counts from estimates.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
4. Phase 4 — Permissions and integration failure. Resolve and enforce repository/worktree root containment, including symlinks, before every file operation. Respect filesystem permissions, require explicit tool/shell approval and bound command duration. Keep provider secrets and confidential file contents out of terminal diagnostics and run traces; model output cannot expand tool access. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Save run state and tool receipts without secrets. Resume against the recorded branch/HEAD or stop for reconciliation; retain patches and worktrees until the owner accepts or discards them. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Two workers touching the same file produce a conflict for review; an interrupted tool call resumes without rerunning an unknown external action. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

## Paid-product capabilities outside this build
- harness tuning: the planning, recovery and verification work that separates a demo agent from one that finishes long tasks
- browser-use verification of the app you just changed
- review clustering that groups a large diff into risk-ranked units
- one bill across every frontier provider instead of six metered API accounts
- SSO, SAML/SCIM, zero data retention and org-level model deny lists

## Implementation prompt
WORKING SLICE
Build a terminal planner that assigns bounded repository tasks to workers with separate contexts. Offer read/search first, inspectable patches second, and explicitly approved shell commands; resume a saved run with a configurable model adapter and budget.

Build this scoped AdaL-inspired workflow with a documented data model and visible failure states.

Architecture
- Node.js and TypeScript with a terminal interface and provider SDK adapters.
- SQLite run metadata, append-only JSONL tool traces and isolated Git worktrees for explicitly assigned workers.

Prerequisites and limits
Node, Git, a selected repository, a provider API key and a per-run budget. Start with a fixture repository and a read-only tool set; browser workers are outside this release.
Outside this release: Enterprise identity, provider billing consolidation and autonomous browser verification are separate projects.

Data model and correctness
repository roots, branch/HEAD snapshots, agent runs, assigned worker scopes, tool calls and usage receipts
Invariant: Workers cannot write outside their assigned root or overwrite another worker's change; reported usage distinguishes provider counts from estimates.
Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.

Security and privacy
Resolve and enforce repository/worktree root containment, including symlinks, before every file operation. Respect filesystem permissions, require explicit tool/shell approval and bound command duration. Keep provider secrets and confidential file contents out of terminal diagnostics and run traces; model output cannot expand tool access.

Recovery and export
Save run state and tool receipts without secrets. Resume against the recorded branch/HEAD or stop for reconciliation; retain patches and worktrees until the owner accepts or discards them.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Build a terminal planner that assigns bounded repository tasks to workers with separate contexts. Offer read/search first, inspectable patches second, and explicitly approved shell commands; resume a saved run with a configurable model adapter and budget. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Enterprise identity, provider billing consolidation and autonomous browser verification are separate projects.
2. Phase 2 — Durable model. Model repository roots, branch/HEAD snapshots, agent runs, assigned worker scopes, tool calls and usage receipts Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Workers cannot write outside their assigned root or overwrite another worker's change; reported usage distinguishes provider counts from estimates.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
4. Phase 4 — Permissions and integration failure. Resolve and enforce repository/worktree root containment, including symlinks, before every file operation. Respect filesystem permissions, require explicit tool/shell approval and bound command duration. Keep provider secrets and confidential file contents out of terminal diagnostics and run traces; model output cannot expand tool access. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Save run state and tool receipts without secrets. Resume against the recorded branch/HEAD or stop for reconciliation; retain patches and worktrees until the owner accepts or discards them. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Two workers touching the same file produce a conflict for review; an interrupted tool call resumes without rerunning an unknown external action. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
Two workers touching the same file produce a conflict for review; an interrupted tool call resumes without rerunning an unknown external action.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: repository roots, branch/HEAD snapshots, agent runs, assigned worker scopes, tool calls and usage receipts
Project rule — preserve this invariant: Workers cannot write outside their assigned root or overwrite another worker's change; reported usage distinguishes provider counts from estimates.
Project rule — acceptance evidence: Two workers touching the same file produce a conflict for review; an interrupted tool call resumes without rerunning an unknown external action.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
