# AGENTS.md — Build guide for Octoparse

## Project scope
Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log.

Catalogue verdict: kinda. The visible visual web scraping loop is buildable, but a credible replacement needs more than the first screen. Octoparse earns its keep through connectors, auth, reliability, so expect a weekend or multi-day build and a narrower personal scope.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- A supported Node release, PostgreSQL, HTTPS for a shared deployment, and a backup destination. Begin with one workspace and explicit owner/member permissions.
- Implementation components: Next.js App Router and TypeScript for server-rendered pages and validated mutations. PostgreSQL with Drizzle migrations; Better Auth sessions for a small private workspace. A TypeScript worker with a typed tool registry and a durable per-step execution log.
- Scope boundary: Proxy evasion and a maintained universal scraper catalog are excluded.

## Stack and architecture
- Next.js App Router and TypeScript for server-rendered pages and validated mutations.
- PostgreSQL with Drizzle migrations; Better Auth sessions for a small private workspace.
- A TypeScript worker with a typed tool registry and a durable per-step execution log.
- Domain model: approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes

## Security and data integrity
- Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission.
- Correctness boundary: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
- Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
- Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching.

## Agent implementation rules
- Project rule — data model: approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes
- Project rule — preserve this invariant: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
- Project rule — acceptance evidence: Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two.

## Optional agent skills and references
- Optional external skill: [supabase-postgres-best-practices](https://github.com/supabase/agent-skills/blob/main/skills/supabase-postgres-best-practices/SKILL.md) — Review PostgreSQL schemas, queries, indexes, pooling, concurrency and row-level security. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [agent-browser](https://github.com/vercel-labs/agent-browser/blob/main/skills/agent-browser/SKILL.md) — Automate browser interaction using accessibility snapshots, element references and reproducible navigation workflows. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate the actual Octoparse-inspired workflow with owned or clearly labeled sample data: Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log.
- Publish a reproducible walkthrough with this observable result: Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two.
- Explain who can operate this scoped tool, its setup and ongoing costs, and these remaining product gaps: Proxy evasion and a maintained universal scraper catalog are excluded. Avoid guaranteed savings, performance scores or implied endorsement.

## Engineering roadmap
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Proxy evasion and a maintained universal scraper catalog are excluded.
2. Phase 2 — Durable model. Model approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
4. Phase 4 — Permissions and integration failure. Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

## Paid-product capabilities outside this build
- hundreds of maintained connectors
- OAuth app verification
- durable execution at scale
- schema drift handling and enterprise controls

## Implementation prompt
WORKING SLICE
Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log.

Build this scoped Octoparse-inspired workflow with a documented data model and visible failure states.

Architecture
- Next.js App Router and TypeScript for server-rendered pages and validated mutations.
- PostgreSQL with Drizzle migrations; Better Auth sessions for a small private workspace.
- A TypeScript worker with a typed tool registry and a durable per-step execution log.

Prerequisites and limits
A supported Node release, PostgreSQL, HTTPS for a shared deployment, and a backup destination. Begin with one workspace and explicit owner/member permissions.
Outside this release: Proxy evasion and a maintained universal scraper catalog are excluded.

Data model and correctness
approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes
Invariant: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.

Security and privacy
Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission.

Recovery and export
Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Proxy evasion and a maintained universal scraper catalog are excluded.
2. Phase 2 — Durable model. Model approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
4. Phase 4 — Permissions and integration failure. Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [supabase-postgres-best-practices](https://github.com/supabase/agent-skills/blob/main/skills/supabase-postgres-best-practices/SKILL.md) — Review PostgreSQL schemas, queries, indexes, pooling, concurrency and row-level security. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [agent-browser](https://github.com/vercel-labs/agent-browser/blob/main/skills/agent-browser/SKILL.md) — Automate browser interaction using accessibility snapshots, element references and reproducible navigation workflows. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes
Project rule — preserve this invariant: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
Project rule — acceptance evidence: Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
