Octoparse

Desktop and cloud web scraping with templates and scheduled runs

KINDA · partial replacement
price $83/mosubscription / year $996estimated build time weekend to multi-dayreplaced by 0 people

The visible visual web scraping loop is buildable, but a credible replacement needs more than the first screen. Octoparse earns its keep through connectors, auth, reliability, so expect a weekend or multi-day build and a narrower personal scope.

Build verification: not recorded. How we judge buildability

What you give up

  • hundreds of maintained connectors
  • OAuth app verification
  • durable execution at scale
  • schema drift handling and enterprise controls

Why people still pay

Octoparse: One automation is easy. Subscribers pay for maintained integrations, credentials, retries, observability, and someone else owning breakage.

Your build guide

The stack, security requirements, and agent rules for a focused replacement.

Before you start

  • A supported Node release, PostgreSQL, HTTPS for a shared deployment, and a backup destination. Begin with one workspace and explicit owner/member permissions.
  • Implementation components: Next.js App Router and TypeScript for server-rendered pages and validated mutations. PostgreSQL with Drizzle migrations; Better Auth sessions for a small private workspace. A TypeScript worker with a typed tool registry and a durable per-step execution log.
  • Scope boundary: Proxy evasion and a maintained universal scraper catalog are excluded.
01
Next.js App Router and TypeScript for server-rendered pages and validated mutations.
02
PostgreSQL with Drizzle migrations; Better Auth sessions for a small private workspace.
03
A TypeScript worker with a typed tool registry and a durable per-step execution log.
04
Domain model: approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes
engineering roadmap

Implementation plan

1

Phase 1

Scope and fixtures. Implement this bounded workflow: Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Proxy evasion and a maintained universal scraper catalog are excluded.

2

Phase 2

Durable model. Model approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.

3

Phase 3

Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.

4

Phase 4

Permissions and integration failure. Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.

5

Phase 5

Portable handoff. Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.

6

Phase 6

Acceptance scenarios. Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

the pro prompt
download AGENTS.md
WORKING SLICE
Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log.

Build this scoped Octoparse-inspired workflow with a documented data model and visible failure states.

Architecture
- Next.js App Router and TypeScript for server-rendered pages and validated mutations.
- PostgreSQL with Drizzle migrations; Better Auth sessions for a small private workspace.
- A TypeScript worker with a typed tool registry and a durable per-step execution log.

Prerequisites and limits
A supported Node release, PostgreSQL, HTTPS for a shared deployment, and a backup destination. Begin with one workspace and explicit owner/member permissions.
Outside this release: Proxy evasion and a maintained universal scraper catalog are excluded.

Data model and correctness
approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes
Invariant: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.

Security and privacy
Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission.

Recovery and export
Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Proxy evasion and a maintained universal scraper catalog are excluded.
2. Phase 2 — Durable model. Model approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Checkpoint each step with input hashes and external receipts. Separate safe retries from unknown side effects; resume from the last confirmed step with a run budget and cancellation.
4. Phase 4 — Permissions and integration failure. Authorize every record read and mutation on the server using its workspace membership; validate payloads, protect mutations against CSRF, and escape user-authored HTML. Allowlist tools, destinations and credential scopes. External writes, shell commands and messages require explicit policy approval; imported content cannot grant permission. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Export versioned JSON plus attachments and a readable CSV summary. Restore into a separate database and compare record IDs and attachment checksums before switching. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [supabase-postgres-best-practices](https://github.com/supabase/agent-skills/blob/main/skills/supabase-postgres-best-practices/SKILL.md) — Review PostgreSQL schemas, queries, indexes, pooling, concurrency and row-level security. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [agent-browser](https://github.com/vercel-labs/agent-browser/blob/main/skills/agent-browser/SKILL.md) — Automate browser interaction using accessibility snapshots, element references and reproducible navigation workflows. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: approved crawl targets, selector recipes, pagination rules, extraction runs, source snapshots and row hashes
Project rule — preserve this invariant: Do not bypass authentication barriers, CAPTCHAs or access controls; selector drift must fail visibly rather than return mislabeled data.
Project rule — acceptance evidence: Change a required selector in a fixture page and stop with a diagnostic; restarting page three cannot duplicate rows already captured from pages one and two.

$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md · generated from this app's build plan

share on X ↗

Alternatives to building your own

EasySpiderA finished visual desktop scraper with selectors, pagination, loops, and scheduled local runs.44kjul 2026open source↗MaxunA browser-based robot recorder with schedules and exports, but your Docker host becomes the cloud.17kjul 2026open source↗

no votes, no pay-to-list · just what's real

Octoparse pricing

planmonthlyannual (per mo)what you get
free$0$010 tasks; 1 device; 2 concurrent local runs; local-only execution; 10,000 rows/export and 50,000 rows/month.
standard$83$69100 tasks; up to 3 concurrent cloud runs; unlimited exports; scheduling and data-export API.Annual rate saves about 16%; 5-day money-back guarantee.
professional$299$249250 tasks; up to 20 concurrent cloud runs; advanced API; cloud monitoring and backup.Annual rate saves about 16%; 5-day money-back guarantee.
enterprise——750+ tasks and 40+ concurrent cloud runs; team collaboration and dedicated success management.Quote required.

free tier10 tasks, 1 device, 2 concurrent local runs, 10,000 rows per export and 50,000 exported rows per month; no cloud runs, scheduling or API

billingmonthly + quarterly + annual; annual saves about 16%

hidden costsResidential proxies cost $3/GB; pay-per-result templates cost $0.001-$3 per 1,000 results; CAPTCHA solving costs $1-$1.50 per 1,000; crawler setup starts at $399 and managed data service at $599.

pricing sources checked 2026-08-11 · pricing source ↗

Questions about Octoparse

Can you build your own Octoparse with AI?

Partly. The visible visual web scraping loop is buildable, but a credible replacement needs more than the first screen. Octoparse earns its keep through connectors, auth, reliability, so expect a weekend or multi-day build and a narrower personal scope.

What does the Octoparse build prompt cover?

The prompt starts with this scope: Configure one permitted site extraction recipe, preview selected fields and run a bounded browser job through its pagination. Record source URLs and errors for each row, then export CSV with a resumable job log. Full-product capabilities excluded from the comparison include: hundreds of maintained connectors; OAuth app verification; durable execution at scale. Follow the implementation plan and its prerequisites before expanding the build.

How do I use the prompt, AGENTS.md and agent skills?

Start with the Octoparse prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.

How long will this Octoparse project take?

The catalogue estimate is weekend to multi-day for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.

What would I give up by replacing Octoparse?

hundreds of maintained connectors; OAuth app verification; durable execution at scale; schema drift handling and enterprise controls. Octoparse: One automation is easy. Subscribers pay for maintained integrations, credentials, retries, observability, and someone else owning breakage.

What price is this guide comparing against?

The recorded Standard plan is $83/mo (monthly), checked 2026-08-11. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.

What can I use instead of building Octoparse?

EasySpider: A finished visual desktop scraper with selectors, pagination, loops, and scheduled local runs. Maxun: A browser-based robot recorder with schedules and exports, but your Docker host becomes the cloud. Check each option's license, hosting needs and feature limits.

Every week, more subscriptions die.

New verdicts, new prompts, the week's most-doomed apps.
One email. Unsubscribe in one click.

last week:100 Questions · KINDA1of10 · KINDA1Password · KINDA+1090 more

free forever · no scanner spam · the prompt stays on the site, the deaths come to you

$weekly: what got a verdict, what died.