Oncrawl
Crawl one owned site locally and join pages to user-supplied analytics exports
A consolation build is possible, but the paid product's decisive value sits outside a solo rebuild. For Oncrawl, crawl one owned site locally and join pages to user-supplied analytics exports. The hard boundary is large-scale crawler, log analysis, data connectors, and enterprise support, plus crawl scale, rule depth, and operational polish.
Build verification: not recorded. How we judge buildability
What you give up
- large-scale crawler, log analysis, data connectors, and enterprise support
- massive hosted crawl capacity
- proprietary scoring
- continuous monitoring
- agency reporting and support
Why people still pay
People still pay for Oncrawl because a crawler is buildable; professionals pay for years of edge-case handling and reports they can trust with clients. The recurring cost buys robots handling, rendering, canonicalization, deduplication, crawl traps, rule maintenance, scheduling, storage, and false positives, not just the visible interface.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Python 3.12
- Playwright browsers
- permission to crawl the target site
- local disk space
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule, data: Model sites, crawl pages, keyword snapshots, audit findings, and report dates; keep stable source IDs and timestamps.
Project rule, behavior: Collect status, title, description, headings, canonical, robots, links, images, structured data, and rendered text.
Project rule, recovery: Require an explicit ownership or permission acknowledgement before a crawl starts.
Implementation plan
Phase 1, architecture and data
Use Python 3.12, FastAPI, SQLite, Playwright, and an HTMX interface. Model sites, crawl pages, keyword snapshots, audit findings, and report dates; keep stable source IDs and timestamps.
Phase 2, implement
Collect status, title, description, headings, canonical, robots, links, images, structured data, and rendered text.
Phase 3, implement
Detect duplicates, orphan candidates, broken links, redirect chains, missing metadata, and indexability conflicts.
Phase 4, review and output
Show every issue with affected URLs, evidence, severity, and a concrete remediation note. Export crawl data and issues to CSV plus a self-contained HTML report.
Phase 5, recovery and acceptance
Require an explicit ownership or permission acknowledgement before a crawl starts. Verify this invariant with a saved fixture: A failed crawl retains its last dated snapshot; a report must not imply backlink or keyword coverage the input lacks. State the practical limit: large-scale crawler, log analysis, data connectors, and enterprise support.
Build me a focused technical audits and on-page optimization workflow for the personal core of Oncrawl. Requirements: - Use Python 3.12, FastAPI, SQLite, Playwright, and an HTMX interface. Model sites, crawl pages, keyword snapshots, audit findings, and report dates; keep stable source IDs and timestamps. - Paid product context: Crawl one owned site locally and join pages to user-supplied analytics exports. Build only this DIY scope: Crawl one owned site locally, inspect HTML and rendered pages, join pages to user-supplied analytics exports, explain prioritized issues, and export a reproducible audit. - Collect status, title, description, headings, canonical, robots, links, images, structured data, and rendered text. - Detect duplicates, orphan candidates, broken links, redirect chains, missing metadata, and indexability conflicts. - Show every issue with affected URLs, evidence, severity, and a concrete remediation note. Export crawl data and issues to CSV plus a self-contained HTML report. - Use a local web page with input, progress, review, and export views. Required input or access: Playwright browsers; permission to crawl the target site. - Recovery: Require an explicit ownership or permission acknowledgement before a crawl starts. - Acceptance: with one labelled sample, show the input, saved intermediate state, and exported result; verify this invariant: A failed crawl retains its last dated snapshot; a report must not imply backlink or keyword coverage the input lacks. - Out of scope: large-scale crawler, log analysis, data connectors, and enterprise support; massive hosted crawl capacity. Keep this a personal, inspectable workflow. - Include a README with setup, a sample input, required keys or permissions, data location, and the supported scope.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy AGENTS.md · generated from this app's build plan
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Alternatives to building your own
no votes, no pay-to-list · just what's real
Oncrawl pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| custom | — | — | Custom crawl volume, projects, log/analytics integrations, data retention, users, and support; no current public numeric allowance.Contact sales. |
free tierno free tier
billingcustom quote; public monthly and annual rates are not published
hidden costsPrice scales with crawl volume, data integrations, and modules.
pricing sources checked 2026-08-14 · pricing source ↗
Questions about Oncrawl
Can you build your own Oncrawl with AI?
A full replacement is not the recommended project. A consolation build is possible, but the paid product's decisive value sits outside a solo rebuild. For Oncrawl, crawl one owned site locally and join pages to user-supplied analytics exports. The hard boundary is large-scale crawler, log analysis, data connectors, and enterprise support, plus crawl scale, rule depth, and operational polish.
What does the Oncrawl build prompt cover?
The prompt starts with this scope: Crawl one owned site locally, inspect HTML and rendered pages, join pages to user-supplied analytics exports, explain prioritized issues, and export a reproducible audit. Full-product capabilities excluded from the comparison include: large-scale crawler, log analysis, data connectors, and enterprise support; massive hosted crawl capacity; proprietary scoring. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Oncrawl prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Oncrawl project take?
The catalogue estimate is closest consolation build: one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Oncrawl?
large-scale crawler, log analysis, data connectors, and enterprise support; massive hosted crawl capacity; proprietary scoring; continuous monitoring; agency reporting and support. People still pay for Oncrawl because a crawler is buildable; professionals pay for years of edge-case handling and reports they can trust with clients. The recurring cost buys robots handling, rendering, canonicalization, deduplication, crawl traps, rule maintenance, scheduling, storage, and false positives, not just the visible interface.
What can I use instead of building Oncrawl?
FreeCrawl: A local crawler that joins GSC and GA4 data, keeps projects and exports the lot; no cloud cluster, no invoice. Check each option's license, hosting needs and feature limits.