NaturalReader
Read user-provided text and documents aloud with a small set of licensed or local voices
The core loop is buildable, but a dependable replacement becomes a real weekend or multi-day project. For NaturalReader, read user-provided text and documents aloud with a small set of licensed or local voices. The hard boundary is voice catalog, ocr, document formats, mobile apps, and commercial licensing options, plus models, compute, rights, and safety operations.
Build verification: not recorded. How we judge buildability
What you give up
- voice catalog, OCR, document formats, mobile apps, and commercial licensing options
- frontier voice or avatar model
- licensed voice catalog
- real-time rendering fleet
- moderation, consent verification, and enterprise rights
Why people still pay
People still pay for NaturalReader because customers pay for output quality, production speed, licensed voices, consent workflows, and a provider that carries the operational risk. The recurring cost buys model licensing, consent records, impersonation risk, watermarking, GPU queues, media storage, abuse response, and rapid model changes, not just the visible interface.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Python, FastAPI, SQLite, FFmpeg/ffprobe and a React review interface.
- Before starting: Installed FFmpeg/ffprobe, writable media storage and a short recording whose use is authorized.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store Document, Paragraph, PronunciationOverride, SpeechJob and AudioSegment; edited paragraphs invalidate only their derived audio and playback order follows the document revision.
Project rule — scope and recovery: Start with one supported voice engine and disclose synthesis. Do not clone a real person's voice or imply reading/OCR is flawless; scanned pages need a reviewed text-extraction step.
Project rule — acceptance: Correct a proper-name pronunciation midway through a chapter; regenerate affected segments without duplicating or skipping surrounding paragraphs.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: modern-python — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: vercel-react-best-practices — keep the proposed React work/review views responsive and avoid unnecessary rendering or data-fetch waterfalls. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Import user-owned text, choose a licensed voice, listen by paragraph and export synthetic audio with source/voice metadata. Confirm setup: Installed FFmpeg/ffprobe, writable media storage and a short recording whose use is authorized.
Phase 2
Implement persistence and write-time invariants before decorating the UI: Store Document, Paragraph, PronunciationOverride, SpeechJob and AudioSegment; edited paragraphs invalidate only their derived audio and playback order follows the document revision.
Phase 3
Connect the working view to real saved state. Keep source files immutable and store edit decisions separately. Run media tools with argument arrays, bounded file size/runtime and unique job directories; only rename a completed export into its final location.
Phase 4
Expose the app-specific limits and recovery path in context: Start with one supported voice engine and disclose synthesis. Do not clone a real person's voice or imply reading/OCR is flawless; scanned pages need a reviewed text-extraction step.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Correct a proper-name pronunciation midway through a chapter; regenerate affected segments without duplicating or skipping surrounding paragraphs. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build the following focused alternative to NaturalReader. This is a deliberately limited personal or small-team substitute, not parity with the paid service. WORKING SLICE Import user-owned text, choose a licensed voice, listen by paragraph and export synthetic audio with source/voice metadata. SETUP AND ARCHITECTURE Use Python, FastAPI, SQLite, FFmpeg/ffprobe and a React review interface. Prerequisites: Installed FFmpeg/ffprobe, writable media storage and a short recording whose use is authorized. Before integrating anything, record actual versions and permissions, plus model files or provider limits only where used, in the README; make unavailable dependencies visible rather than simulating success. DOMAIN MODEL AND INVARIANTS Store Document, Paragraph, PronunciationOverride, SpeechJob and AudioSegment; edited paragraphs invalidate only their derived audio and playback order follows the document revision. IMPLEMENTATION CONTRACT Keep source files immutable and store edit decisions separately. Run media tools with argument arrays, bounded file size/runtime and unique job directories; only rename a completed export into its final location. Provide an input/setup view, the main work view, and a review/export view appropriate to this workflow. Preserve the last saved state if a job or save fails. Include empty, loading, permission-denied, partial and retryable-error states. Log identifiers and error categories without secret values or unnecessary private content. APP-SPECIFIC BOUNDARY AND RECOVERY Start with one supported voice engine and disclose synthesis. Do not clone a real person's voice or imply reading/OCR is flawless; scanned pages need a reviewed text-extraction step. ACCEPTANCE SCENARIO Correct a proper-name pronunciation midway through a chapter; regenerate affected segments without duplicating or skipping surrounding paragraphs. Also reopen the app after an interrupted operation, confirm the saved record/export remains inspectable, and document the recovery action. These are implementation acceptance requirements, not a claim that this guide has been tested. DELIVERY Deliver a runnable repository with migrations or project-format versioning, a non-sensitive example, environment/permission setup, the exact manual acceptance steps, and a backup/export-and-restore walkthrough. Implement the working slice before optional integrations; list any deferred paid-product capabilities honestly. Do not add capabilities outside the working slice just to resemble the original product. PROJECT RULES FOR AGENTS.md Keep the domain invariants above executable at the write boundary. Propose scope changes before adding providers or permissions. Never fabricate source evidence, publish results, identity matches or successful delivery. Preserve user originals and require an explicit confirmation for destructive changes or external publication.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Alternatives to building your own
all 4 free alternatives to NaturalReader →· no votes, no pay-to-list · just what's real
NaturalReader pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free | $0/user | $0/user | Unlimited system Free Voices; 20,000 Lite-voice characters/day; 4,000 shared Plus/Cloned characters/day; no MP3 conversion or OCR. |
| lite | $13.90/user | $6.58/user | Unlimited Lite listening; 1,000,000 Lite-voice MP3 characters/month; no Plus/Pro voices or voice cloning.Annual total is $79. |
| plus | $20.90/user | $9.92/user | 500,000 shared Plus/Cloned listening characters/day; 1,000,000 Lite MP3 characters/month + 1,000,000 shared Plus/Cloned MP3 characters/month; clone up to 2 voices.Annual total is $119. |
| pro | $25.90/user | $13.25/user | 500,000 shared Plus/Cloned/Pro listening characters/day; 1,000,000 Lite MP3 characters/month + 1,000,000 shared premium MP3 characters/month; clone up to 2 voices.Annual total is $159. |
| lite edu | — | $49.92 | Annual packages: 50 users $599; 100 $700; 150 $800; 200 $900; 250 $1,000; 500 $1,800; 1,000 $2,500; 1,500 $2,900; 2,000 $3,300.Annual-only; price field is the effective monthly cost of the 50-user starting package, not a per-user rate. |
| plus edu | — | $24.92 | Annual packages: 5 users $299; 10 $499; 20 $899; 30 $1,200; 40 $1,400; 50 $1,500; above 50 starts at $25/user/year.Annual-only; price field is the effective monthly cost of the 5-user starting package. |
| site license | — | $275 | Minimum 2,000 enrolled users; starts at $3,300/year.Annual-only; pricing is based on total enrollment. |
free tierunlimited system voices; 20,000 Lite-voice characters/day; 4,000 Plus/Cloned characters/day; no MP3 or OCR
billingindividual plans are monthly + annual; EDU and Site License plans are annual-only
hidden costsAudio made under Personal/EDU plans is licensed for personal use only; commercial rights require the separately billed Commercial AI Voice Generator. App Store/Google Play prices, taxes, and regional fees can differ.
pricing sources checked 2026-08-12 · pricing source ↗
Questions about NaturalReader
Can you build your own NaturalReader with AI?
Partly. The core loop is buildable, but a dependable replacement becomes a real weekend or multi-day project. For NaturalReader, read user-provided text and documents aloud with a small set of licensed or local voices. The hard boundary is voice catalog, ocr, document formats, mobile apps, and commercial licensing options, plus models, compute, rights, and safety operations.
What does the NaturalReader build prompt cover?
The prompt starts with this scope: Import user-owned text, choose a licensed voice, listen by paragraph and export synthetic audio with source/voice metadata. Full-product capabilities excluded from the comparison include: voice catalog, OCR, document formats, mobile apps, and commercial licensing options; frontier voice or avatar model; licensed voice catalog. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the NaturalReader prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this NaturalReader project take?
The catalogue estimate is closest consolation build: one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing NaturalReader?
voice catalog, OCR, document formats, mobile apps, and commercial licensing options; frontier voice or avatar model; licensed voice catalog; real-time rendering fleet; moderation, consent verification, and enterprise rights. People still pay for NaturalReader because customers pay for output quality, production speed, licensed voices, consent workflows, and a provider that carries the operational risk. The recurring cost buys model licensing, consent records, impersonation risk, watermarking, GPU queues, media storage, abuse response, and rapid model changes, not just the visible interface.
What price is this guide comparing against?
The recorded Lite plan is $13.9/mo per seat (monthly per user), checked 2026-08-12. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building NaturalReader?
Koodo Reader: Reads nearly every ebook format aloud; the fancy hosted voice is paid, the local and system voices are not. Readest: A book-and-PDF reader with built-in narration and no premium-voice bait-and-switch. Thorium Reader: A free desktop reader built by accessibility people: ebooks, PDFs and text-to-speech, with no account or ad slot. Compare all listed options at https://howtovibecodeit.dev/naturalreader/alternatives. Check each option's license, hosting needs and feature limits.