Sonix
Transcription, translation, subtitles, and collaborative media review
The visible automated transcription loop is buildable, but a credible replacement needs more than the first screen. Sonix earns its keep through capture, integrations, reliability, so expect a weekend or multi-day build and a narrower personal scope.
Build verification: not recorded. How we judge buildability
What you give up
- live multi-speaker accuracy
- calendar and CRM integrations
- cross-call team analytics
- meeting-bot auto-join
Why people still pay
Sonix: Customers pay for automatic capture, dependable speaker handling, search across calls, and notes arriving without manual file wrangling.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- A Python virtual environment, writable input/output directories and sufficient disk for both originals and outputs. Bind the service to localhost. Install FFmpeg and confirm codec support for the intended inputs. Download a compatible speech model and record its version; diarization, if added, has separate model and hardware requirements.
- Implementation components: Python, FastAPI and server-rendered HTML with HTMX for a local interface. SQLite for manifests and job state, with an explicit worker process and immutable source files. FFmpeg/ffprobe for explicit media operations and browser previews; never interpolate user filenames into shell commands. A locally installed faster-whisper model for transcription; optional model API only after source-text preview and consent.
- Scope boundary: live multi-speaker accuracy; calendar and CRM integrations
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Optional external skill: modern-python — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: web-design-guidelines — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: sharp-edges — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: audio/video imports, languages, model versions, timed words/segments, speaker labels and subtitle exports
Project rule — preserve this invariant: Detected language and speaker labels are uncertain until reviewed; retranscription creates a new revision rather than replacing corrections.
Project rule — acceptance evidence: An unsupported language is reported explicitly; editing one segment preserves valid neighboring timecodes and export reopens as SRT/VTT.
Implementation plan
Phase 1
Scope and fixtures. Implement this bounded workflow: Transcribe user-owned recordings in a selected supported language, edit time-aligned text and export subtitles or a searchable transcript. Keep manual corrections independent from generated model output. Record prerequisites, select representative user-owned fixtures and document the unsupported features: live multi-speaker accuracy; calendar and CRM integrations
Phase 2
Durable model. Model audio/video imports, languages, model versions, timed words/segments, speaker labels and subtitle exports Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Detected language and speaker labels are uncertain until reviewed; retranscription creates a new revision rather than replacing corrections.
Phase 3
Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals.
Phase 4
Permissions and integration failure. Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
Phase 5
Portable handoff. Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
Phase 6
Acceptance scenarios. An unsupported language is reported explicitly; editing one segment preserves valid neighboring timecodes and export reopens as SRT/VTT. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.
WORKING SLICE Transcribe user-owned recordings in a selected supported language, edit time-aligned text and export subtitles or a searchable transcript. Keep manual corrections independent from generated model output. Build this scoped Sonix-inspired workflow with a documented data model and visible failure states. Architecture - Python, FastAPI and server-rendered HTML with HTMX for a local interface. - SQLite for manifests and job state, with an explicit worker process and immutable source files. - FFmpeg/ffprobe for explicit media operations and browser previews; never interpolate user filenames into shell commands. - A locally installed faster-whisper model for transcription; optional model API only after source-text preview and consent. Prerequisites and limits A Python virtual environment, writable input/output directories and sufficient disk for both originals and outputs. Bind the service to localhost. Install FFmpeg and confirm codec support for the intended inputs. Download a compatible speech model and record its version; diarization, if added, has separate model and hardware requirements. Outside this release: live multi-speaker accuracy; calendar and CRM integrations Data model and correctness audio/video imports, languages, model versions, timed words/segments, speaker labels and subtitle exports Invariant: Detected language and speaker labels are uncertain until reviewed; retranscription creates a new revision rather than replacing corrections. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals. Security and privacy Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Recovery and export Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Implementation order 1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Transcribe user-owned recordings in a selected supported language, edit time-aligned text and export subtitles or a searchable transcript. Keep manual corrections independent from generated model output. Record prerequisites, select representative user-owned fixtures and document the unsupported features: live multi-speaker accuracy; calendar and CRM integrations 2. Phase 2 — Durable model. Model audio/video imports, languages, model versions, timed words/segments, speaker labels and subtitle exports Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Detected language and speaker labels are uncertain until reviewed; retranscription creates a new revision rather than replacing corrections. 3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals. 4. Phase 4 — Permissions and integration failure. Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results. 5. Phase 5 — Portable handoff. Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README. 6. Phase 6 — Acceptance scenarios. An unsupported language is reported explicitly; editing one segment preserves valid neighboring timecodes and export reopens as SRT/VTT. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder. Acceptance An unsupported language is reported explicitly; editing one segment preserves valid neighboring timecodes and export reopens as SRT/VTT. Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees. Optional agent guidance Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Project rule — data model: audio/video imports, languages, model versions, timed words/segments, speaker labels and subtitle exports Project rule — preserve this invariant: Detected language and speaker labels are uncertain until reviewed; retranscription creates a new revision rather than replacing corrections. Project rule — acceptance evidence: An unsupported language is reported explicitly; editing one segment preserves valid neighboring timecodes and export reopens as SRT/VTT.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md · generated from this app's build plan
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Sonix pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free trial | $0 | $0 | 30 transcription minutes total.No card required; trial, not a permanent free tier. |
| pay as you go | — | — | 1 user; 5 GB storage; transcription charged at $10 per audio/video hour.No subscription fee; usage is $10/hour. |
| core | $25/workspace | $22.92/workspace | 5 transcription/translation hours and 5 AI Workspace hours per month; 25 GB storage; 1 included user.Annual plan gives 1 month free; annual equivalent calculated from that published term. |
| advanced | $50/workspace | $45.83/workspace | 20 transcription/translation hours and 25 AI Workspace hours per month; 50 GB storage; 1 included user.Annual plan gives 1 month free; annual equivalent calculated from that published term. |
| pro | $80/workspace | $73.33/workspace | 40 transcription/translation hours and 100 AI Workspace hours per month; 100 GB storage; 1 included user.Annual plan gives 1 month free; annual equivalent calculated from that published term. |
| enterprise | — | — | Custom hours; 1 TB or more storage; unlimited team members.Contact sales. |
free tierno free tier
billingpay as you go, monthly, or annual; annual subscriptions give 1 month free; 30-minute trial
hidden costsExtra transcription hours cost $10/hour on every subscription. Extra seats cost $25/month each but add no hours; automated translation, alignment, and burned-in subtitles can also consume separately charged hours.
pricing sources checked 2026-08-14 · pricing source ↗
Questions about Sonix
Can you build your own Sonix with AI?
Partly. The visible automated transcription loop is buildable, but a credible replacement needs more than the first screen. Sonix earns its keep through capture, integrations, reliability, so expect a weekend or multi-day build and a narrower personal scope.
What does the Sonix build prompt cover?
The prompt starts with this scope: Transcribe user-owned recordings in a selected supported language, edit time-aligned text and export subtitles or a searchable transcript. Keep manual corrections independent from generated model output. Full-product capabilities excluded from the comparison include: live multi-speaker accuracy; calendar and CRM integrations; cross-call team analytics. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Sonix prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Sonix project take?
The catalogue estimate is multi-day for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Sonix?
live multi-speaker accuracy; calendar and CRM integrations; cross-call team analytics; meeting-bot auto-join. Sonix: Customers pay for automatic capture, dependable speaker handling, search across calls, and notes arriving without manual file wrangling.
What can I use instead of building Sonix?
The prior-art section lists whisper.cpp as starting points. Review their current scope, license and maintenance before adopting one.