MacWhisper
Mac app for local Whisper transcription, subtitles, and batch processing
MacWhisper's solo core is compact: build a private local transcription pipeline that accepts an audio file, transcribes it, creates structured notes, and exports Markdown. A competent builder can reach a useful personal version in one sitting, while the paid product mainly wins on capture, integrations, reliability.
Build verification: not recorded. How we judge buildability
What you give up
- cross-call team analytics
- meeting-bot auto-join
- live multi-speaker accuracy
- calendar and CRM integrations
Why people still pay
MacWhisper: Customers pay for automatic capture, dependable speaker handling, search across calls, and notes arriving without manual file wrangling.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- A Python virtual environment, writable input/output directories and sufficient disk for both originals and outputs. Bind the service to localhost. Install FFmpeg and confirm codec support for the intended inputs. Download a compatible speech model and record its version; diarization, if added, has separate model and hardware requirements. This guide is a local transcription workbench rather than a clone of the native Mac application.
- Implementation components: Python, FastAPI and server-rendered HTML with HTMX for a local interface. SQLite for manifests and job state, with an explicit worker process and immutable source files. FFmpeg/ffprobe for explicit media operations and browser previews; never interpolate user filenames into shell commands. A locally installed faster-whisper model for transcription; optional model API only after source-text preview and consent.
- Scope boundary: cross-call team analytics; meeting-bot auto-join
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Optional external skill: modern-python — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: web-design-guidelines — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: sharp-edges — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: audio imports, model configurations, transcription jobs, segment timestamps, speaker labels and subtitle revisions
Project rule — preserve this invariant: Speaker identities remain user-assigned labels; model changes create a new transcript revision and do not overwrite manual corrections.
Project rule — acceptance evidence: Cancel one batch job without losing other outputs; a corrected subtitle reopens with valid nondecreasing timecodes and an unchanged source hash.
Implementation plan
Phase 1
Scope and fixtures. Implement this bounded workflow: Import a recording, transcribe locally, correct timestamped text and export SRT, VTT and Markdown. Add a batch queue and manual speaker renaming before optional diarization. Record prerequisites, select representative user-owned fixtures and document the unsupported features: cross-call team analytics; meeting-bot auto-join
Phase 2
Durable model. Model audio imports, model configurations, transcription jobs, segment timestamps, speaker labels and subtitle revisions Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Speaker identities remain user-assigned labels; model changes create a new transcript revision and do not overwrite manual corrections.
Phase 3
Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals.
Phase 4
Permissions and integration failure. Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
Phase 5
Portable handoff. Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
Phase 6
Acceptance scenarios. Cancel one batch job without losing other outputs; a corrected subtitle reopens with valid nondecreasing timecodes and an unchanged source hash. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.
WORKING SLICE Import a recording, transcribe locally, correct timestamped text and export SRT, VTT and Markdown. Add a batch queue and manual speaker renaming before optional diarization. Build this scoped MacWhisper-inspired workflow with a documented data model and visible failure states. Architecture - Python, FastAPI and server-rendered HTML with HTMX for a local interface. - SQLite for manifests and job state, with an explicit worker process and immutable source files. - FFmpeg/ffprobe for explicit media operations and browser previews; never interpolate user filenames into shell commands. - A locally installed faster-whisper model for transcription; optional model API only after source-text preview and consent. Prerequisites and limits A Python virtual environment, writable input/output directories and sufficient disk for both originals and outputs. Bind the service to localhost. Install FFmpeg and confirm codec support for the intended inputs. Download a compatible speech model and record its version; diarization, if added, has separate model and hardware requirements. This guide is a local transcription workbench rather than a clone of the native Mac application. Outside this release: cross-call team analytics; meeting-bot auto-join Data model and correctness audio imports, model configurations, transcription jobs, segment timestamps, speaker labels and subtitle revisions Invariant: Speaker identities remain user-assigned labels; model changes create a new transcript revision and do not overwrite manual corrections. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals. Security and privacy Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Recovery and export Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Implementation order 1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Import a recording, transcribe locally, correct timestamped text and export SRT, VTT and Markdown. Add a batch queue and manual speaker renaming before optional diarization. Record prerequisites, select representative user-owned fixtures and document the unsupported features: cross-call team analytics; meeting-bot auto-join 2. Phase 2 — Durable model. Model audio imports, model configurations, transcription jobs, segment timestamps, speaker labels and subtitle revisions Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Speaker identities remain user-assigned labels; model changes create a new transcript revision and do not overwrite manual corrections. 3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals. 4. Phase 4 — Permissions and integration failure. Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results. 5. Phase 5 — Portable handoff. Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README. 6. Phase 6 — Acceptance scenarios. Cancel one batch job without losing other outputs; a corrected subtitle reopens with valid nondecreasing timecodes and an unchanged source hash. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder. Acceptance Cancel one batch job without losing other outputs; a corrected subtitle reopens with valid nondecreasing timecodes and an unchanged source hash. Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees. Optional agent guidance Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission. Project rule — data model: audio imports, model configurations, transcription jobs, segment timestamps, speaker labels and subtitle revisions Project rule — preserve this invariant: Speaker identities remain user-assigned labels; model changes create a new transcript revision and do not overwrite manual corrections. Project rule — acceptance evidence: Cancel one batch job without losing other outputs; a corrected subtitle reopens with valid nondecreasing timecodes and an unchanged source hash.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Alternatives to building your own
BBuzzMacWhisper’s basic job, minus the Mac-only part and the invoice.open source↗SSubtitle EditA subtitle workbench with Whisper bolted in; ugly enough to be trustworthy, useful enough to stay installed.open source↗VVibeA cross-platform desktop transcriber with subtitles, batch jobs, translation and optional local summaries.open source↗all 4 free alternatives to MacWhisper →· no votes, no pay-to-list · just what's real
MacWhisper pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free for mac | $0/user | $0/user | Local transcription with smaller Whisper models and core export tools; no published numeric file/minute cap |
| macwhisper pro direct | — | — | Lifetime updates; soft activation limit of 3 devices€64 one-time, converted to about $73.82 at ECB EUR/USD 1.1534. |
| mac app store subscription | — | — | Weekly, monthly, and yearly subscription options exist; exact US prices were not exposed on the official public pagesRecurring App Store purchase rather than the direct lifetime license. |
| macwhisper for ios | $0/user | $0/user | All iOS app features and local models are free; no subscription required |
free tierMac free tier: local transcription with smaller models and no published numeric file/minute cap; iOS features and local models are free.
billingdirect Mac license is one-time; Mac App Store edition offers weekly/monthly/yearly subscriptions; iOS is free
hidden costsCloud transcription and LLM actions can require the user's own paid API key or provider account. Direct Pro activation is softly limited to 3 devices; volume-license packs are sold separately.
pricing sources checked 2026-08-14 · pricing source ↗
Questions about MacWhisper
Can you build your own MacWhisper with AI?
The verdict is yes for the scoped workflow. MacWhisper's solo core is compact: build a private local transcription pipeline that accepts an audio file, transcribes it, creates structured notes, and exports Markdown. A competent builder can reach a useful personal version in one sitting, while the paid product mainly wins on capture, integrations, reliability.
What does the MacWhisper build prompt cover?
The prompt starts with this scope: Import a recording, transcribe locally, correct timestamped text and export SRT, VTT and Markdown. Add a batch queue and manual speaker renaming before optional diarization. Full-product capabilities excluded from the comparison include: cross-call team analytics; meeting-bot auto-join; live multi-speaker accuracy. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the MacWhisper prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this MacWhisper project take?
The catalogue estimate is one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing MacWhisper?
cross-call team analytics; meeting-bot auto-join; live multi-speaker accuracy; calendar and CRM integrations. MacWhisper: Customers pay for automatic capture, dependable speaker handling, search across calls, and notes arriving without manual file wrangling.
What can I use instead of building MacWhisper?
Buzz: MacWhisper’s basic job, minus the Mac-only part and the invoice. Subtitle Edit: A subtitle workbench with Whisper bolted in; ugly enough to be trustworthy, useful enough to stay installed. Vibe: A cross-platform desktop transcriber with subtitles, batch jobs, translation and optional local summaries. Compare all listed options at https://howtovibecodeit.dev/macwhisper/alternatives. Check each option's license, hosting needs and feature limits.