Wispr Flow
Hotkey → record → Whisper → paste at cursor
Hotkey → record → Whisper → paste at cursor. One of the most-cloned apps of the trend for a reason.
Build verification: not recorded. How we judge buildability
What you give up
- their tuned auto-editing voice model
- per-app tone formatting
- the mobile keyboard
- polish on edge cases (accents, noise)
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Swift, SwiftUI/AppKit and SQLite on macOS, with secrets in Keychain where needed.
- Before starting: A Mac, the current Xcode toolchain and a documented permission checklist; distribution signing is a separate delivery task.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store SessionState, CapturedAppIdentity, AudioTempPath and PasteboardRevision; cancellation terminates work and late results cannot change a cancelled session or paste into a different app.
Project rule — scope and recovery: Retain the pinned whisper.cpp local path and no-cloud initial scope. Permission denial, five-minute stop, temp-file cleanup and manual insertion are required; no guarantee that every destination accepts simulated paste.
Project rule — acceptance: Switch apps while recognition runs and copy new clipboard text before insertion; keep the transcript ready for manual action and preserve the new clipboard value.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: swiftui-expert-skill — design native SwiftUI state, accessible controls and permission/error views for the proposed Apple-platform interface. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: swift-concurrency — isolate capture, background work and UI updates, and make cancellation invalidate late callbacks. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Start/stop local dictation from a shortcut, review the result and insert only into the originally selected app with a copy-only fallback. Confirm setup: A Mac, the current Xcode toolchain and a documented permission checklist; distribution signing is a separate delivery task.
Phase 2
Implement persistence and write-time invariants before decorating the UI: Store SessionState, CapturedAppIdentity, AudioTempPath and PasteboardRevision; cancellation terminates work and late results cannot change a cancelled session or paste into a different app.
Phase 3
Connect the working view to real saved state. Ask for each operating-system permission when its feature is used. Handle denial and revocation without a loop; show permission status and keep local history deletable.
Phase 4
Expose the app-specific limits and recovery path in context: Retain the pinned whisper.cpp local path and no-cloud initial scope. Permission denial, five-minute stop, temp-file cleanup and manual insertion are required; no guarantee that every destination accepts simulated paste.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Switch apps while recognition runs and copy new clipboard text before insertion; keep the transcript ready for manual action and preserve the new clipboard value. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build me a macOS menu-bar dictation tool with local transcription like Wispr Flow. Requirements: - Use Swift, AppKit, AVFoundation, and a pinned whisper.cpp command-line build. Store settings under Application Support/LocalDictation and keep the downloaded model at a configured local path. No cloud provider or API key is needed for the initial version. - Implement states idle, recording, transcribing, ready, cancelled, and error. A configurable global shortcut starts/stops recording; an optional hold-to-talk mode stops on release. Show a recording indicator with elapsed time and auto-stop after five minutes. - Capture microphone audio with AVAudioEngine and convert to the mono sample format expected by the pinned transcription binary. Write a unique temporary audio file, spawn the binary with an argument array, and capture bounded stdout/stderr. Never invoke a shell with dictated text or filenames. - Escape cancels recording or transcription. Cancellation terminates the child process, deletes its temporary audio, and marks the session cancelled so a late callback cannot paste text. Denied microphone or Accessibility permission must show an actionable state with no recording underway. - Capture the frontmost app identity when recording starts. After transcription, display editable text and an Insert button; optional auto-insert is allowed only if the same app remains frontmost. Otherwise keep the result for manual insertion. Do not inspect or store unrelated screen content. - Insert via the pasteboard and a simulated Cmd-V when Accessibility permission allows it. Offer copy-only fallback. If restoring the previous clipboard is enabled, restore only if the pasteboard change count still matches this app write, so a new user copy is not overwritten. - Keep audio only until transcription or cancellation completes; transcript history is off by default. Apply only deterministic whitespace/punctuation cleanup in v1, retaining the raw transcript for review. No accounts or telemetry. Out of scope: mobile keyboard, per-app AI tone changes, and guaranteed insertion into every app. - Acceptance: start/stop, five-minute stop, denied permission, Escape during recognition, focus change before completion, and clipboard change during insertion all produce safe states. README covers Xcode build, model download/checksum, microphone/Accessibility permissions, privacy behavior, and manual fallback. EDITORIAL IMPLEMENTATION CONTRACT Working slice: Start/stop local dictation from a shortcut, review the result and insert only into the originally selected app with a copy-only fallback. Data and invariants: Store SessionState, CapturedAppIdentity, AudioTempPath and PasteboardRevision; cancellation terminates work and late results cannot change a cancelled session or paste into a different app. Boundary and recovery: Retain the pinned whisper.cpp local path and no-cloud initial scope. Permission denial, five-minute stop, temp-file cleanup and manual insertion are required; no guarantee that every destination accepts simulated paste. Acceptance walkthrough: Switch apps while recognition runs and copy new clipboard text before insertion; keep the transcript ready for manual action and preserve the new clipboard value. Record actual dependency versions, permissions and provider access in setup instructions. Preserve originals, expose partial failures and document backup/restore. These are acceptance requirements, not a claim of a completed or production-certified build. Add the domain, recovery and acceptance rules to AGENTS.md so future edits preserve them.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
Alternatives to building your own
all 8 free alternatives to Wispr Flow →· no votes, no pay-to-list · just what's real
Wispr Flow pricing
| plan | monthly | annual (per mo) | what you get |
|---|---|---|---|
| free | $0 | $0 | Desktop: 2,000 dictated words/week; iPhone: 1,000 words/week; Android: unlimited words; recordings under 5 minutes do not count.Meeting-notetaker weekly cap is mentioned but its number is not public. |
| pro | $15/user | $12/user | Unlimited dictation across supported devices; 1 paid user.20% annual discount. |
| enterprise | — | — | Custom paid-user count; admin seats are free and no minimum seat count is published.Volume discounts are custom. |
free tierDesktop 2,000 words/week; iPhone 1,000 words/week; Android unlimited; recordings under 5 minutes do not count.
billingmonthly + annual (20% off); no commitment and cancel anytime
hidden costsEnterprise admin seats are free, but regular users are billed; volume discounts are custom. Students and educators with accepted .edu eligibility can receive Pro free, while nonprofit discounts are custom.
pricing sources checked 2026-08-11 · pricing source ↗
Questions about Wispr Flow
Can you build your own Wispr Flow with AI?
The verdict is yes for the scoped workflow. Hotkey → record → Whisper → paste at cursor. One of the most-cloned apps of the trend for a reason.
What does the Wispr Flow build prompt cover?
The prompt starts with this scope: Start/stop local dictation from a shortcut, review the result and insert only into the originally selected app with a copy-only fallback. Full-product capabilities excluded from the comparison include: their tuned auto-editing voice model; per-app tone formatting; the mobile keyboard. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the Wispr Flow prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this Wispr Flow project take?
The catalogue estimate is one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing Wispr Flow?
their tuned auto-editing voice model; per-app tone formatting; the mobile keyboard; polish on edge cases (accents, noise). Compare these limits with the workflow you actually need.
What price is this guide comparing against?
The recorded paid plan is $15/mo (monthly), checked 2026-08-07. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building Wispr Flow?
Handy: The literal replacement: press, record, transcribe locally, paste. FluidVoice: Local Mac dictation with on-device cleanup and no weekly word allowance. OpenWhispr: Hotkey, local model, cleaned text at the cursor, plus searchable history and meeting capture. Compare all listed options at https://howtovibecodeit.dev/wispr-flow/alternatives. Check each option's license, hosting needs and feature limits.