ekto
Real-time voice translation: speak in one language, get the other side rendered back in near real time.
The pipeline is no longer exotic: capture mic audio, segment it with a voice activity detector, transcribe with Whisper, translate, speak it back with a local TTS voice. An agent can wire that into a working local app in a weekend and it will genuinely translate a conversation. What it will not do out of the gate is stay graceful for an hour: chunk boundaries clip words, speaker turns bleed together, latency creeps as the buffer grows, and the sentence by sentence pacing that makes these apps usable in real conversation is a tuning problem, not a coding problem. You also get no phone app, which is where voice translation actually happens. Fine for a desk setup and travel prep, unconvincing when you are holding it out to a stranger in a market.
Build verification: not recorded. How we judge buildability
What you give up
- Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes
- Clean sentence by sentence pacing and turn detection, which is most of the perceived quality
- A mobile app, so no translating anything while standing up
- Offline or low-bandwidth behavior tuned for actual travel
- Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment
Why people still pay
Because voice translation is judged entirely on the seconds between someone finishing a sentence and you hearing it, and on whether it still works on minute 40. A local build nails the demo and then frays: barge-in, background noise, two people talking over each other, the phone locking. Paying gets you a phone in your pocket that handles those cases without you adding VAD thresholds mid-conversation.
Your build guide
The stack, security requirements, and agent rules for a focused replacement.
Before you start
- Runtime and tools: Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor.
- Before starting: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
Use these project rules and optional skill references alongside the prompt. Review each skill before adding it to your agent; the AGENTS.md export includes the same guidance.
Project rule — domain: Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback.
Project rule — scope and recovery: Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.
Project rule — acceptance: Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation.
Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.
Recommended skill: modern-python — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Recommended skill: web-design-guidelines — review keyboard access, focus, validation, error recovery and the readable work/review interface or HTML report. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
Implementation plan
Phase 1
Pin the working slice and create its example input: Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review. Confirm setup: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
Phase 2
Implement persistence and write-time invariants before decorating the UI: Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback.
Phase 3
Connect the working view to real saved state. Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words.
Phase 4
Expose the app-specific limits and recovery path in context: Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.
Phase 5
Walk through this concrete acceptance case and preserve its exported evidence: Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.
Build the following focused alternative to ekto. This is a deliberately limited personal or small-team substitute, not parity with the paid service. WORKING SLICE Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review. SETUP AND ARCHITECTURE Use Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor. Prerequisites: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key. Before integrating anything, record actual versions and permissions, plus model files or provider limits only where used, in the README; make unavailable dependencies visible rather than simulating success. DOMAIN MODEL AND INVARIANTS Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback. IMPLEMENTATION CONTRACT Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words. Provide an input/setup view, the main work view, and a review/export view appropriate to this workflow. Preserve the last saved state if a job or save fails. Include empty, loading, permission-denied, partial and retryable-error states. Log identifiers and error categories without secret values or unnecessary private content. APP-SPECIFIC BOUNDARY AND RECOVERY Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation. ACCEPTANCE SCENARIO Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation. Also reopen the app after an interrupted operation, confirm the saved record/export remains inspectable, and document the recovery action. These are implementation acceptance requirements, not a claim that this guide has been tested. DELIVERY Deliver a runnable repository with migrations or project-format versioning, a non-sensitive example, environment/permission setup, the exact manual acceptance steps, and a backup/export-and-restore walkthrough. Implement the working slice before optional integrations; list any deferred paid-product capabilities honestly. Do not add capabilities outside the working slice just to resemble the original product. PROJECT RULES FOR AGENTS.md Keep the domain invariants above executable at the write boundary. Propose scope changes before adding providers or permissions. Never fabricate source evidence, publish results, identity matches or successful delivery. Preserve user originals and require an explicit confirmation for destructive changes or external publication.
$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy or download AGENTS.md
prompt copied. want to know what dies next week?
new verdicts + top votes, weekly. free. one-click out.
No prior-art project is listed yet. Compare the scoped build with the paid product before choosing.
Questions about ekto
Can you build your own ekto with AI?
Partly. The pipeline is no longer exotic: capture mic audio, segment it with a voice activity detector, transcribe with Whisper, translate, speak it back with a local TTS voice. An agent can wire that into a working local app in a weekend and it will genuinely translate a conversation. What it will not do out of the gate is stay graceful for an hour: chunk boundaries clip words, speaker turns bleed together, latency creeps as the buffer grows, and the sentence by sentence pacing that makes these apps usable in real conversation is a tuning problem, not a coding problem. You also get no phone app, which is where voice translation actually happens. Fine for a desk setup and travel prep, unconvincing when you are holding it out to a stranger in a market.
What does the ekto build prompt cover?
The prompt starts with this scope: Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review. Full-product capabilities excluded from the comparison include: Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes; Clean sentence by sentence pacing and turn detection, which is most of the perceived quality; A mobile app, so no translating anything while standing up. Follow the implementation plan and its prerequisites before expanding the build.
How do I use the prompt, AGENTS.md and agent skills?
Start with the ekto prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.
How long will this ekto project take?
The catalogue estimate is a weekend for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.
What would I give up by replacing ekto?
Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes; Clean sentence by sentence pacing and turn detection, which is most of the perceived quality; A mobile app, so no translating anything while standing up; Offline or low-bandwidth behavior tuned for actual travel; Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment. Because voice translation is judged entirely on the seconds between someone finishing a sentence and you hearing it, and on whether it still works on minute 40. A local build nails the demo and then frays: barge-in, background noise, two people talking over each other, the phone locking. Paying gets you a phone in your pocket that handles those cases without you adding VAD thresholds mid-conversation.
What price is this guide comparing against?
The recorded Monthly Unlimited PRO plan is $29.99/mo (monthly subscription), checked 2026-08-18. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.
What can I use instead of building ekto?
No alternative is listed in this entry yet. That is a gap in this catalogue, not proof that no suitable product exists. Compare the paid product and the proposed scope before committing to a build.