# AGENTS.md — Build guide for ekto

## Project scope
Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review.

Catalogue verdict: kinda. The pipeline is no longer exotic: capture mic audio, segment it with a voice activity detector, transcribe with Whisper, translate, speak it back with a local TTS voice. An agent can wire that into a working local app in a weekend and it will genuinely translate a conversation. What it will not do out of the gate is stay graceful for an hour: chunk boundaries clip words, speaker turns bleed together, latency creeps as the buffer grows, and the sentence by sentence pacing that makes these apps usable in real conversation is a tuning problem, not a coding problem. You also get no phone app, which is where voice translation actually happens. Fine for a desk setup and travel prep, unconvincing when you are holding it out to a stranger in a market.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- Runtime and tools: Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor.
- Before starting: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.

## Stack and architecture
- Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor
- Data design: Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback.
- Setup: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key

## Security and data integrity
- Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words.
- Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.
- Keep secrets outside client bundles and exported projects; document what leaves the device and make retention/deletion controls visible.

## Agent implementation rules
- Project rule — domain: Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback.
- Project rule — scope and recovery: Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.
- Project rule — acceptance: Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation.
- Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.

## Optional agent skills and references
- Recommended skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — structure the Python worker or explicitly optional read-only utility with pinned dependencies, typed boundaries and clear failure handling. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
- Recommended skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — review keyboard access, focus, validation, error recovery and the readable work/review interface or HTML report. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate this working slice using synthetic or explicitly authorized non-sensitive examples: Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review.
- Share a synthetic example export and the acceptance walkthrough; keep real customer, health, financial and source data private: Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation.
- State the limits before asking someone to replace their existing tool: Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.

## Engineering roadmap
1. Phase 1 — Pin the working slice and create its example input: Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review. Confirm setup: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key.
2. Phase 2 — Implement persistence and write-time invariants before decorating the UI: Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback.
3. Phase 3 — Connect the working view to real saved state. Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words.
4. Phase 4 — Expose the app-specific limits and recovery path in context: Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.
5. Phase 5 — Walk through this concrete acceptance case and preserve its exported evidence: Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.

## Paid-product capabilities outside this build
- Long session reliability: memory growth, drifting segmentation and dropped turns after the first 20 minutes
- Clean sentence by sentence pacing and turn detection, which is most of the perceived quality
- A mobile app, so no translating anything while standing up
- Offline or low-bandwidth behavior tuned for actual travel
- Latency budgets someone else already fought for: streaming partial results instead of waiting for a full segment

## Implementation prompt
Build the following focused alternative to ekto. This is a deliberately limited personal or small-team substitute, not parity with the paid service.

WORKING SLICE
Capture a short spoken turn, display the transcript, translate it into one chosen language and play an optional synthetic rendition after review.

SETUP AND ARCHITECTURE
Use Python, FastAPI, SQLite, FFmpeg and a local faster-whisper worker with a React transcript editor. Prerequisites: An installed speech model, adequate local disk space and a recording made with participant permission; any optional hosted model needs a separately disclosed API key. Before integrating anything, record actual versions and permissions, plus model files or provider limits only where used, in the README; make unavailable dependencies visible rather than simulating success.

DOMAIN MODEL AND INVARIANTS
Store AudioTurn, TranscriptRevision, TranslationRevision and SpeechOutput; each translation references the exact source turn and cancellation invalidates queued playback.

IMPLEMENTATION CONTRACT
Keep original timing alongside corrected text. Speech recognition does not itself establish speaker identity; permit manual speaker labels. Show undecodable audio and uncertain passages without inventing words. Provide an input/setup view, the main work view, and a review/export view appropriate to this workflow. Preserve the last saved state if a job or save fails. Include empty, loading, permission-denied, partial and retryable-error states. Log identifiers and error categories without secret values or unnecessary private content.

APP-SPECIFIC BOUNDARY AND RECOVERY
Begin with turn-based translation to avoid unproven simultaneous latency. Configure supported language/voice models and licenses explicitly; never use this as a substitute for critical interpretation.

ACCEPTANCE SCENARIO
Pause mid-sentence, correct a proper name and cancel the next turn; only the reviewed translation is spoken and no stale audio plays after cancellation. Also reopen the app after an interrupted operation, confirm the saved record/export remains inspectable, and document the recovery action. These are implementation acceptance requirements, not a claim that this guide has been tested.

DELIVERY
Deliver a runnable repository with migrations or project-format versioning, a non-sensitive example, environment/permission setup, the exact manual acceptance steps, and a backup/export-and-restore walkthrough. Implement the working slice before optional integrations; list any deferred paid-product capabilities honestly. Do not add capabilities outside the working slice just to resemble the original product.

PROJECT RULES FOR AGENTS.md
Keep the domain invariants above executable at the write boundary. Propose scope changes before adding providers or permissions. Never fabricate source evidence, publish results, identity matches or successful delivery. Preserve user originals and require an explicit confirmation for destructive changes or external publication.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
