# AGENTS.md — Build guide for Vapi

## Project scope
Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup.

Catalogue verdict: kinda. A focused inbound agent is a credible weekend build with LiveKit Agents or Pipecat. Replacing Vapi as a product is much larger: it combines realtime orchestration, provider abstraction, phone-number and SIP workflows, outbound campaigns, tool calls, testing, observability, scaling, support, and compliance options. Self-hosting can remove Vapi's $0.05 per-minute platform fee, but model, carrier, number, infrastructure, failed-call, monitoring, and engineering costs remain.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- A user-owned server, LiveKit/SIP configuration, one provisioned inbound number/trunk, provider credentials, consent language and a human fallback destination. Confirm installed SDK APIs before implementation.
- Implementation components: Python and a documented supported LiveKit Agents release with one named worker. LiveKit Server, Redis and SIP services; explicit STT/LLM/TTS/VAD adapters and SQLite call outcome records.
- Scope boundary: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities.

## Stack and architecture
- Python and a documented supported LiveKit Agents release with one named worker.
- LiveKit Server, Redis and SIP services; explicit STT/LLM/TTS/VAD adapters and SQLite call outcome records.
- Domain model: inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records

## Security and data integrity
- Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls.
- Correctness boundary: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
- Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect.
- Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call.

## Agent implementation rules
- Project rule — data model: inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records
- Project rule — preserve this invariant: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
- Project rule — acceptance evidence: Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result.

## Optional agent skills and references
- Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate the actual Vapi-inspired workflow with owned or clearly labeled sample data: Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup.
- Publish a reproducible walkthrough with this observable result: Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result.
- Explain who can operate this scoped tool, its setup and ongoing costs, and these remaining product gaps: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities. Avoid guaranteed savings, performance scores or implied endorsement.

## Engineering roadmap
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities.
2. Phase 2 — Durable model. Model inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect.
4. Phase 4 — Permissions and integration failure. Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

## Paid-product capabilities outside this build
- managed realtime infrastructure, scaling, provider failover, and support
- assistant, squad, workflow, phone-number, and provider configuration APIs
- outbound campaigns, carrier integrations, spam mitigation, and production telephony debugging
- hosted simulations, evaluations, call logs, recordings, analytics, and retention controls
- enterprise SLAs, compliance add-ons, role controls, and data-residency options

## Implementation prompt
WORKING SLICE
Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup.

Build this scoped Vapi-inspired workflow with a documented data model and visible failure states.

Architecture
- Python and a documented supported LiveKit Agents release with one named worker.
- LiveKit Server, Redis and SIP services; explicit STT/LLM/TTS/VAD adapters and SQLite call outcome records.

Prerequisites and limits
A user-owned server, LiveKit/SIP configuration, one provisioned inbound number/trunk, provider credentials, consent language and a human fallback destination. Confirm installed SDK APIs before implementation.
Outside this release: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities.

Data model and correctness
inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records
Invariant: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect.

Security and privacy
Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls.

Recovery and export
Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Build one inbound support voice agent using a supported LiveKit Agents release, a configured SIP path and explicit STT/LLM/TTS providers. Handle interruptions, offer a human fallback and allow one bounded server-side lookup. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Telephony provisioning, emergency handling and guaranteed conversational accuracy are separate responsibilities.
2. Phase 2 — Durable model. Model inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Persist the call ID and tool receipts, cancel speech on interruption and bound tool latency. Treat uncertain tool outcomes as a handoff condition rather than repeating a side effect.
4. Phase 4 — Permissions and integration failure. Authenticate SIP/agent connections and constrain callable tools and destinations. Caller speech is untrusted input and cannot grant access or override policy. The single lookup reads an allowlisted, non-sensitive support knowledge source only; no account-specific data or external writes are exposed. Authenticated account lookups and write actions require a separate scope and server-verified caller authorization. Keep provider keys server-side, bound call duration/spend and apply explicit recording/transcript consent and retention controls. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Retain only the chosen call records under a clear retention policy; export outcome/transcript data with access control. A worker restart must not resurrect a finished call. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: inbound call sessions, participant consent, agent state, tool approvals, interruption events and outcome records
Project rule — preserve this invariant: A caller's speech cannot expand tool permissions; this release exposes only an allowlisted non-sensitive support lookup, excludes account-data access and external writes, and records only under an explicit consent/retention policy.
Project rule — acceptance evidence: Interrupt the agent mid-sentence and stop obsolete speech; a failed lookup yields a handoff or honest failure rather than an invented account result.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
