# AGENTS.md — Build guide for Vyravid

## Project scope
Create a script, divide it into editable scenes, request or import authorized voice/image assets and assemble a simple timeline. Preview each scene and render a local video with explicit attribution and provider-cost records.

Catalogue verdict: kinda. A local version of the core workflow is buildable: project dashboard, script-to-scenes generation, voiceover, media generation, timeline assembly, and local ffmpeg rendering. The hard parts are not the CRUD shell but the media edge cases, provider integrations, long-running render jobs, and the editor polish that makes the timeline pleasant instead of fragile.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- A supported Node release, a writable local data directory and a separate backup location. Bind to localhost; remote use requires authentication and HTTPS first. FFmpeg, permitted media and provider credentials only for explicitly enabled generation steps.
- Implementation components: Node.js, TypeScript and Express with server-rendered HTML and small browser modules. SQLite through better-sqlite3 with migrations, prepared statements and a single background worker. FFmpeg rendering, a configured TTS provider and optional image-generation adapter; immutable local asset manifests.
- Scope boundary: someone else's servers & support; provider-specific voice/image/video integration details

## Stack and architecture
- Node.js, TypeScript and Express with server-rendered HTML and small browser modules.
- SQLite through better-sqlite3 with migrations, prepared statements and a single background worker.
- FFmpeg rendering, a configured TTS provider and optional image-generation adapter; immutable local asset manifests.
- Domain model: scripts, scene IDs, voiceover chunks, asset provenance, timeline intervals and render manifests

## Security and data integrity
- Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs.
- Correctness boundary: Generated assets are not silently regenerated on retry; scene duration follows actual audio and missing assets block final export.
- Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals.
- Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together.

## Agent implementation rules
- Project rule — data model: scripts, scene IDs, voiceover chunks, asset provenance, timeline intervals and render manifests
- Project rule — preserve this invariant: Generated assets are not silently regenerated on retry; scene duration follows actual audio and missing assets block final export.
- Project rule — acceptance evidence: Replace one voiceover chunk and update only its affected timeline interval; an interrupted render preserves all source assets and leaves no misleading final file.

## Optional agent skills and references
- Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate the actual Vyravid-inspired workflow with owned or clearly labeled sample data: Create a script, divide it into editable scenes, request or import authorized voice/image assets and assemble a simple timeline. Preview each scene and render a local video with explicit attribution and provider-cost records.
- Publish a reproducible walkthrough with this observable result: Replace one voiceover chunk and update only its affected timeline interval; an interrupted render preserves all source assets and leaves no misleading final file.
- Explain who can operate this scoped tool, its setup and ongoing costs, and these remaining product gaps: someone else's servers & support; provider-specific voice/image/video integration details Avoid guaranteed savings, performance scores or implied endorsement.

## Engineering roadmap
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Create a script, divide it into editable scenes, request or import authorized voice/image assets and assemble a simple timeline. Preview each scene and render a local video with explicit attribution and provider-cost records. Record prerequisites, select representative user-owned fixtures and document the unsupported features: someone else's servers & support; provider-specific voice/image/video integration details
2. Phase 2 — Durable model. Model scripts, scene IDs, voiceover chunks, asset provenance, timeline intervals and render manifests Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Generated assets are not silently regenerated on retry; scene duration follows actual audio and missing assets block final export.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals.
4. Phase 4 — Permissions and integration failure. Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Replace one voiceover chunk and update only its affected timeline interval; an interrupted render preserves all source assets and leaves no misleading final file. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

## Paid-product capabilities outside this build
- someone else's servers & support
- provider-specific voice/image/video integration details
- timeline rendering polish
- custom voice and character workflows

## Implementation prompt
WORKING SLICE
Create a script, divide it into editable scenes, request or import authorized voice/image assets and assemble a simple timeline. Preview each scene and render a local video with explicit attribution and provider-cost records.

Build this scoped Vyravid-inspired workflow with a documented data model and visible failure states.

Architecture
- Node.js, TypeScript and Express with server-rendered HTML and small browser modules.
- SQLite through better-sqlite3 with migrations, prepared statements and a single background worker.
- FFmpeg rendering, a configured TTS provider and optional image-generation adapter; immutable local asset manifests.

Prerequisites and limits
A supported Node release, a writable local data directory and a separate backup location. Bind to localhost; remote use requires authentication and HTTPS first. FFmpeg, permitted media and provider credentials only for explicitly enabled generation steps.
Outside this release: someone else's servers & support; provider-specific voice/image/video integration details

Data model and correctness
scripts, scene IDs, voiceover chunks, asset provenance, timeline intervals and render manifests
Invariant: Generated assets are not silently regenerated on retry; scene duration follows actual audio and missing assets block final export.
Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals.

Security and privacy
Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs.

Recovery and export
Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Create a script, divide it into editable scenes, request or import authorized voice/image assets and assemble a simple timeline. Preview each scene and render a local video with explicit attribution and provider-cost records. Record prerequisites, select representative user-owned fixtures and document the unsupported features: someone else's servers & support; provider-specific voice/image/video integration details
2. Phase 2 — Durable model. Model scripts, scene IDs, voiceover chunks, asset provenance, timeline intervals and render manifests Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Generated assets are not silently regenerated on retry; scene duration follows actual audio and missing assets block final export.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Save a job manifest with input hash, parameters and state. Write to temporary outputs, then atomically finalize only successful results; resume unfinished jobs without replacing originals.
4. Phase 4 — Permissions and integration failure. Bound file sizes and processing time, reject path traversal, and use argument arrays for subprocesses. Treat imported text as data and redact confidential source content from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Export sources, manifests and outputs with checksums. Keep failed-job diagnostics and allow retry into a new output path; restore the database and file directory together. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. Replace one voiceover chunk and update only its affected timeline interval; an interrupted render preserves all source assets and leaves no misleading final file. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
Replace one voiceover chunk and update only its affected timeline interval; an interrupted render preserves all source assets and leaves no misleading final file.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: scripts, scene IDs, voiceover chunks, asset provenance, timeline intervals and render manifests
Project rule — preserve this invariant: Generated assets are not silently regenerated on retry; scene duration follows actual audio and missing assets block final export.
Project rule — acceptance evidence: Replace one voiceover chunk and update only its affected timeline interval; an interrupted render preserves all source assets and leaves no misleading final file.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
