# AGENTS.md — Build guide for TranscriptAPI

## Project scope
Expose a small FastAPI endpoint that accepts a YouTube URL and retrieves an available caption track through a supported library. Return timestamped JSON, language and whether captions are manual or generated.

Catalogue verdict: kinda. The happy path is genuinely a one-sitting build: the open-source youtube-transcript-api library fetches a caption track in a few lines, and wrapping it in FastAPI gives you a working personal endpoint. The honest gap is reliability. YouTube rate-limits and IP-blocks caption scraping at any real volume, so the DIY version works until it suddenly does not, and there is no fix without a rotating proxy pool you must rent and operate. The paid product is not selling the parsing; it is selling the unblocked pipe, plus search and playlist endpoints on top.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- Python, permission to retrieve the requested public caption tracks and access compatible with the selected library. Start private and low-volume; no rotation to bypass provider restrictions.
- Implementation components: Python, FastAPI, uvicorn and the documented youtube-transcript-api package interface. A small private HTTP endpoint with a bounded optional cache and typed caption/language/error responses.
- Scope boundary: Platform reliability, all-video coverage and speech transcription of inaccessible media are not guaranteed.

## Stack and architecture
- Python, FastAPI, uvicorn and the documented youtube-transcript-api package interface.
- A small private HTTP endpoint with a bounded optional cache and typed caption/language/error responses.
- Domain model: requested video IDs, permitted caption tracks, language selections, fetch status and timestamped transcript caches

## Security and data integrity
- Accept only a validated video ID or supported YouTube URL shape, never an arbitrary fetch URL. Bound request rate, upstream timeout and response size; bind uvicorn to 127.0.0.1 for local use. Remote deployment requires authenticated HTTPS, a server-side token check and per-token limits; a private label is not authorization. Do not expose cookies or access credentials in errors.
- Correctness boundary: No caption access means an explicit unavailable response; do not bypass private, age or geographic access restrictions or fabricate a transcript.
- Validate the video ID before fetching, bound request time and distinguish missing captions from upstream failures. Cache only permitted successful results with fetch time and language.
- Return portable timestamped JSON and preserve a clear error status. Cached captions are disposable; no local transcript is invented when the upstream response is unavailable.

## Agent implementation rules
- Project rule — data model: requested video IDs, permitted caption tracks, language selections, fetch status and timestamped transcript caches
- Project rule — preserve this invariant: No caption access means an explicit unavailable response; do not bypass private, age or geographic access restrictions or fabricate a transcript.
- Project rule — acceptance evidence: A malformed URL is rejected before network work; a video without captions returns a typed unavailable result and repeated allowed requests use a dated cache.

## Optional agent skills and references
- Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate the actual TranscriptAPI-inspired workflow with owned or clearly labeled sample data: Expose a small FastAPI endpoint that accepts a YouTube URL and retrieves an available caption track through a supported library. Return timestamped JSON, language and whether captions are manual or generated.
- Publish a reproducible walkthrough with this observable result: A malformed URL is rejected before network work; a video without captions returns a typed unavailable result and repeated allowed requests use a dated cache.
- Explain who can operate this scoped tool, its setup and ongoing costs, and these remaining product gaps: Platform reliability, all-video coverage and speech transcription of inaccessible media are not guaranteed. Avoid guaranteed savings, performance scores or implied endorsement.

## Engineering roadmap
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Expose a small FastAPI endpoint that accepts a YouTube URL and retrieves an available caption track through a supported library. Return timestamped JSON, language and whether captions are manual or generated. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Platform reliability, all-video coverage and speech transcription of inaccessible media are not guaranteed.
2. Phase 2 — Durable model. Model requested video IDs, permitted caption tracks, language selections, fetch status and timestamped transcript caches Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: No caption access means an explicit unavailable response; do not bypass private, age or geographic access restrictions or fabricate a transcript.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Validate the video ID before fetching, bound request time and distinguish missing captions from upstream failures. Cache only permitted successful results with fetch time and language.
4. Phase 4 — Permissions and integration failure. Accept only a validated video ID or supported YouTube URL shape, never an arbitrary fetch URL. Bound request rate, upstream timeout and response size; bind uvicorn to 127.0.0.1 for local use. Remote deployment requires authenticated HTTPS, a server-side token check and per-token limits; a private label is not authorization. Do not expose cookies or access credentials in errors. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Return portable timestamped JSON and preserve a clear error status. Cached captions are disposable; no local transcript is invented when the upstream response is unavailable. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. A malformed URL is rejected before network work; a video without captions returns a typed unavailable result and repeated allowed requests use a dated cache. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

## Paid-product capabilities outside this build
- staying unblocked when YouTube rate-limits caption fetches
- search, channel and playlist endpoints
- reliability at bulk volume
- an SLA and support when YouTube changes something

## Implementation prompt
WORKING SLICE
Expose a small FastAPI endpoint that accepts a YouTube URL and retrieves an available caption track through a supported library. Return timestamped JSON, language and whether captions are manual or generated.

Build this scoped TranscriptAPI-inspired workflow with a documented data model and visible failure states.

Architecture
- Python, FastAPI, uvicorn and the documented youtube-transcript-api package interface.
- A small private HTTP endpoint with a bounded optional cache and typed caption/language/error responses.

Prerequisites and limits
Python, permission to retrieve the requested public caption tracks and access compatible with the selected library. Start private and low-volume; no rotation to bypass provider restrictions.
Outside this release: Platform reliability, all-video coverage and speech transcription of inaccessible media are not guaranteed.

Data model and correctness
requested video IDs, permitted caption tracks, language selections, fetch status and timestamped transcript caches
Invariant: No caption access means an explicit unavailable response; do not bypass private, age or geographic access restrictions or fabricate a transcript.
Validate the video ID before fetching, bound request time and distinguish missing captions from upstream failures. Cache only permitted successful results with fetch time and language.

Security and privacy
Accept only a validated video ID or supported YouTube URL shape, never an arbitrary fetch URL. Bound request rate, upstream timeout and response size; bind uvicorn to 127.0.0.1 for local use. Remote deployment requires authenticated HTTPS, a server-side token check and per-token limits; a private label is not authorization. Do not expose cookies or access credentials in errors.

Recovery and export
Return portable timestamped JSON and preserve a clear error status. Cached captions are disposable; no local transcript is invented when the upstream response is unavailable.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Expose a small FastAPI endpoint that accepts a YouTube URL and retrieves an available caption track through a supported library. Return timestamped JSON, language and whether captions are manual or generated. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Platform reliability, all-video coverage and speech transcription of inaccessible media are not guaranteed.
2. Phase 2 — Durable model. Model requested video IDs, permitted caption tracks, language selections, fetch status and timestamped transcript caches Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: No caption access means an explicit unavailable response; do not bypass private, age or geographic access restrictions or fabricate a transcript.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Validate the video ID before fetching, bound request time and distinguish missing captions from upstream failures. Cache only permitted successful results with fetch time and language.
4. Phase 4 — Permissions and integration failure. Accept only a validated video ID or supported YouTube URL shape, never an arbitrary fetch URL. Bound request rate, upstream timeout and response size; bind uvicorn to 127.0.0.1 for local use. Remote deployment requires authenticated HTTPS, a server-side token check and per-token limits; a private label is not authorization. Do not expose cookies or access credentials in errors. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Return portable timestamped JSON and preserve a clear error status. Cached captions are disposable; no local transcript is invented when the upstream response is unavailable. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. A malformed URL is rejected before network work; a video without captions returns a typed unavailable result and repeated allowed requests use a dated cache. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
A malformed URL is rejected before network work; a video without captions returns a typed unavailable result and repeated allowed requests use a dated cache.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [modern-python](https://github.com/trailofbits/skills/blob/main/plugins/modern-python/skills/modern-python/SKILL.md) — Set up Python projects with pyproject.toml, dependency management, linting, typing and automated checks. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: requested video IDs, permitted caption tracks, language selections, fetch status and timestamped transcript caches
Project rule — preserve this invariant: No caption access means an explicit unavailable response; do not bypass private, age or geographic access restrictions or fabricate a transcript.
Project rule — acceptance evidence: A malformed URL is rejected before network work; a video without captions returns a typed unavailable result and repeated allowed requests use a dated cache.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
