# AGENTS.md — Build guide for Saply

## Project scope
Import a text-based PDF or DOCX, review extracted candidate fields and render an approved Word template using docxtemplater/PizZip. Tailoring may reorder or rephrase verified facts but cannot add experience.

Catalogue verdict: kinda. A capable coding agent can build the personal core: extract a clean, text-based CV, structure or tailor it with an LLM, and render a branded DOCX. It will work on tidy CVs and silently drop data on the rest. Saply's value is extraction accuracy across thousands of real-world CV layouts, tuned so no field goes missing on a client-facing document, plus the integrations that keep recruiters inside Word, their email, and their ATS.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- A supported Node release, a writable local data directory and a separate backup location. Bind to localhost; remote use requires authentication and HTTPS first. Approved DOCX templates, text-based resume fixtures and an optional model key with permission to send the chosen resume text.
- Implementation components: Node.js, TypeScript and Express with server-rendered HTML and small browser modules. SQLite through better-sqlite3 with migrations, prepared statements and a single background worker. A PDF/DOCX text extractor, schema-constrained optional model adapter, docxtemplater and PizZip for reviewed template rendering.
- Scope boundary: Scanned resumes require an explicit OCR step; hiring suitability and perfect layout extraction are not promised.

## Stack and architecture
- Node.js, TypeScript and Express with server-rendered HTML and small browser modules.
- SQLite through better-sqlite3 with migrations, prepared statements and a single background worker.
- A PDF/DOCX text extractor, schema-constrained optional model adapter, docxtemplater and PizZip for reviewed template rendering.
- Domain model: candidate source files, extracted sections with provenance, structured resume JSON, template fields and DOCX jobs

## Security and data integrity
- Reject unexpected origins and unbounded request bodies even on localhost. Keep credentials outside the database export and redact sensitive text from logs.
- Correctness boundary: Uncertain extraction is shown before rendering; template expressions are constrained and output cannot silently omit unmapped source sections.
- Use short SQLite transactions and persist job state before starting work. Give retries stable operation IDs; report incomplete or unknown results instead of silently repeating them.
- Use a consistent SQLite backup and an attachment manifest. Export portable JSON/CSV, then restore to a new directory without overwriting the original data.

## Agent implementation rules
- Project rule — data model: candidate source files, extracted sections with provenance, structured resume JSON, template fields and DOCX jobs
- Project rule — preserve this invariant: Uncertain extraction is shown before rendering; template expressions are constrained and output cannot silently omit unmapped source sections.
- Project rule — acceptance evidence: A two-column CV produces an extraction review with missing fields flagged; the exported DOCX retains template headers and every approved experience row.

## Optional agent skills and references
- Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [docx](https://github.com/anthropics/skills/blob/main/skills/docx/SKILL.md) — Generate or edit Word documents, document XML, tracked changes and comments. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
- Optional external skill: [pdf](https://github.com/anthropics/skills/blob/main/skills/pdf/SKILL.md) — Process PDFs through extraction, generation, page operations, form filling and OCR workflows. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate the actual Saply-inspired workflow with owned or clearly labeled sample data: Import a text-based PDF or DOCX, review extracted candidate fields and render an approved Word template using docxtemplater/PizZip. Tailoring may reorder or rephrase verified facts but cannot add experience.
- Publish a reproducible walkthrough with this observable result: A two-column CV produces an extraction review with missing fields flagged; the exported DOCX retains template headers and every approved experience row.
- Explain who can operate this scoped tool, its setup and ongoing costs, and these remaining product gaps: Scanned resumes require an explicit OCR step; hiring suitability and perfect layout extraction are not promised. Avoid guaranteed savings, performance scores or implied endorsement.

## Engineering roadmap
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Import a text-based PDF or DOCX, review extracted candidate fields and render an approved Word template using docxtemplater/PizZip. Tailoring may reorder or rephrase verified facts but cannot add experience. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Scanned resumes require an explicit OCR step; hiring suitability and perfect layout extraction are not promised.
2. Phase 2 — Durable model. Model candidate source files, extracted sections with provenance, structured resume JSON, template fields and DOCX jobs Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Uncertain extraction is shown before rendering; template expressions are constrained and output cannot silently omit unmapped source sections.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Use short SQLite transactions and persist job state before starting work. Give retries stable operation IDs; report incomplete or unknown results instead of silently repeating them.
4. Phase 4 — Permissions and integration failure. Reject unexpected origins and unbounded request bodies even on localhost. Keep credentials outside the database export and redact sensitive text from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Use a consistent SQLite backup and an attachment manifest. Export portable JSON/CSV, then restore to a new directory without overwriting the original data. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. A two-column CV produces an extraction review with missing fields flagged; the exported DOCX retains template headers and every approved experience row. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

## Paid-product capabilities outside this build
- extraction accuracy across thousands of real-world CV layouts, tuned on years of data so nothing is silently dropped
- OCR for scanned and image-based CVs
- the AI agent that edits any CV in plain language directly inside Word and Google Docs
- Word, Google Docs, email, and ATS integrations (Bullhorn, Carerix, Spott, Loxo)
- EU tender templates (Europass, DIGIT-TM III, ITUSS21) plus anonymisation and translation
- ISO 27001 certified handling of security-sensitive candidate data
- bulk processing, team workflows, SLA, and enterprise support

## Implementation prompt
WORKING SLICE
Import a text-based PDF or DOCX, review extracted candidate fields and render an approved Word template using docxtemplater/PizZip. Tailoring may reorder or rephrase verified facts but cannot add experience.

Build this scoped Saply-inspired workflow with a documented data model and visible failure states.

Architecture
- Node.js, TypeScript and Express with server-rendered HTML and small browser modules.
- SQLite through better-sqlite3 with migrations, prepared statements and a single background worker.
- A PDF/DOCX text extractor, schema-constrained optional model adapter, docxtemplater and PizZip for reviewed template rendering.

Prerequisites and limits
A supported Node release, a writable local data directory and a separate backup location. Bind to localhost; remote use requires authentication and HTTPS first. Approved DOCX templates, text-based resume fixtures and an optional model key with permission to send the chosen resume text.
Outside this release: Scanned resumes require an explicit OCR step; hiring suitability and perfect layout extraction are not promised.

Data model and correctness
candidate source files, extracted sections with provenance, structured resume JSON, template fields and DOCX jobs
Invariant: Uncertain extraction is shown before rendering; template expressions are constrained and output cannot silently omit unmapped source sections.
Use short SQLite transactions and persist job state before starting work. Give retries stable operation IDs; report incomplete or unknown results instead of silently repeating them.

Security and privacy
Reject unexpected origins and unbounded request bodies even on localhost. Keep credentials outside the database export and redact sensitive text from logs.

Recovery and export
Use a consistent SQLite backup and an attachment manifest. Export portable JSON/CSV, then restore to a new directory without overwriting the original data.

Implementation order
1. Phase 1 — Scope and fixtures. Implement this bounded workflow: Import a text-based PDF or DOCX, review extracted candidate fields and render an approved Word template using docxtemplater/PizZip. Tailoring may reorder or rephrase verified facts but cannot add experience. Record prerequisites, select representative user-owned fixtures and document the unsupported features: Scanned resumes require an explicit OCR step; hiring suitability and perfect layout extraction are not promised.
2. Phase 2 — Durable model. Model candidate source files, extracted sections with provenance, structured resume JSON, template fields and DOCX jobs Add migrations or a versioned document format, explicit validation, stable IDs and a visible import-error report. Preserve this rule: Uncertain extraction is shown before rendering; template expressions are constrained and output cannot silently omit unmapped source sections.
3. Phase 3 — Complete the first useful path. Implement the workflow's input, review and output interface, with clear controls and explicit empty/error states. Use short SQLite transactions and persist job state before starting work. Give retries stable operation IDs; report incomplete or unknown results instead of silently repeating them.
4. Phase 4 — Permissions and integration failure. Reject unexpected origins and unbounded request bodies even on localhost. Keep credentials outside the database export and redact sensitive text from logs. Request integration credentials and permissions only for the enabled feature; show a disconnected state instead of mock results.
5. Phase 5 — Portable handoff. Use a consistent SQLite backup and an attachment manifest. Export portable JSON/CSV, then restore to a new directory without overwriting the original data. Include setup, operating limits, fixture walkthrough and shutdown/restart instructions in the README.
6. Phase 6 — Acceptance scenarios. A two-column CV produces an extraction review with missing fields flagged; the exported DOCX retains template headers and every approved experience row. Repeat the workflow after restart and with a denied permission or unavailable dependency; show recoverable failure rather than a success placeholder.

Acceptance
A two-column CV produces an extraction review with missing fields flagged; the exported DOCX retains template headers and every approved experience row.
Use real source data or clearly labeled fixtures. Explain unsupported input and provider failures; do not fabricate analytics, delivery receipts, accuracy claims or security guarantees.

Optional agent guidance
Optional external skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — Review web interfaces for accessibility, keyboard focus, forms, navigation and interaction quality. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — Review security-sensitive APIs and configuration for dangerous defaults and easy-to-misuse interfaces. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [docx](https://github.com/anthropics/skills/blob/main/skills/docx/SKILL.md) — Generate or edit Word documents, document XML, tracked changes and comments. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Optional external skill: [pdf](https://github.com/anthropics/skills/blob/main/skills/pdf/SKILL.md) — Process PDFs through extraction, generation, page operations, form filling and OCR workflows. Review its instructions and compatibility before use; it does not grant deployment, data-access or publication permission.
Project rule — data model: candidate source files, extracted sections with provenance, structured resume JSON, template fields and DOCX jobs
Project rule — preserve this invariant: Uncertain extraction is shown before rendering; template expressions are constrained and output cannot silently omit unmapped source sections.
Project rule — acceptance evidence: A two-column CV produces an extraction review with missing fields flagged; the exported DOCX retains template headers and every approved experience row.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
