# AGENTS.md — Build guide for ArtificialWatch

## Project scope
Poll the model catalogues accessible to the user's keys, seed a silent baseline and alert only when a new identifier survives the configured confirmation window.

Catalogue verdict: kinda. The core loop is a genuine one-sitting build: poll the model-list endpoint of every provider you hold a key for, diff the ids against a local table, push the new ones to your phone. What does not one-shot is the three things a launch alert is actually judged on. Coverage · your script sees only the labs you have accounts with, while the watchlist here runs to 48 models including Chinese labs, restricted previews and things that have not shipped at all, which no API returns. Telephony · a call that rings until you answer means Twilio, a purchased number and US A2P 10DLC registration, which is days of paperwork before a line of code. And uptime, which is the whole product · a poller on a laptop that slept through the drop is worth nothing, and the second-sweep debounce that keeps preview aliases from crying wolf is the part you only tune after it has already cried wolf twice.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- Runtime and tools: TypeScript, Node, SQLite, a bounded scheduled worker and a read-only status dashboard.
- Before starting: An always-on host, explicit allowed targets, an alert destination and synthetic healthy/failing response fixtures.

## Stack and architecture
- TypeScript, Node, SQLite, a bounded scheduled worker and a read-only status dashboard
- Data design: Store Provider, PollResult, ModelObservation, Candidate and AlertReceipt; distinguish a failed poll from model removal and keep aliases/snapshots separate from a newly observed family.
- Setup: An always-on host, explicit allowed targets, an alert destination and synthetic healthy/failing response fixtures

## Security and data integrity
- Persist samples, incident transitions and notification receipts. Missing samples are gaps rather than success; apply hysteresis and keep alert state across restarts. Run the checker separately from the monitored service.
- Model-list visibility is account-dependent and does not establish launch availability. Keep provider-specific rate limits, alias filters and startup gaps visible; notification delivery is not guaranteed.
- Keep secrets outside client bundles and exported projects; document what leaves the device and make retention/deletion controls visible.

## Agent implementation rules
- Project rule — domain: Store Provider, PollResult, ModelObservation, Candidate and AlertReceipt; distinguish a failed poll from model removal and keep aliases/snapshots separate from a newly observed family.
- Project rule — scope and recovery: Model-list visibility is account-dependent and does not establish launch availability. Keep provider-specific rate limits, alias filters and startup gaps visible; notification delivery is not guaranteed.
- Project rule — acceptance: Seed twenty IDs, fail the next poll and then return one new ID twice; emit one alert for the confirmed candidate and none for the initial seed or outage.
- Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.

## Optional agent skills and references
- Recommended skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — review configuration and API defaults against the app-specific invariants and recovery boundaries above; this is not a security certification. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
- Recommended skill: [copywriting](https://github.com/coreyhaines31/marketingskills/blob/main/skills/copywriting/SKILL.md) — write clear, evidence-grounded draft copy or notifications without invented claims; sending/publishing remains separately approved. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate this working slice using synthetic or explicitly authorized non-sensitive examples: Poll the model catalogues accessible to the user's keys, seed a silent baseline and alert only when a new identifier survives the configured confirmation window.
- Share a synthetic example export and the acceptance walkthrough; keep real customer, health, financial and source data private: Seed twenty IDs, fail the next poll and then return one new ID twice; emit one alert for the confirmed candidate and none for the initial seed or outage.
- State the limits before asking someone to replace their existing tool: Model-list visibility is account-dependent and does not establish launch availability. Keep provider-specific rate limits, alias filters and startup gaps visible; notification delivery is not guaranteed.

## Engineering roadmap
1. Phase 1 — Pin the working slice and create its example input: Poll the model catalogues accessible to the user's keys, seed a silent baseline and alert only when a new identifier survives the configured confirmation window. Confirm setup: An always-on host, explicit allowed targets, an alert destination and synthetic healthy/failing response fixtures.
2. Phase 2 — Implement persistence and write-time invariants before decorating the UI: Store Provider, PollResult, ModelObservation, Candidate and AlertReceipt; distinguish a failed poll from model removal and keep aliases/snapshots separate from a newly observed family.
3. Phase 3 — Connect the working view to real saved state. Persist samples, incident transitions and notification receipts. Missing samples are gaps rather than success; apply hysteresis and keep alert state across restarts. Run the checker separately from the monitored service.
4. Phase 4 — Expose the app-specific limits and recovery path in context: Model-list visibility is account-dependent and does not establish launch availability. Keep provider-specific rate limits, alias filters and startup gaps visible; notification delivery is not guaranteed.
5. Phase 5 — Walk through this concrete acceptance case and preserve its exported evidence: Seed twenty IDs, fail the next poll and then return one new ID twice; emit one alert for the confirmed candidate and none for the initial seed or outage. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.

## Paid-product capabilities outside this build
- coverage of labs you hold no key for · 48 tracked models including Chinese labs and restricted previews
- the phone call that rings until you answer, and SMS · Twilio plus US A2P 10DLC registration
- the pre-launch watchlist and live Polymarket odds on models that have not shipped, which no API can return
- debounce and alias filtering tuned so dated snapshots and -preview ids do not fire false alarms
- someone else owning the uptime · your poller sleeps when your machine does

## Implementation prompt
Build me a new-AI-model launch alerter to replace ArtificialWatch. Requirements:

- Node 22 + node-cron + better-sqlite3, one process on a small VPS under pm2 with a supervised process; a sleeping or unavailable host misses observations.
- Every 60 seconds, GET the model-list endpoints for the keys in .env: OpenAI
  /v1/models, Anthropic /v1/models, Google generativelanguage /v1beta/models,
  and OpenRouter /api/v1/models, which covers labs I have no account with.
- Store every model id ever seen in SQLite. A candidate is an identifier newly observed in that table, not proof of a public launch
  · seed it on first run so the first sweep is silent.
- Debounce: an id fires only after two consecutive sweeps, and ids matching a
  regex list in config.json (dated snapshots, -preview, -latest) never fire.
- Alert by POSTing to an ntfy.sh topic: model id as the title, provider plus
  context window and per-token price as the body, link to the provider's docs.
- Append each confirmed launch to launches.md as `YYYY-MM-DD · provider · id`.
- One page on localhost:8080: last 50 launches, last good sweep per provider,
  red banner when a provider has errored 10 minutes · a silent poller is worse
  than none.
- Out of scope: SMS and phone calls (Twilio plus US A2P 10DLC registration is
  paperwork, not code) and any watchlist of unshipped models.
- README with the four .env keys and the pm2 command.

EDITORIAL IMPLEMENTATION CONTRACT
Working slice: Poll the model catalogues accessible to the user's keys, seed a silent baseline and alert only when a new identifier survives the configured confirmation window.
Data and invariants: Store Provider, PollResult, ModelObservation, Candidate and AlertReceipt; distinguish a failed poll from model removal and keep aliases/snapshots separate from a newly observed family.
Boundary and recovery: Model-list visibility is account-dependent and does not establish launch availability. Keep provider-specific rate limits, alias filters and startup gaps visible; notification delivery is not guaranteed.
Acceptance walkthrough: Seed twenty IDs, fail the next poll and then return one new ID twice; emit one alert for the confirmed candidate and none for the initial seed or outage.
Record actual dependency versions, permissions and provider access in setup instructions. Preserve originals, expose partial failures and document backup/restore. These are acceptance requirements, not a claim of a completed or production-certified build. Add the domain, recovery and acceptance rules to AGENTS.md so future edits preserve them.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
