# AGENTS.md — Build guide for Better Stack / UptimeRobot paid

## Project scope
Probe allowed targets from a separate host, open incidents after defined failure thresholds and send down/recovery notifications with visible monitoring gaps.

Catalogue verdict: yes. A cron loop, a fetch, an alert webhook, and a status page. The $30/mo is for the dashboard gloss.
Use the implementation prompt below to define the deliverable. Complete each phase's acceptance checks before extending the scope.

## Working agreement
- Inspect the repository and its existing instructions before choosing paths, dependencies or commands. Keep one coherent stack and explain changes to the proposed architecture.
- Plan a vertical slice that accepts a real input and produces the useful output described below. Persist only the state the prompt calls for; respect memory-only and upstream-managed workflows. Use fixtures only when they are clearly labelled.
- After scaffolding, document the actual install, development, check and build commands in README and keep them synchronized with the package or project manifest. Do not report commands as successful unless they ran.
- Work in small steps. At handoff, list implemented flows, checks actually performed, remaining blockers, and any credentials or provider setup the owner must supply.
- Do not publish, spend money, contact customers, delete source data or run irreversible migrations without the project owner's authorization.

## Prerequisites
- Runtime and tools: TypeScript, Node, SQLite, a bounded scheduled worker and a read-only status dashboard.
- Before starting: An always-on host, explicit allowed targets, an alert destination and synthetic healthy/failing response fixtures.

## Stack and architecture
- TypeScript, Node, SQLite, a bounded scheduled worker and a read-only status dashboard
- Data design: Store Monitor, Probe, FailureStreak, Incident and AlertOutbox; restarting preserves the active incident and missing probes remain unknown time.
- Setup: An always-on host, explicit allowed targets, an alert destination and synthetic healthy/failing response fixtures

## Security and data integrity
- Persist samples, incident transitions and notification receipts. Missing samples are gaps rather than success; apply hysteresis and keep alert state across restarts. Run the checker separately from the monitored service.
- Keep the existing bounded Node/Express implementation and explicit host allowlist. Observed check success is not an SLA; single-region probes cannot prove global availability.
- Keep secrets outside client bundles and exported projects; document what leaves the device and make retention/deletion controls visible.

## Agent implementation rules
- Project rule — domain: Store Monitor, Probe, FailureStreak, Incident and AlertOutbox; restarting preserves the active incident and missing probes remain unknown time.
- Project rule — scope and recovery: Keep the existing bounded Node/Express implementation and explicit host allowlist. Observed check success is not an SLA; single-region probes cannot prove global availability.
- Project rule — acceptance: Simulate two failures, restart during the outage and recover; produce one incident lifecycle, with redirect-to-private-host probes rejected and scheduler gaps excluded from successful checks.
- Project rule — delivery: document real setup commands and permissions; do not claim a build, accuracy level, performance result or security certification that has not been demonstrated.

## Optional agent skills and references
- Recommended skill: [web-design-guidelines](https://github.com/vercel-labs/agent-skills/blob/main/skills/web-design-guidelines/SKILL.md) — review keyboard access, focus, validation, error recovery and the readable work/review interface or HTML report. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.
- Recommended skill: [sharp-edges](https://github.com/trailofbits/skills/blob/main/plugins/sharp-edges/skills/sharp-edges/SKILL.md) — review configuration and API defaults against the app-specific invariants and recovery boundaries above; this is not a security certification. Follow the maintainer's installation instructions and match its requirements to the chosen runtime.

Read the linked SKILL.md and its dependencies before adding a skill. Select only the skills matching this project's runtime and task; their documentation does not supply API access, credentials or approval to perform external actions. Pin the reviewed revision where the tool supports it. Follow the chosen agent's documented project-level installation mechanism.

## Distribution ideas
These are optional planning notes. Obtain the owner's approval before publishing or contacting anyone.
- Demonstrate this working slice using synthetic or explicitly authorized non-sensitive examples: Probe allowed targets from a separate host, open incidents after defined failure thresholds and send down/recovery notifications with visible monitoring gaps.
- Share a synthetic example export and the acceptance walkthrough; keep real customer, health, financial and source data private: Simulate two failures, restart during the outage and recover; produce one incident lifecycle, with redirect-to-private-host probes rejected and scheduler gaps excluded from successful checks.
- State the limits before asking someone to replace their existing tool: Keep the existing bounded Node/Express implementation and explicit host allowlist. Observed check success is not an SLA; single-region probes cannot prove global availability.

## Engineering roadmap
1. Phase 1 — Pin the working slice and create its example input: Probe allowed targets from a separate host, open incidents after defined failure thresholds and send down/recovery notifications with visible monitoring gaps. Confirm setup: An always-on host, explicit allowed targets, an alert destination and synthetic healthy/failing response fixtures.
2. Phase 2 — Implement persistence and write-time invariants before decorating the UI: Store Monitor, Probe, FailureStreak, Incident and AlertOutbox; restarting preserves the active incident and missing probes remain unknown time.
3. Phase 3 — Connect the working view to real saved state. Persist samples, incident transitions and notification receipts. Missing samples are gaps rather than success; apply hysteresis and keep alert state across restarts. Run the checker separately from the monitored service.
4. Phase 4 — Expose the app-specific limits and recovery path in context: Keep the existing bounded Node/Express implementation and explicit host allowlist. Observed check success is not an SLA; single-region probes cannot prove global availability.
5. Phase 5 — Walk through this concrete acceptance case and preserve its exported evidence: Simulate two failures, restart during the outage and recover; produce one incident lifecycle, with redirect-to-private-host probes rejected and scheduler gaps excluded from successful checks. Finish the README and backup/restore instructions; report unfinished capabilities explicitly.

## Paid-product capabilities outside this build
- global multi-region probes
- on-call scheduling & escalation policies
- incident timelines and postmortem tooling
- phone-call alerts

## Implementation prompt
Build me a small uptime monitor with incident notifications like UptimeRobot. Requirements:

- Use Node.js, Express, better-sqlite3, and native fetch. Read monitors.json entries containing name, HTTPS URL, interval, timeout, expected status, and optional body keyword. Store probe results, incident transitions, and an alert outbox in ./data/uptime.db.
- Run one scheduler with bounded concurrency and no overlapping checks for the same monitor. Use AbortController timeouts and cap response bytes when checking a body keyword. Record timeout, DNS, TLS, status mismatch, and keyword mismatch as distinct failure reasons.
- Require an explicit allowed-host list for configured targets. Validate DNS results and every redirect hop against public addresses and the allowed scope; block loopback, private, link-local, and metadata endpoints. Do not let public requests add arbitrary probe targets.
- A monitor starts unknown. Two consecutive failed probes open one incident; a successful probe closes it and clears the failure streak. Persist state so restarting does not create a second incident. A scheduler gap is unknown time, not proof the target stayed up.
- Write incident changes and notification outbox rows together. Send Slack-compatible webhook notifications with a URL from .env; retry transport failures and show exhausted attempts. Include failure category and observed outage duration without secret query strings.
- Serve a public status page only for monitors marked public. Show current state, last observed time, latency sparkline, and observed successful-check percentage for 24h/7d/30d, with successful/total check counts. Label missing monitoring periods and do not call the percentage an SLA.
- Keep 90 days of raw probes and prune in bounded batches. Add protected local configuration validation, JSON/CSV export, no telemetry, and a systemd unit. Run on a host separate from the targets it monitors.
- Acceptance: scripted status/timeout fixtures trigger one down and one recovery alert; a restart retains the open incident; an internal-address redirect is blocked; a monitoring gap stays unknown. README covers configuration, HTTPS, webhook setup, backups, and single-region limitations. Out of scope: global probes, on-call routing, phone alerts.

EDITORIAL IMPLEMENTATION CONTRACT
Working slice: Probe allowed targets from a separate host, open incidents after defined failure thresholds and send down/recovery notifications with visible monitoring gaps.
Data and invariants: Store Monitor, Probe, FailureStreak, Incident and AlertOutbox; restarting preserves the active incident and missing probes remain unknown time.
Boundary and recovery: Keep the existing bounded Node/Express implementation and explicit host allowlist. Observed check success is not an SLA; single-region probes cannot prove global availability.
Acceptance walkthrough: Simulate two failures, restart during the outage and recover; produce one incident lifecycle, with redirect-to-private-host probes rejected and scheduler gaps excluded from successful checks.
Record actual dependency versions, permissions and provider access in setup instructions. Preserve originals, expose partial failures and document backup/restore. These are acceptance requirements, not a claim of a completed or production-certified build. Add the domain, recovery and acceptance rules to AGENTS.md so future edits preserve them.

## Completion evidence
Demonstrate the prompt's acceptance scenarios against the scoped workflow. Include setup from a clean checkout and failure recovery. Check persistence across restart and export/restore only for the state the prompt says to store; for memory-only tools, confirm that temporary content is discarded as specified. Record actual results and remaining limitations. A detailed plan alone does not establish a working replacement.
