Pangram

A classifier that scores whether a piece of text was written by an AI model.

NOT REALLY · consider alternatives
price $20/mosubscription / year $240estimated build time one sittingreplaced by 0 people

The interface is a text box and a percentage, which is exactly the kind of thing that tricks you into thinking it is a weekend project. The product is not the text box: it is a classifier trained on a very large, continuously refreshed corpus of human writing paired with output from every model release, tuned hard against false positives because accusing a real person of cheating is the failure mode that ends the company. You can absolutely build a local detector from perplexity and burstiness features in an afternoon, and it will be confidently wrong often enough to be useless for any decision that matters. Nobody outside your own head will accept your homemade score, and the calibration drifts every time a new frontier model ships. Build it to understand the problem, not to rely on it.

Build verification: not recorded. How we judge buildability

What you give up

  • Calibration: a real false positive rate you can quote, instead of a vibe
  • Coverage of new models, which changes every few weeks whether you update or not
  • Sentence-level and mixed-authorship detection rather than one blunt document score
  • Any external credibility, since a self-built score persuades exactly zero teachers, editors or clients
  • Throughput, batch uploads, API access and document parsing

Why people still pay

Because the number has to be defensible to someone else. Institutions and publishers are not buying a classifier, they are buying a third party willing to stand behind a false positive rate, plus continuous retraining against whatever model came out last month. A local heuristic detector gives you a plausible-sounding percentage with no error bars, which is worse than nothing when the outcome is an accusation.

Your build guide

The stack, security requirements, and agent rules for a focused replacement.

Before you start

  • Local Python with torch and transformers
  • A couple of GB of disk for small model weights, CPU works but is slow
  • Your own labelled samples of human and AI text if you want any idea of accuracy
01
Stack, no substitutions: Python 3.11, FastAPI + uvicorn, transformers + torch, a single server-rendered HTML page with vanilla JS. No build step, no database, no Docker.
02
Model source documents, editable pages or notes, internal links, revisions, and exports; retain source IDs and timestamps.
03
Interface for this AI text detection workflow: a local web page with input, progress, review, and export views.
engineering roadmap

Implementation plan

1

Phase 1, architecture and data

Stack, no substitutions: Python 3.11, FastAPI + uvicorn, transformers + torch, a single server-rendered HTML page with vanilla JS. No build step, no database, no Docker. Model source documents, editable pages or notes, internal links, revisions, and exports; retain source IDs and timestamps.

2

Phase 2, implement

Features to compute per submission:.

3

Phase 3, implement

Burstiness: standard deviation of per-sentence mean log-probability.

4

Phase 4, review and output

Rank-based signal: fraction of tokens that were in the model's top-10 predictions. Scoring runs entirely locally with gpt2 (small) loaded once at startup via transformers. Download on first run and cache it.

5

Phase 5, recovery and acceptance

On invalid input or interrupted processing, keep the original record, show the failed step, and permit a safe retry. Verify this invariant with a saved fixture: A classifier score cannot be presented as proof of authorship; a failed detector run remains unknown and never overwrites source text. State the practical limit: Calibration: a real false positive rate you can quote, instead of a vibe.

the pro prompt
Build a local AI-text-detection playground so I can see how weak naive detection actually is. No cloud calls, no accounts, no telemetry.

Stack, no substitutions: Python 3.11, FastAPI + uvicorn, transformers + torch, a single server-rendered HTML page with vanilla JS. No build step, no database, no Docker.

Core loop:
1. One page with a large textarea and an Analyze button.
2. POST /analyze takes the text and returns a JSON report.
3. Scoring runs entirely locally with gpt2 (small) loaded once at startup via transformers. Download on first run and cache it.

Features to compute per submission:
- Mean token log-probability and perplexity under gpt2.
- Burstiness: standard deviation of per-sentence mean log-probability.
- Rank-based signal: fraction of tokens that were in the model's top-10 predictions.
- Style stats: sentence length mean and variance, type-token ratio, punctuation counts, count of common LLM filler phrases from a hardcoded list.

Output: a 0-100 score from a simple weighted formula over those features with the weights in a single WEIGHTS dict at the top of the file, per-sentence heat colouring in the page, and a full feature table so I can see what drove the score.

Calibration, and be honest about it:
- Add a CLI command: python calibrate.py --human dir --ai dir that runs the features over two folders of .txt files, prints mean and spread per class, and prints the ROC AUC of the current weights.
- The web page must show a fixed banner: this is an uncalibrated heuristic, not evidence, false positives are common.

Explicitly out of scope: user accounts, saved history, PDF or DOCX parsing, any hosted API, model fine-tuning, mixed-authorship segmentation, batch uploads.

Deliver: main.py, calibrate.py, features.py, templates/index.html, requirements.txt, and a README with run instructions and one paragraph stating plainly why this cannot match a trained commercial detector. No .env needed since there are no secrets; say so in the README.

$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy AGENTS.md · generated from this app's build plan

prior art · use these instead of building, if you'd rather

No prior-art project is listed yet. Compare the scoped build with the paid product before choosing.

share on X ↗

Questions about Pangram

Can you build your own Pangram with AI?

A full replacement is not the recommended project. The interface is a text box and a percentage, which is exactly the kind of thing that tricks you into thinking it is a weekend project. The product is not the text box: it is a classifier trained on a very large, continuously refreshed corpus of human writing paired with output from every model release, tuned hard against false positives because accusing a real person of cheating is the failure mode that ends the company. You can absolutely build a local detector from perplexity and burstiness features in an afternoon, and it will be confidently wrong often enough to be useless for any decision that matters. Nobody outside your own head will accept your homemade score, and the calibration drifts every time a new frontier model ships. Build it to understand the problem, not to rely on it.

What does the Pangram build prompt cover?

The prompt starts with this scope: Paste text, score it locally with a small language model's token log-probabilities plus a few style statistics, and get a hand-wavy human-or-machine guess with a loud accuracy disclaimer. Full-product capabilities excluded from the comparison include: Calibration: a real false positive rate you can quote, instead of a vibe; Coverage of new models, which changes every few weeks whether you update or not; Sentence-level and mixed-authorship detection rather than one blunt document score. Follow the implementation plan and its prerequisites before expanding the build.

How do I use the prompt, AGENTS.md and agent skills?

Start with the Pangram prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.

How long will this Pangram project take?

The catalogue estimate is one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.

What would I give up by replacing Pangram?

Calibration: a real false positive rate you can quote, instead of a vibe; Coverage of new models, which changes every few weeks whether you update or not; Sentence-level and mixed-authorship detection rather than one blunt document score; Any external credibility, since a self-built score persuades exactly zero teachers, editors or clients; Throughput, batch uploads, API access and document parsing. Because the number has to be defensible to someone else. Institutions and publishers are not buying a classifier, they are buying a third party willing to stand behind a false positive rate, plus continuous retraining against whatever model came out last month. A local heuristic detector gives you a plausible-sounding percentage with no error bars, which is worse than nothing when the outcome is an accusation.

What price is this guide comparing against?

The recorded Individual plan is $20/mo (monthly, 300,000 words/month), checked 2026-08-18. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.

What can I use instead of building Pangram?

No alternative is listed in this entry yet. That is a gap in this catalogue, not proof that no suitable product exists. Compare the paid product and the proposed scope before committing to a build.

Every week, more subscriptions die.

New verdicts, new prompts, the week's most-doomed apps.
One email. Unsubscribe in one click.

last week:100 Questions · KINDA1of10 · KINDA1Password · KINDA+1090 more

free forever · no scanner spam · the prompt stays on the site, the deaths come to you

$weekly: what got a verdict, what died.