Idea
A public scanner on the web app: enter a company's domain and we try to find the changelog(s) associated with it — the interactive version of the find-a-changelog guide (#2252), and a natural upgrade to /submit.
Flow sketch:
- Index first, zero fetches. Resolve the domain against what we already track (
/v1/lookups/by-domain, org catalog). If it's indexed: show the org page, sources, and .atom links. Most lookups should end here for free.
- Heuristic scan (bounded, cacheable) for unknown domains — this is mostly machinery we already have in
packages/adapters / the discovery pre-checks:
/.well-known/releases.json + changelog.json, <link rel="changelog">
- guessable URLs (
/changelog, /updates, /releases, changelog.{domain}, …)
- feed autodiscovery + provider fingerprinting (Mintlify/ReadMe/Fern/etc. known paths)
- GitHub org guess from homepage links
- Cheap AI pass only when heuristics are ambiguous: a Haiku-class verdict on "is this page actually a changelog?" (same shape as the marketing classifier / URL evaluation we already run — reuse an existing model lane, no per-feature model var).
- Outcome ties into submission: found → pre-filled submit/track request (the
tracking_requested_at demand signal from the listing lane); already indexed → org page; nothing found → plain "couldn't find one" plus the manifest pitch for owners.
Doing it responsibly
This is the actual design problem — an open "scan any domain" endpoint is a free fetch proxy and an AI-spend faucet if built naively:
- Cache by domain (KV/D1, generous TTL). A domain gets scanned once, not per visitor.
- Rate limits per-IP and per-domain, CF-native, mirroring
POST /v1/listing/{validate,activate} (10/min IP, per-domain cap).
- Bounded fetch budget per scan: a fixed small number of requests, target domain only (plus at most one GitHub API call) — never user-supplied arbitrary URLs, so it can't be used as a proxy. SSRF guards as in webhook-url-safety.
- Opt-out preflight fail-closed: robots.txt / Content-Signal check before fetching content pages, same posture as
local-ingest/preflight.ts.
- AI spend: only on the ambiguous tail, Haiku-class, behind the existing spend-cap gate; heuristics-only is an acceptable degraded mode.
- Kill switch flag — this one genuinely earns a flag (public, external-fetching, AI-spending).
- Possibly require a signed-in account for the live-scan tier while keeping index lookups anonymous.
Why bother
- Converts guide/SEO traffic into submissions and follows: the guide teaches the method, the scanner runs it for you on the spot.
- The
tracking_requested_at signal from scans of unindexed domains is curator-facing demand data we don't get today.
- Nearly all the discovery logic exists (evaluate/discovery pre-checks, well-known parsing, provider table); the new work is the public surface, the caching/limits, and the responsible-fetch envelope.
Not scoped yet — filing so it doesn't get lost. Related: #1947 (self-serve listing), the guides PR #2252.
Idea
A public scanner on the web app: enter a company's domain and we try to find the changelog(s) associated with it — the interactive version of the find-a-changelog guide (#2252), and a natural upgrade to
/submit.Flow sketch:
/v1/lookups/by-domain, org catalog). If it's indexed: show the org page, sources, and.atomlinks. Most lookups should end here for free.packages/adapters/ the discovery pre-checks:/.well-known/releases.json+changelog.json,<link rel="changelog">/changelog,/updates,/releases,changelog.{domain}, …)tracking_requested_atdemand signal from the listing lane); already indexed → org page; nothing found → plain "couldn't find one" plus the manifest pitch for owners.Doing it responsibly
This is the actual design problem — an open "scan any domain" endpoint is a free fetch proxy and an AI-spend faucet if built naively:
POST /v1/listing/{validate,activate}(10/min IP, per-domain cap).local-ingest/preflight.ts.Why bother
tracking_requested_atsignal from scans of unindexed domains is curator-facing demand data we don't get today.Not scoped yet — filing so it doesn't get lost. Related: #1947 (self-serve listing), the guides PR #2252.