Skip to content

Public changelog scanner: enter a domain, find its changelogs (heuristics + cheap AI) #2253

Description

@zachdunn

Idea

A public scanner on the web app: enter a company's domain and we try to find the changelog(s) associated with it — the interactive version of the find-a-changelog guide (#2252), and a natural upgrade to /submit.

Flow sketch:

  1. Index first, zero fetches. Resolve the domain against what we already track (/v1/lookups/by-domain, org catalog). If it's indexed: show the org page, sources, and .atom links. Most lookups should end here for free.
  2. Heuristic scan (bounded, cacheable) for unknown domains — this is mostly machinery we already have in packages/adapters / the discovery pre-checks:
    • /.well-known/releases.json + changelog.json, <link rel="changelog">
    • guessable URLs (/changelog, /updates, /releases, changelog.{domain}, …)
    • feed autodiscovery + provider fingerprinting (Mintlify/ReadMe/Fern/etc. known paths)
    • GitHub org guess from homepage links
  3. Cheap AI pass only when heuristics are ambiguous: a Haiku-class verdict on "is this page actually a changelog?" (same shape as the marketing classifier / URL evaluation we already run — reuse an existing model lane, no per-feature model var).
  4. Outcome ties into submission: found → pre-filled submit/track request (the tracking_requested_at demand signal from the listing lane); already indexed → org page; nothing found → plain "couldn't find one" plus the manifest pitch for owners.

Doing it responsibly

This is the actual design problem — an open "scan any domain" endpoint is a free fetch proxy and an AI-spend faucet if built naively:

  • Cache by domain (KV/D1, generous TTL). A domain gets scanned once, not per visitor.
  • Rate limits per-IP and per-domain, CF-native, mirroring POST /v1/listing/{validate,activate} (10/min IP, per-domain cap).
  • Bounded fetch budget per scan: a fixed small number of requests, target domain only (plus at most one GitHub API call) — never user-supplied arbitrary URLs, so it can't be used as a proxy. SSRF guards as in webhook-url-safety.
  • Opt-out preflight fail-closed: robots.txt / Content-Signal check before fetching content pages, same posture as local-ingest/preflight.ts.
  • AI spend: only on the ambiguous tail, Haiku-class, behind the existing spend-cap gate; heuristics-only is an acceptable degraded mode.
  • Kill switch flag — this one genuinely earns a flag (public, external-fetching, AI-spending).
  • Possibly require a signed-in account for the live-scan tier while keeping index lookups anonymous.

Why bother

  • Converts guide/SEO traffic into submissions and follows: the guide teaches the method, the scanner runs it for you on the spot.
  • The tracking_requested_at signal from scans of unindexed domains is curator-facing demand data we don't get today.
  • Nearly all the discovery logic exists (evaluate/discovery pre-checks, well-known parsing, provider table); the new work is the public surface, the caching/limits, and the responsible-fetch envelope.

Not scoped yet — filing so it doesn't get lost. Related: #1947 (self-serve listing), the guides PR #2252.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions