Skip to content

Auto-refetch change-driven sources after a provider outage clears #2194

Description

@zachdunn

Gap

When a Firecrawl webhook delivery's ingest fails (e.g. the provider-quota shutoff of 2026-07-23→28, #2168/#2191), the change event is consumed but the extraction never lands — and nothing retries it. Change-driven sources only run again when the page next changes, so the failed ingest is silently missing until then (or forever, for a page that rarely changes). The staleness digest's new "Provider outage aftermath" section (#2191) tells an operator to re-fetch manually, but the system could heal itself.

Concrete instance: firecrawl-changelog (failed 2026-07-23) and beehiiv product-updates (failed 2026-07-26) sat broken until manual re-fetches on 2026-07-30. (Both turned out lossless after re-fetch, but only because the content was still on the page index.)

Ideas (pick one)

  • Auto-refetch on recovery: when scanProviderHealth computes outageActive: false while entries exist (the aftermath state), dispatch one deterministic update run per affected source instead of only emailing — bounded, idempotent, and uses the existing startDeterministicUpdate gate (spend cap + kill switch still apply).
  • Persist failed webhook payloads for replay: stash the monitor.page diff (R2, content-hash keyed like released-raw) when ingest fails, with a replay hook. Heavier; probably not worth it given the diff can usually be reconstructed from a full re-scrape.

The first option is small and rides entirely on existing machinery.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions