Gap
When a Firecrawl webhook delivery's ingest fails (e.g. the provider-quota shutoff of 2026-07-23→28, #2168/#2191), the change event is consumed but the extraction never lands — and nothing retries it. Change-driven sources only run again when the page next changes, so the failed ingest is silently missing until then (or forever, for a page that rarely changes). The staleness digest's new "Provider outage aftermath" section (#2191) tells an operator to re-fetch manually, but the system could heal itself.
Concrete instance: firecrawl-changelog (failed 2026-07-23) and beehiiv product-updates (failed 2026-07-26) sat broken until manual re-fetches on 2026-07-30. (Both turned out lossless after re-fetch, but only because the content was still on the page index.)
Ideas (pick one)
- Auto-refetch on recovery: when
scanProviderHealth computes outageActive: false while entries exist (the aftermath state), dispatch one deterministic update run per affected source instead of only emailing — bounded, idempotent, and uses the existing startDeterministicUpdate gate (spend cap + kill switch still apply).
- Persist failed webhook payloads for replay: stash the
monitor.page diff (R2, content-hash keyed like released-raw) when ingest fails, with a replay hook. Heavier; probably not worth it given the diff can usually be reconstructed from a full re-scrape.
The first option is small and rides entirely on existing machinery.
Gap
When a Firecrawl webhook delivery's ingest fails (e.g. the provider-quota shutoff of 2026-07-23→28, #2168/#2191), the change event is consumed but the extraction never lands — and nothing retries it. Change-driven sources only run again when the page next changes, so the failed ingest is silently missing until then (or forever, for a page that rarely changes). The staleness digest's new "Provider outage aftermath" section (#2191) tells an operator to re-fetch manually, but the system could heal itself.
Concrete instance:
firecrawl-changelog(failed 2026-07-23) and beehiivproduct-updates(failed 2026-07-26) sat broken until manual re-fetches on 2026-07-30. (Both turned out lossless after re-fetch, but only because the content was still on the page index.)Ideas (pick one)
scanProviderHealthcomputesoutageActive: falsewhile entries exist (the aftermath state), dispatch one deterministic update run per affected source instead of only emailing — bounded, idempotent, and uses the existingstartDeterministicUpdategate (spend cap + kill switch still apply).monitor.pagediff (R2, content-hash keyed like released-raw) when ingest fails, with a replay hook. Heavier; probably not worth it given the diff can usually be reconstructed from a full re-scrape.The first option is small and rides entirely on existing machinery.