Summary
POST /Cluster/{id}/operation/ is now 31% failures (1,443 of 4,627 calls in 7d), up from 10% (135 of 1,379) over the preceding 23 days. Essentially all of the increase is 404s: 71 → 1,399.
These 404s raise no error and produce no toast, because getSearchByValue deliberately treats 404 as an empty result. That is why none of the existing poll guards (#1527, #1546, #1569, #1602, #1603) can ever engage on them — and why this has been invisible in Error Tracking while still costing ~1,400 doomed control-plane round-trips a week.
The signature
Worst single session in the window: 993 × 404 over 112 unbroken minutes, inter-request gaps pinned at 5.1s, all against one cluster, 100% 404 across 335+ consecutive calls — while other sessions in the same window got clean 200s on other clusters. So this is a stable per-cluster condition, not a transient.
window calls 404 403 500 400 200 fail%
168h..now 4,627 1,399 19 8 15 3,149 31%
720h..168h 1,379 71 25 27 1 1,237 10%
Breadth is flat — 26 sessions / 83 views in 7d vs 23 sessions / 69 views in the prior 23d. What changed is volume per affected session: 3 → 54. This is a latent defect whose cost scales with how long a tab stays open, not a code regression. One long-lived tab produced ~71% of the week's total.
Note POST /HDBInstance/{id}/operation has zero 404s across the same 30 days (2% failures, all 500/400/403/401). The 404s are cluster-path-only, which is what pins the source below.
Mechanism
useChallengeCertificates(cluster.id) → getChallengeCertificatesQueryOptions polls every 5s, unconditionally, whenever the cluster Domains table is mounted:
src/features/cluster/domains/queries/getChallengeCertificates.ts — refetchInterval: pollUnlessForbidden(5000), enabled: !!clusterId
- it calls
getSearchByValue with forceFabricConnect: true on a clu- id, so getInstanceClient routes it to apiClient.baseURL + /Cluster/{id}/operation (src/config/getInstanceClient.ts:58-59)
getSearchByValue POSTs search_by_value for data.ChallengeCertificate to / → the observed POST /Cluster/{id}/operation/
The 404 is swallowed on purpose (src/integrations/api/instance/database/getSearchByValue.ts):
{ validateStatus: (status) => status >= 200 && status < 400 || status === 404 }
if (response.status === 404) { return { data: [] }; }
A cluster with no data.ChallengeCertificate table therefore answers 404 on every tick, the query succeeds with [], query.state.error is never set, and pollUnlessForbidden cannot stop the timer. The poll re-fires every 5 seconds for as long as the tab is open.
The only consumer of the data is isGeneratingCert in src/features/cluster/domains/constants/tableDefinition.tsx:25-29:
const isGeneratingCert = !!challengeCertificates?.some((cert) => !cert.issueDate || cert.inProgress);
So a 5s cadence exists solely to watch an in-progress certificate — and it runs at full rate precisely when there is nothing to watch.
Why this needs a design decision, not a one-liner
Nothing invalidates the [clusterId, 'ChallengeCertificate'] query key anywhere in the codebase. The 5s poll is currently the only way a newly-started challenge is discovered, so simply stopping or slowing the idle poll would delay the certificate-progress UI in the one flow that matters (bind domain with generateDomainCerts, via setDomainIdsOnCluster).
Options, roughly in increasing order of correctness:
- Idle backoff — keep 5s while
isGeneratingCert, drop to 30–60s otherwise. ~6–12× fewer doomed requests, but adds up to 60s discovery latency after starting a challenge.
- Invalidate on mutation, then idle backoff — have
setDomainIdsOnCluster invalidate [clusterId, 'ChallengeCertificate'] on success so discovery is immediate, which makes an aggressive idle interval (or false) safe. Preferred.
- Distinguish "table absent" from "no rows" — a 404 here means the table has never existed on that cluster; that cannot change until a challenge is created, which option 2 already signals. Treating it as a terminal state would let the poll stop outright.
Option 2 is the smallest change that removes the waste without regressing the cert flow.
Verification notes
- All figures from RUM app
f590deee-…, @type:resource aggregates (not event samples, so no group-by truncation).
- Attribution to
getChallengeCertificates is by elimination: it is the only cluster-scoped 5s POST to /Cluster/{id}/operation/, and getSearchByValue is the only caller that tolerates a 404 instead of throwing. The 5.1s observed cadence and the cluster-only 404 distribution both match. Not confirmed against a live reproduction — the Domains page needs a cluster with a missing ChallengeCertificate table to exercise.
- No customer, org, cluster, instance or session identifiers are included above.
Filed by the daily automated Datadog RUM review (2026-08-17).
Summary
POST /Cluster/{id}/operation/is now 31% failures (1,443 of 4,627 calls in 7d), up from 10% (135 of 1,379) over the preceding 23 days. Essentially all of the increase is 404s: 71 → 1,399.These 404s raise no error and produce no toast, because
getSearchByValuedeliberately treats 404 as an empty result. That is why none of the existing poll guards (#1527, #1546, #1569, #1602, #1603) can ever engage on them — and why this has been invisible in Error Tracking while still costing ~1,400 doomed control-plane round-trips a week.The signature
Worst single session in the window: 993 × 404 over 112 unbroken minutes, inter-request gaps pinned at 5.1s, all against one cluster, 100% 404 across 335+ consecutive calls — while other sessions in the same window got clean
200s on other clusters. So this is a stable per-cluster condition, not a transient.Breadth is flat — 26 sessions / 83 views in 7d vs 23 sessions / 69 views in the prior 23d. What changed is volume per affected session: 3 → 54. This is a latent defect whose cost scales with how long a tab stays open, not a code regression. One long-lived tab produced ~71% of the week's total.
Note
POST /HDBInstance/{id}/operationhas zero 404s across the same 30 days (2% failures, all 500/400/403/401). The 404s are cluster-path-only, which is what pins the source below.Mechanism
useChallengeCertificates(cluster.id)→getChallengeCertificatesQueryOptionspolls every 5s, unconditionally, whenever the cluster Domains table is mounted:src/features/cluster/domains/queries/getChallengeCertificates.ts—refetchInterval: pollUnlessForbidden(5000),enabled: !!clusterIdgetSearchByValuewithforceFabricConnect: trueon aclu-id, sogetInstanceClientroutes it toapiClient.baseURL + /Cluster/{id}/operation(src/config/getInstanceClient.ts:58-59)getSearchByValuePOSTssearch_by_valuefordata.ChallengeCertificateto/→ the observedPOST /Cluster/{id}/operation/The 404 is swallowed on purpose (
src/integrations/api/instance/database/getSearchByValue.ts):A cluster with no
data.ChallengeCertificatetable therefore answers 404 on every tick, the query succeeds with[],query.state.erroris never set, andpollUnlessForbiddencannot stop the timer. The poll re-fires every 5 seconds for as long as the tab is open.The only consumer of the data is
isGeneratingCertinsrc/features/cluster/domains/constants/tableDefinition.tsx:25-29:So a 5s cadence exists solely to watch an in-progress certificate — and it runs at full rate precisely when there is nothing to watch.
Why this needs a design decision, not a one-liner
Nothing invalidates the
[clusterId, 'ChallengeCertificate']query key anywhere in the codebase. The 5s poll is currently the only way a newly-started challenge is discovered, so simply stopping or slowing the idle poll would delay the certificate-progress UI in the one flow that matters (bind domain withgenerateDomainCerts, viasetDomainIdsOnCluster).Options, roughly in increasing order of correctness:
isGeneratingCert, drop to 30–60s otherwise. ~6–12× fewer doomed requests, but adds up to 60s discovery latency after starting a challenge.setDomainIdsOnClusterinvalidate[clusterId, 'ChallengeCertificate']on success so discovery is immediate, which makes an aggressive idle interval (orfalse) safe. Preferred.Option 2 is the smallest change that removes the waste without regressing the cert flow.
Verification notes
f590deee-…,@type:resourceaggregates (not event samples, so no group-by truncation).getChallengeCertificatesis by elimination: it is the only cluster-scoped 5sPOSTto/Cluster/{id}/operation/, andgetSearchByValueis the only caller that tolerates a 404 instead of throwing. The 5.1s observed cadence and the cluster-only 404 distribution both match. Not confirmed against a live reproduction — the Domains page needs a cluster with a missingChallengeCertificatetable to exercise.Filed by the daily automated Datadog RUM review (2026-08-17).