9262d43a - Report an indexer outage once instead of on every failed operation - #131
Conversation
- Track outage state in the Ponder client: the first final network failure logs error, following ones warn, and the first successful response logs one recovery line. - Treat a fallback URL equal to the primary (ignoring a trailing slash) as no fallback. - Export isIndexerNetworkError and log indexer network failures in ApiService and SocialMediaService as warn. - Inline the error cause in ApiService log messages, since the production formatter drops the second logger argument.
A failed attempt can emit a result before its network error (a non-2xx response carrying both data and errors), and RetryLink forwards that result from every attempt. Reporting reachability on the first result could therefore clear an active outage and log a spurious recovery.
|
EN: DE: DetailsReview passes
One reported point was rejected with reason instead of being applied: shared module-level outage state across concurrent operations. Gates at the final head
|
EN:
During an indexer outage the API currently logs two error lines per failed operation and update cycle, for as long as the outage lasts. This PR logs one
errorline when an outage starts,warnfor every further failed operation and one recovery line when the indexer answers again. A fallback URL equal to the primary is now treated as "no fallback", andApiService/SocialMediaServicelog indexer network failures aswarnwith their cause inlined. Normal operation is unchanged.DE:
Bei einem Indexer-Ausfall schreibt die API heute zwei Error-Zeilen pro fehlgeschlagener Operation und Update-Zyklus, solange der Ausfall dauert. Dieser PR loggt ein
errorbeim Beginn eines Ausfalls,warnfür jede weitere fehlgeschlagene Operation und eine Recovery-Zeile, sobald der Indexer wieder antwortet. Eine Fallback-URL gleich der primären gilt nun als „kein Fallback“, undApiService/SocialMediaServiceloggen Indexer-Netzwerkfehler alswarnmit inline Ursache. Der Normalbetrieb bleibt unverändert.Details
Problem
ApolloLink.from([errorLink, routingLink, retryLink, httpLink]):errorLinkonly sees a network error afterRetryLinkused its 3 attempts.CONFIG_INDEXER_FALLBACK_URLequal toCONFIG_INDEXER_URL(there is no second indexer; unsetting it would fall back to the dev indexer default inapi.config.ts). With equal URLssentToFallbackwas true for every request, so every network error went straight tologger.error.ApiApolloConfigerror, aFailed to update … after …ms:error fromApiService(without cause text) and, for the social-media queries, anError while sending … updateserror. In production a 5-second reverse-proxy restart produced ~20 error lines; on the dev environment one day of outages produced ~27k.api.main.ts) prints onlyinfo.message, so the second argument oflogger.error(msg, err)is dropped:ApiServicelines carried no cause.Changes
api.apollo.config.tsoutageLinksees each operation's final outcome once (after retries and a possible fallback attempt). A final network failure logserrorif no outage is active and marks it active; while active it logswarn. The first operation that completes with a result logs one[Ponder] Indexer reachable again after Ns (M failed operations)line and clears the state. Reachability is reported oncomplete, not onnext: for a non-2xx response carryingdataanderrors, Apollo emits the result before the network error, andRetryLinkforwards that result from every attempt. A completed result carrying GraphQL errors counts as reachable; GraphQL-error logging is unchanged.HAS_FALLBACK: a fallback only exists ifCONFIG.indexerFallbackis set and differs fromCONFIG.indexerignoring trailing/. Without one,errorLinkexplicitly does nothing for network errors (no switch, no extraforward). With a distinct fallback the existing warn +activateFallback()+forward(operation)path is unchanged; the final failure (also the one of the forwarded fallback attempt, whichonErrornever saw before) goes through the outage logic.isIndexerNetworkError(err):ApolloErrorwithnetworkErrorset.api.service.tsFailed to update … after …ms: <cause>:warnfor indexer network errors,errorotherwise;throw errkept.${err?.message ?? err}) inFailed to update social media,Error in updateWorkflow,Error getting block number.socialmedia/socialmedia.service.tswarnfor indexer network errors,errorotherwise (Telegram/Twitter send failures stayerror).sendUpdatesis unchanged since it never queries the indexer.Expected production effect
During an indexer outage: one
errorline per outage instead of two per operation and cycle;ApiService/SocialMediaServiceindexer lines becomewarnwith their cause; onelogline on recovery. Normal operation unchanged.Verification
yarn install --frozen-lockfile && yarn build: green, no lockfile changes.npx prettier --check api.apollo.config.ts api.service.ts socialmedia/socialmedia.service.ts(repo prettier 3.3.2): green.yarn lintwas not used as evidence: its globs{src,apis,libs,test}/**don't match this repo's root-level layout, so it would not check these files.rootDirpoints to a non-existentsrc/, CI only runs build + lint), consistent with fix: reclassify Apollo network-error log severity #117, fix: attribute network errors to the URL the request was sent to #121 and Retry transient Ponder GraphQL network errors #125 which changed the same file. Instead, an uncommitted harness loads the real compiledPONDER_CLIENT(dist/api.apollo.config.js) against local stub HTTP servers that toggle between 502 and a valid GraphQL response, and captures the log lines viaLogger.overrideLogger. One process per scenario, run at the final head (23/23 checks pass):aequal URLs: 3 cycles × 4 operations down → exactly oneerror, 11warn, no fallback switch; indexer back → exactly one recovery line, then silence; second outage → new singleerror.dURLs that differ only by a trailing/: same asa.edistinct URLs: primary down →warn+Switching to fallback, query served by the fallback; fallback down too → oneerror, restwarn; fallback back → one recovery line.e2distinct URLs, both down from the start: primarywarn+ switch, final failure oneerror, restwarn.fisIndexerNetworkError: true for a rejected query on 502, false for a plainError.gindexer back but answering with a GraphQL error: GraphQL error logged as before, recovery line,isIndexerNetworkErrorfalse.hduring an outage the indexer answers 401 withdata+errors: no spurious recovery line, a following 502 is stillwarnof the same outage, recovery once it answers normally.Harness source
Harness output