Skip to content

9262d43a - Report an indexer outage once instead of on every failed operation - #131

Merged
davidleomay merged 3 commits into
d-EURO:developfrom
Daniel-DFX:fix/indexer-outage-logging
Sep 23, 2026
Merged

davidleomay merged 3 commits into
d-EURO:developfrom
Daniel-DFX:fix/indexer-outage-logging

Conversation

@Daniel-DFX

@Daniel-DFX Daniel-DFX commented Sep 22, 2026 •

Copy link
Copy Markdown

EN:
During an indexer outage the API currently logs two error lines per failed operation and update cycle, for as long as the outage lasts. This PR logs one error line when an outage starts, warn for every further failed operation and one recovery line when the indexer answers again. A fallback URL equal to the primary is now treated as "no fallback", and ApiService/SocialMediaService log indexer network failures as warn with their cause inlined. Normal operation is unchanged.

DE:
Bei einem Indexer-Ausfall schreibt die API heute zwei Error-Zeilen pro fehlgeschlagener Operation und Update-Zyklus, solange der Ausfall dauert. Dieser PR loggt ein error beim Beginn eines Ausfalls, warn für jede weitere fehlgeschlagene Operation und eine Recovery-Zeile, sobald der Indexer wieder antwortet. Eine Fallback-URL gleich der primären gilt nun als „kein Fallback“, und ApiService/SocialMediaService loggen Indexer-Netzwerkfehler als warn mit inline Ursache. Der Normalbetrieb bleibt unverändert.

Details

Problem

  • ApolloLink.from([errorLink, routingLink, retryLink, httpLink]): errorLink only sees a network error after RetryLink used its 3 attempts.
  • The deployment environments set CONFIG_INDEXER_FALLBACK_URL equal to CONFIG_INDEXER_URL (there is no second indexer; unsetting it would fall back to the dev indexer default in api.config.ts). With equal URLs sentToFallback was true for every request, so every network error went straight to logger.error.
  • Per update cycle, each failed operation then produced an ApiApolloConfig error, a Failed to update … after …ms: error from ApiService (without cause text) and, for the social-media queries, an Error while sending … updates error. In production a 5-second reverse-proxy restart produced ~20 error lines; on the dev environment one day of outages produced ~27k.
  • The production formatter (api.main.ts) prints only info.message, so the second argument of logger.error(msg, err) is dropped: ApiService lines carried no cause.

Changes

api.apollo.config.ts

  • New outermost outageLink sees each operation's final outcome once (after retries and a possible fallback attempt). A final network failure logs error if no outage is active and marks it active; while active it logs warn. The first operation that completes with a result logs one [Ponder] Indexer reachable again after Ns (M failed operations) line and clears the state. Reachability is reported on complete, not on next: for a non-2xx response carrying data and errors, Apollo emits the result before the network error, and RetryLink forwards that result from every attempt. A completed result carrying GraphQL errors counts as reachable; GraphQL-error logging is unchanged.
  • HAS_FALLBACK: a fallback only exists if CONFIG.indexerFallback is set and differs from CONFIG.indexer ignoring trailing /. Without one, errorLink explicitly does nothing for network errors (no switch, no extra forward). With a distinct fallback the existing warn + activateFallback() + forward(operation) path is unchanged; the final failure (also the one of the forwarded fallback attempt, which onError never saw before) goes through the outage logic.
  • Exported isIndexerNetworkError(err): ApolloError with networkError set.

api.service.ts

  • Failed to update … after …ms: <cause>: warn for indexer network errors, error otherwise; throw err kept.
  • Cause inlined (${err?.message ?? err}) in Failed to update social media, Error in updateWorkflow, Error getting block number.

socialmedia/socialmedia.service.ts

  • Saving, frontend code, trade, bridge and mint update errors: warn for indexer network errors, error otherwise (Telegram/Twitter send failures stay error). sendUpdates is unchanged since it never queries the indexer.

Expected production effect

During an indexer outage: one error line per outage instead of two per operation and cycle; ApiService/SocialMediaService indexer lines become warn with their cause; one log line on recovery. Normal operation unchanged.

Verification

  • yarn install --frozen-lockfile && yarn build: green, no lockfile changes.
  • npx prettier --check api.apollo.config.ts api.service.ts socialmedia/socialmedia.service.ts (repo prettier 3.3.2): green. yarn lint was not used as evidence: its globs {src,apis,libs,test}/** don't match this repo's root-level layout, so it would not check these files.
  • No committed tests: the repo has no working test setup (Jest rootDir points to a non-existent src/, CI only runs build + lint), consistent with fix: reclassify Apollo network-error log severity #117, fix: attribute network errors to the URL the request was sent to #121 and Retry transient Ponder GraphQL network errors #125 which changed the same file. Instead, an uncommitted harness loads the real compiled PONDER_CLIENT (dist/api.apollo.config.js) against local stub HTTP servers that toggle between 502 and a valid GraphQL response, and captures the log lines via Logger.overrideLogger. One process per scenario, run at the final head (23/23 checks pass):
    • a equal URLs: 3 cycles × 4 operations down → exactly one error, 11 warn, no fallback switch; indexer back → exactly one recovery line, then silence; second outage → new single error.
    • d URLs that differ only by a trailing /: same as a.
    • e distinct URLs: primary down → warn + Switching to fallback, query served by the fallback; fallback down too → one error, rest warn; fallback back → one recovery line.
    • e2 distinct URLs, both down from the start: primary warn + switch, final failure one error, rest warn.
    • f isIndexerNetworkError: true for a rejected query on 502, false for a plain Error.
    • g indexer back but answering with a GraphQL error: GraphQL error logged as before, recovery line, isIndexerNetworkError false.
    • h during an outage the indexer answers 401 with data + errors: no spurious recovery line, a following 502 is still warn of the same outage, recovery once it answers normally.
Harness source
// Uncommitted harness: loads the real compiled PONDER_CLIENT against local stub
// indexers whose responses toggle between 502 and a valid GraphQL response.
// Usage: node harness.js <scenario> (run from the repo root after `yarn build`)
const http = require('http');
const path = require('path');
const REPO = process.cwd();
const { Logger } = require(path.join(REPO, 'node_modules/@nestjs/common'));
const { gql } = require(path.join(REPO, 'node_modules/@apollo/client/core'));

// Capture every log line as "<level> [<context>]: <message>" (same shape as production)
const lines = [];
const capture = (level) => (message, ...rest) => {
	const ctx = typeof rest[rest.length - 1] === 'string' ? rest[rest.length - 1] : '?';
	const line = `${level} [${ctx}]: ${message}`;
	lines.push({ level, ctx, line });
	console.log('  ' + line);
};
Logger.overrideLogger({ log: capture('info'), error: capture('error'), warn: capture('warn'), debug: () => {}, verbose: () => {} });

function stub(name) {
	const s = { name, up: true, hits: 0 };
	s.server = http.createServer((req, res) => {
		req.resume();
		req.on('end', () => {
			s.hits++;
			if (s.partial) {
				res.writeHead(401, { 'content-type': 'application/json' });
				return res.end(JSON.stringify({ data: { ping: null }, errors: [{ message: 'unauthorized' }] }));
			}
			if (s.gqlError) {
				res.writeHead(200, { 'content-type': 'application/json' });
				return res.end(JSON.stringify({ data: null, errors: [{ message: 'Cannot query field "ping"' }] }));
			}
			if (!s.up) {
				res.writeHead(502, { 'content-type': 'text/html' });
				return res.end('<html>502 Bad Gateway</html>');
			}
			res.writeHead(200, { 'content-type': 'application/json' });
			res.end(JSON.stringify({ data: { ping: name } }));
		});
	});
	return new Promise((r) => s.server.listen(0, '127.0.0.1', () => r(s)));
}
const url = (s, slash = false) => `http://127.0.0.1:${s.server.address().port}${slash ? '/' : ''}`;

function load(indexer, fallback) {
	Object.assign(process.env, {
		RPC_URL_MAINNET: 'http://127.0.0.1:1',
		RPC_URL_POLYGON: 'http://127.0.0.1:1',
		COINGECKO_BASE_URL: 'http://127.0.0.1:1',
		CONFIG_INDEXER_URL: indexer,
		CONFIG_INDEXER_FALLBACK_URL: fallback,
	});
	return require(path.join(REPO, 'dist/api.apollo.config.js'));
}

async function run(client, op) {
	try {
		const r = await client.query({ query: gql(`query ${op} { ping }`), fetchPolicy: 'no-cache' });
		return `ok(${r.data.ping})`;
	} catch (e) {
		return `rejected(${e.constructor.name})`;
	}
}

async function cycles(client, n, tag) {
	for (let c = 1; c <= n; c++) {
		for (const op of ['GetMinters', 'GetPositions', 'GetPrices', 'GetSavings']) {
			const r = await run(client, op);
			console.log(`    ${tag} cycle ${c} ${op} -> ${r}`);
		}
	}
}

function count(from, level, ctx = 'ApiApolloConfig') {
	return lines.slice(from).filter((l) => l.level === level && l.ctx === ctx).length;
}
function check(label, cond) {
	console.log(`  CHECK ${cond ? 'PASS' : 'FAIL'}: ${label}`);
	if (!cond) process.exitCode = 1;
}

async function equalUrls(trailingSlashDiff) {
	const a = await stub('primary');
	const { PONDER_CLIENT } = trailingSlashDiff ? load(url(a, true), url(a)) : load(url(a), url(a));
	console.log(`  indexer=${process.env.CONFIG_INDEXER_URL} fallback=${process.env.CONFIG_INDEXER_FALLBACK_URL}`);

	console.log('  -- outage 1: indexer down, 3 cycles x 4 operations');
	a.up = false;
	let mark = lines.length;
	await cycles(PONDER_CLIENT, 3, 'down');
	check('exactly one error line from ApiApolloConfig', count(mark, 'error') === 1);
	check('11 warn lines for the other failed operations', count(mark, 'warn') === 11);
	check('no fallback switch', !lines.slice(mark).some((l) => l.line.includes('Switching to fallback')));

	console.log('  -- indexer back: 2 cycles');
	a.up = true;
	mark = lines.length;
	await cycles(PONDER_CLIENT, 2, 'up');
	check('exactly one recovery log line, nothing else', count(mark, 'info') === 1 && lines.length - mark === 1);

	console.log('  -- outage 2');
	a.up = false;
	mark = lines.length;
	await cycles(PONDER_CLIENT, 1, 'down');
	check('second outage: new single error line', count(mark, 'error') === 1 && count(mark, 'warn') === 3);
	a.server.close();
}

async function distinctUrls() {
	const a = await stub('primary');
	const b = await stub('fallback');
	const { PONDER_CLIENT } = load(url(a), url(b));
	console.log(`  indexer=${url(a)} fallback=${url(b)}`);

	console.log('  -- primary down, fallback up');
	a.up = false;
	let mark = lines.length;
	const r = await run(PONDER_CLIENT, 'GetMinters');
	console.log(`    GetMinters -> ${r}`);
	check('served by fallback', r === 'ok(fallback)');
	check('warn + Switching to fallback, no error', count(mark, 'warn') === 2 && lines.some((l) => l.line.includes('Switching to fallback')) && count(mark, 'error') === 0);

	console.log('  -- fallback down too (fallback window active)');
	b.up = false;
	mark = lines.length;
	await cycles(PONDER_CLIENT, 2, 'both-down');
	check('exactly one error, rest warn', count(mark, 'error') === 1 && count(mark, 'warn') === 7);

	console.log('  -- fallback back');
	b.up = true;
	mark = lines.length;
	await cycles(PONDER_CLIENT, 1, 'fallback-up');
	check('exactly one recovery line', count(mark, 'info') === 1 && lines.length - mark === 1);
	a.server.close();
	b.server.close();
}

async function bothDownFromStart() {
	const a = await stub('primary');
	const b = await stub('fallback');
	const { PONDER_CLIENT } = load(url(a), url(b));
	a.up = false;
	b.up = false;
	console.log('  -- primary and fallback down from the start');
	const mark = lines.length;
	await cycles(PONDER_CLIENT, 1, 'both-down');
	check('primary failure warns + switches, final failure is one error, rest warn', count(mark, 'error') === 1 && lines.slice(mark)[0].level === 'warn' && lines.some((l) => l.line.includes('Switching to fallback')));
	check('fallback received the forwarded attempt', b.hits > 0);
	a.server.close();
	b.server.close();
}

async function classify() {
	const a = await stub('primary');
	const { PONDER_CLIENT, isIndexerNetworkError } = load(url(a), url(a));
	a.up = false;
	let err;
	try {
		await PONDER_CLIENT.query({ query: gql('query GetMinters { ping }'), fetchPolicy: 'no-cache' });
	} catch (e) {
		err = e;
	}
	check('isIndexerNetworkError(rejected query on 502) === true', isIndexerNetworkError(err) === true);
	check('isIndexerNetworkError(new Error()) === false', isIndexerNetworkError(new Error('boom')) === false);
	a.server.close();
}

async function graphqlErrorMeansReachable() {
	const a = await stub('primary');
	const { PONDER_CLIENT, isIndexerNetworkError } = load(url(a), url(a));
	a.up = false;
	await cycles(PONDER_CLIENT, 1, 'down');
	console.log('  -- indexer reachable again but answers with a GraphQL error');
	a.up = true;
	a.gqlError = true;
	const mark = lines.length;
	let err;
	try {
		await PONDER_CLIENT.query({ query: gql('query GetMinters { ping }'), fetchPolicy: 'no-cache' });
	} catch (e) {
		err = e;
	}
	check('GraphQL error logged as before + recovery line, no network-error line', count(mark, 'info') === 1 && count(mark, 'error') === 1 && lines.slice(mark).some((l) => l.line.includes('[GraphQL error')) && count(mark, 'warn') === 0);
	check('isIndexerNetworkError(GraphQL-error rejection) === false', isIndexerNetworkError(err) === false);
	a.server.close();
}

async function failedResponseWithData() {
	const a = await stub('primary');
	const { PONDER_CLIENT } = load(url(a), url(a));
	a.up = false;
	await cycles(PONDER_CLIENT, 1, 'down');
	console.log('  -- outage continues, indexer answers non-2xx with data + errors');
	a.partial = true;
	const mark = lines.length;
	await cycles(PONDER_CLIENT, 1, 'partial');
	check('no spurious recovery line (GraphQL errors logged as before)', count(mark, 'info') === 0);
	console.log('  -- 502 again: outage must still be active');
	a.partial = false;
	const mark1 = lines.length;
	await cycles(PONDER_CLIENT, 1, 'down');
	check('still the same outage: warn only, no new error', count(mark1, 'error') === 0 && count(mark1, 'warn') === 4);
	a.up = true;
	const mark2 = lines.length;
	await cycles(PONDER_CLIENT, 1, 'up');
	check('recovery once the indexer answers normally', count(mark2, 'info') === 1 && lines.length - mark2 === 1);
	a.server.close();
}

const scenarios = { h: failedResponseWithData, g: graphqlErrorMeansReachable,  a: () => equalUrls(false), d: () => equalUrls(true), e: distinctUrls, e2: bothDownFromStart, f: classify };
scenarios[process.argv[2]]().then(() => process.exit());
Harness output
=== scenario a
  indexer=http://127.0.0.1:65440 fallback=http://127.0.0.1:65440
  -- outage 1: indexer down, 3 cycles x 4 operations
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 2 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 2 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 2 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 2 GetSavings -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 3 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 3 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 3 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 3 GetSavings -> rejected(ApolloError)
  CHECK PASS: exactly one error line from ApiApolloConfig
  CHECK PASS: 11 warn lines for the other failed operations
  CHECK PASS: no fallback switch
  -- indexer back: 2 cycles
  info [ApiApolloConfig]: [Ponder] Indexer reachable again after 7s (12 failed operations)
    up cycle 1 GetMinters -> ok(primary)
    up cycle 1 GetPositions -> ok(primary)
    up cycle 1 GetPrices -> ok(primary)
    up cycle 1 GetSavings -> ok(primary)
    up cycle 2 GetMinters -> ok(primary)
    up cycle 2 GetPositions -> ok(primary)
    up cycle 2 GetPrices -> ok(primary)
    up cycle 2 GetSavings -> ok(primary)
  CHECK PASS: exactly one recovery log line, nothing else
  -- outage 2
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  CHECK PASS: second outage: new single error line
exit=0
=== scenario d
  indexer=http://127.0.0.1:65443/ fallback=http://127.0.0.1:65443
  -- outage 1: indexer down, 3 cycles x 4 operations
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 2 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 2 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 2 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 2 GetSavings -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 3 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 3 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 3 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 3 GetSavings -> rejected(ApolloError)
  CHECK PASS: exactly one error line from ApiApolloConfig
  CHECK PASS: 11 warn lines for the other failed operations
  CHECK PASS: no fallback switch
  -- indexer back: 2 cycles
  info [ApiApolloConfig]: [Ponder] Indexer reachable again after 6s (12 failed operations)
    up cycle 1 GetMinters -> ok(primary)
    up cycle 1 GetPositions -> ok(primary)
    up cycle 1 GetPrices -> ok(primary)
    up cycle 1 GetSavings -> ok(primary)
    up cycle 2 GetMinters -> ok(primary)
    up cycle 2 GetPositions -> ok(primary)
    up cycle 2 GetPrices -> ok(primary)
    up cycle 2 GetSavings -> ok(primary)
  CHECK PASS: exactly one recovery log line, nothing else
  -- outage 2
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  CHECK PASS: second outage: new single error line
exit=0
=== scenario e
  indexer=http://127.0.0.1:65445 fallback=http://127.0.0.1:65446
  -- primary down, fallback up
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
  warn [ApiApolloConfig]: [Ponder] Switching to fallback for 10min: http://127.0.0.1:65446
    GetMinters -> ok(fallback)
  CHECK PASS: served by fallback
  CHECK PASS: warn + Switching to fallback, no error
  -- fallback down too (fallback window active)
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    both-down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    both-down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    both-down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    both-down cycle 1 GetSavings -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    both-down cycle 2 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    both-down cycle 2 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    both-down cycle 2 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    both-down cycle 2 GetSavings -> rejected(ApolloError)
  CHECK PASS: exactly one error, rest warn
  -- fallback back
  info [ApiApolloConfig]: [Ponder] Indexer reachable again after 4s (8 failed operations)
    fallback-up cycle 1 GetMinters -> ok(fallback)
    fallback-up cycle 1 GetPositions -> ok(fallback)
    fallback-up cycle 1 GetPrices -> ok(fallback)
    fallback-up cycle 1 GetSavings -> ok(fallback)
  CHECK PASS: exactly one recovery line
exit=0
=== scenario e2
  -- primary and fallback down from the start
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
  warn [ApiApolloConfig]: [Ponder] Switching to fallback for 10min: http://127.0.0.1:65450
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    both-down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    both-down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    both-down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    both-down cycle 1 GetSavings -> rejected(ApolloError)
  CHECK PASS: primary failure warns + switches, final failure is one error, rest warn
  CHECK PASS: fallback received the forwarded attempt
exit=0
=== scenario f
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
  CHECK PASS: isIndexerNetworkError(rejected query on 502) === true
  CHECK PASS: isIndexerNetworkError(new Error()) === false
exit=0
=== scenario g
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  -- indexer reachable again but answers with a GraphQL error
  error [ApiApolloConfig]: [GraphQL error in operation: GetMinters] Cannot query field "ping"
  info [ApiApolloConfig]: [Ponder] Indexer reachable again after 2s (4 failed operations)
  CHECK PASS: GraphQL error logged as before + recovery line, no network-error line
  CHECK PASS: isIndexerNetworkError(GraphQL-error rejection) === false
exit=0
=== scenario h
  error [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  -- outage continues, indexer answers non-2xx with data + errors
  error [ApiApolloConfig]: [GraphQL error in operation: GetMinters] unauthorized
    partial cycle 1 GetMinters -> rejected(ApolloError)
  error [ApiApolloConfig]: [GraphQL error in operation: GetPositions] unauthorized
    partial cycle 1 GetPositions -> rejected(ApolloError)
  error [ApiApolloConfig]: [GraphQL error in operation: GetPrices] unauthorized
    partial cycle 1 GetPrices -> rejected(ApolloError)
  error [ApiApolloConfig]: [GraphQL error in operation: GetSavings] unauthorized
    partial cycle 1 GetSavings -> rejected(ApolloError)
  CHECK PASS: no spurious recovery line (GraphQL errors logged as before)
  -- 502 again: outage must still be active
  warn [ApiApolloConfig]: [Network error in operation: GetMinters] Response not successful: Received status code 502
    down cycle 1 GetMinters -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPositions] Response not successful: Received status code 502
    down cycle 1 GetPositions -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetPrices] Response not successful: Received status code 502
    down cycle 1 GetPrices -> rejected(ApolloError)
  warn [ApiApolloConfig]: [Network error in operation: GetSavings] Response not successful: Received status code 502
    down cycle 1 GetSavings -> rejected(ApolloError)
  CHECK PASS: still the same outage: warn only, no new error
  info [ApiApolloConfig]: [Ponder] Indexer reachable again after 3s (8 failed operations)
    up cycle 1 GetMinters -> ok(primary)
    up cycle 1 GetPositions -> ok(primary)
    up cycle 1 GetPrices -> ok(primary)
    up cycle 1 GetSavings -> ok(primary)
  CHECK PASS: recovery once the indexer answers normally
exit=0

- Track outage state in the Ponder client: the first final network
  failure logs error, following ones warn, and the first successful
  response logs one recovery line.
- Treat a fallback URL equal to the primary (ignoring a trailing slash)
  as no fallback.
- Export isIndexerNetworkError and log indexer network failures in
  ApiService and SocialMediaService as warn.
- Inline the error cause in ApiService log messages, since the
  production formatter drops the second logger argument.
A failed attempt can emit a result before its network error (a non-2xx
response carrying both data and errors), and RetryLink forwards that
result from every attempt. Reporting reachability on the first result
could therefore clear an active outage and log a spurious recovery.
@Daniel-DFX

Copy link
Copy Markdown
Author

EN:
Ready after 4 review passes.
Logs an indexer outage once as an error instead of twice per failed operation and cycle, treats an identical fallback URL as no fallback, and downgrades indexer network failures in ApiService and SocialMediaService to warnings with their cause.

DE:
Bereit nach 4 Review-Durchläufen.
Loggt einen Indexer-Ausfall einmal als Error statt zweimal pro fehlgeschlagener Operation und Zyklus, behandelt eine identische Fallback-URL als keinen Fallback und stuft Indexer-Netzwerkfehler in ApiService und SocialMediaService zu Warnungen mit Ursache herab.

Details

Review passes

  1. Logic: reachability was reported on the first result of an operation, but for a non-2xx response carrying both data and errors Apollo emits that result before the network error, and RetryLink forwards it from every attempt — an active outage could be cleared and a spurious recovery logged. Fixed in 32bbd45: reachability is reported on complete, only if a result was seen. Quality: 0 findings.
  2. Quality: two comments did not match the code — the fall-through comment in errorLink sat above the fallback branch instead of describing the fall-through, and the outage-state comment said "successful response" although a completed result carrying GraphQL errors also counts. Fixed in d5e53ca (comments only). Logic: 0 findings.
  3. Quality and logic: 0 findings.
  4. Second reviewer family, quality and logic at the final head: 0 findings.

One reported point was rejected with reason instead of being applied: shared module-level outage state across concurrent operations. ApiService.updateWorkflow awaits its 14 updates sequentially, so the interleaving described cannot occur here, and the single-flag design is the intended noise reduction.

Gates at the final head d5e53ca

  • Build & Lint: green (run 35745114283, job Build & Lint success). It is the only workflow triggered by pull_request for this base; develop has no classic branch protection (404) and no required-status-check rulesets.
  • Comments: 0 issue comments, 0 reviews, 0 inline comments, 0 review threads — nothing open.
  • Mergeability: mergeable: MERGEABLE, mergeStateStatus: CLEAN.
  • Local: yarn install --frozen-lockfile && yarn build green with no lockfile changes; npx prettier --check green on all three changed files; the harness described in the PR body passes 23/23 checks at this head.
  • All three commits are signed and verified.

@Daniel-DFX
Daniel-DFX marked this pull request as ready for review September 22, 2026 17:40
@davidleomay
davidleomay merged commit 4347ed1 into d-EURO:develop Sep 23, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants