Skip to content

refactor(metrics): Rewrite metric exporter - #309

Merged
ketiltrout merged 1 commit into
chime-upgradefrom
redo_prom_client
Sep 28, 2026
Merged

ketiltrout merged 1 commit into
chime-upgradefrom
redo_prom_client

Conversation

@ketiltrout

@ketiltrout ketiltrout commented Sep 26, 2026 •

Copy link
Copy Markdown
Member

This is a complete rewrite of the metric exporter to fix problems with the old implementation, the chief one being it didn't work with any prometheus_client >= 0.8 (see #198).

The old metrics exporter was a separate thread-based HTTP server created by the qworker within the request_forwarder's context and listening on a second port.

The new way it's done now is the things that used to directly update metrics, now push these updates into redis and there's a separate Sanic worker (the metric aggregator) that pops these updates out of redis and keeps the metrics up-to-date.

Access to the metrics is through a /metrics route in the Sanic server now (no secondary web server used) and this route dispatches a request to the aggregator which responds with the rendered metrics, which are then passed back to the client.

Because there's no separate prom-client HTTP server anymore, the metrics_port config value is no longer used. cocod will emit a warning about that key being ignored if encountered during start-up.

There are no changes to the metrics provided by cocod.

Closes #198
Closes #204

This is a complete rewrite of the metric exporter to fix problems
with the old implementation, the chief one being it didn't work
with any `prometheus_client >= 0.8` (see #198).

The old metrics exporter was a separate thread-based HTTP server
created by the qworker within the request_forwarder's context and
listening on a second port.

The new way it's done now is the things that used to directly
update metrics, now push these updates into redis and there's a separate
Sanic worker (the metric aggregator) that pops these updates out of
redis and keeps the metrics up-to-date.

Access to the metrics is through a `/metrics` route in the Sanic server
now (no secondary web server used) and this route dispatches a request
to the aggregator which responds with the rendered metrics, which are
then passed back to the client.

Because there's no separate prom-client HTTP server anymore, the
`metrics_port` config value is no longer used.  cocod will emit a
warning about that key being ignored if encountered during start-up.
@ketiltrout
ketiltrout added this pull request to stack #311 September 26, 2026 06:57
@ketiltrout
ketiltrout merged commit 22b7050 into chime-upgrade Sep 28, 2026
5 checks passed
@ketiltrout
ketiltrout deleted the redo_prom_client branch September 28, 2026 22:27
ketiltrout added a commit that referenced this pull request Sep 28, 2026
This is the last of the old tests that needed recovery. It was waiting
on the prometheus client fix.

This also moves the creation of the queue update script down to where
it's first used. If coco isn't using a limited queue, it won't be used
at all.

Requires #309 because the test checks the metrics.

Closes #252
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

metrics endpoint crashes Metrics server broken with prometheus-client 0.8

2 participants