This repo contains a Prometheus exporter for Amazon Elastic Container Service (ECS) that publishes ECS task infra metrics in the Prometheus exposition formats.
Run the following container as a sidecar on ECS tasks:
quay.io/prometheuscommunity/ecs-exporter:v0.4.0
An example Fargate task definition that includes the container is available.
To add ECS exporter to your existing ECS task:
- Go to ECS task definitions.
- Click on "Create new revision".
- Scroll down to "Container definitions" and click on "Add container".
- Set "ecs-exporter" as container name.
- Copy the container image URL from above. (Use the tag for the latest release.)
- Add tcp/9779 as a port mapping.
- Click on "Add" to return back to task definition page.
- Click on "Create" to create a new revision.
By default, it publishes Prometheus metrics on ":9779/metrics". The exporter in this repo can be a useful complementary sidecar for the scenario described in this blog post. Adding this sidecar to the ECS task definition would export task-level metrics in addition to the custom metrics described in the blog.
All metrics exported by ecs_exporter are sourced from the ECS task metadata
API
accessible from within every ECS task. AWS can make, and on occasion has
made,
unannounced (or at least unversioned) breaking changes to the data served from
this API, especially the container-level stats served from the /task/stats
endpoint. Metrics emitted by ecs_exporter may spontaneously break as a result,
in which case we may need to make breaking changes to ecs_exporter to keep up.
In light of these conditions, we currently do not have plans to cut a 1.0 release. When necessary, breaking changes will continue to land in minor version releases. The release notes will document any breaking changes as they come.
None. You may
join
to ecs_task_metadata_info to add task-level metadata (such as the task ARN) to
task-level or any other metrics emitted by ecs_exporter.
- container_name: Name of the container (as in the ECS task definition) associated with the metric.
- interface: Network interface device associated with the metric.
The ECS task stats endpoint exposes Docker-compatible Linux cgroup memory
accounting. See the kernel documentation for the underlying cgroup v1 memory
statistics
and cgroup v2 memory
interface,
and Docker's explanation of the difference between raw API usage and the
working-set-like value displayed by docker stats.
Some notes on the metrics:
- Exact metric names and definitions are in the example output.
- The raw "usage" metric includes page cache and accounted kernel memory.
- Memory breakdown metrics overlap and do not sum to a meaningful number.
- Some breakdown metrics are only available with one cgroup version. Unsupported metrics are absent rather than reported as zero. (Currently, Fargate is always on cgroup v1, Managed Instances uses cgroup v2, and EC2 depends on the instance's configuration. AmazonLinux 2 used v1, while AmazonLinux 2023 uses v2, for example.)
- The cgroup v1 memory limit failures metric counts memory charges that encounter the limit, not OOM kills. Its rate indicates limit pressure, but a nonzero value does not mean that an allocation or container failed.
- The container limit metric is sourced from ECS configuration. If none is set, the task-level limit is used instead. That fallback value is shared capacity and must not be summed across containers as if it were independently reserved for each one.
Check out the metrics snapshots which
contain sample metrics emitted by ecs_exporter in the Prometheus text
format
you should expect to see on /metrics. Note that these snapshots behave as if
--web.disable-exporter-metrics were passed when running ecs_exporter, such
that standard client_golang
metrics are not included.