Skip to content

[Bug]: Lack of logrotate configuration causes disk exhaustion and unit crashes for high-traffic ingresses #638

Description

@alejdg

Bug Description

We are experiencing disk exhaustion issues on units running the haproxy charm when handling high-traffic ingresses. Because the charm currently does not expose any configuration for logrotate, it falls back to the default log rotation policy, which only rotates logs once a week.

For high-volume environments, this weekly rotation is far too slow, causing the HAProxy logs to fill up the entire disk before the rotation is ever triggered. When the disk is filled, the charm goes into error and stops serving.

We already expanded the disk size but this is not a sustainable solution, we have and application that stops every 2 days. We need the ability to configure log rotation (e.g., size-based rotation, daily/hourly frequency, or capping retention) directly through the charm.

Impact

Medium (functionality degraded, workaround exists)

Impact Rationale

No response

To Reproduce

  1. Deploy the haproxy charm and relate it to workloads as an ingress.
  2. Direct a high and continuous volume of traffic to the ingress.
  3. Monitor the disk usage on the HAProxy unit.
  4. Observe that the HAProxy log files grow unbounded over the course of the week because log rotation only runs weekly.
  5. Once the disk fills up, observe the unit going into error.

Environment

channel=2.8/edge
revision=537
running on canonical-k8s

Relevant log output

2026-08-14 05:44:13 INFO juju.worker.uniter.operation runhook.go:186 ran "haproxy-route-relation-created" hook (via hook dispatching script: dispatch)
2026-08-14 05:44:14 ERROR unit.ingress-ps7-certification/15.juju-log server.go:405 haproxy-route:1220: 2 validation errors for RequirerApplicationData
service
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
ports
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
Traceback (most recent call last):
  File "/var/lib/juju/agents/unit-ingress-ps7-certification-15/charm/lib/charms/haproxy/v2/haproxy_route.py", line 272, in load
    return cls.model_validate_json(json.dumps(data))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/var/lib/juju/agents/unit-ingress-ps7-certification-15/charm/venv/lib/python3.12/site-packages/pydantic/main.py", line 782, in model_validate_json
    return cls.__pydantic_validator__.validate_json(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
pydantic_core._pydantic_core.ValidationError: 2 validation errors for RequirerApplicationData
service
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
ports
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
2026-08-14 05:44:14 ERROR unit.ingress-ps7-certification/15.juju-log server.go:405 haproxy-route:1220: Invalid requirer application data for remote-dcefa3f1833d4dc18e0cd06dcb274f63
2026-08-14 05:44:14 ERROR unit.ingress-ps7-certification/15.juju-log server.go:405 haproxy-route:1220: 2 validation errors for RequirerApplicationData
service
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
ports
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
Traceback (most recent call last):
  File "/var/lib/juju/agents/unit-ingress-ps7-certification-15/charm/lib/charms/haproxy/v2/haproxy_route.py", line 272, in load
    return cls.model_validate_json(json.dumps(data))
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/var/lib/juju/agents/unit-ingress-ps7-certification-15/charm/venv/lib/python3.12/site-packages/pydantic/main.py", line 782, in model_validate_json
    return cls.__pydantic_validator__.validate_json(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
pydantic_core._pydantic_core.ValidationError: 2 validation errors for RequirerApplicationData
service
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
ports
  Field required [type=missing, input_value={}, input_type=dict]
    For further information visit https://errors.pydantic.dev/2.13/v/missing
2026-08-14 05:44:14 ERROR unit.ingress-ps7-certification/15.juju-log server.go:405 haproxy-route:1220: Invalid requirer application data for remote-dcefa3f1833d4dc18e0cd06dcb274f63
2026-08-14 05:44:16 INFO unit.ingress-ps7-certification/15.juju-log server.go:405 haproxy-route:1220: ha integration is not ready, skipping.

Additional context

No response

Activity

  1. seb4stien commented on Sep 2, 2026

    @seb4stien
    Contributor

    @alejdg what is your current logging solution? Is it something charmed, and if it is, could the rotation be managed at this level?

  2. seb4stien commented on Sep 2, 2026

    @seb4stien
    Contributor

    Can you please check your logrotate configuration? The apt package deployed by the charm deploys a daily rotation from what I see in our staging environment:

    root@juju-cef4f1-stg-haproxy-56:~# cat /etc/logrotate.d/haproxy 
    /var/log/haproxy.log {
        daily
        rotate 7
        missingok
        notifempty
        compress
        delaycompress
        postrotate
            [ ! -x /usr/lib/rsyslog/rsyslog-rotate ] || /usr/lib/rsyslog/rsyslog-rotate
        endscript
    }
    
  3. seb4stien commented on Sep 2, 2026

    @seb4stien
    Contributor

    If urgent, please reach out on mattermost as we are currently facing issues with the github2jira synchronization.

  4. syncronize-issues-to-jira commented on Sep 3, 2026

    @syncronize-issues-to-jira

    Thank you for reporting your feedback to us!

    The internal ticket has been created: https://warthogs.atlassian.net/browse/ISD-6548.

    This message was autogenerated

  5. alexdlukens-canonical commented on Sep 22, 2026

    @alexdlukens-canonical

    We are using 50G root disks on our HAProxy units. The logs are uploaded to Loki running in COS, so we have no need for long-term storage on the charm units themselves. We would want a way to override this log-rotate configuration to enforce a max disk usage, instead of a number of days.

    e.g. something like

    /var/log/haproxy.log {
        hourly
        size 50M # would want this to be configurable
        rotate 10 # would want this to be configurable
        compress
        delaycompress
    }
    
  6. beliaev-maksim commented on Sep 29, 2026

    @beliaev-maksim
    Member

    I assume we even do not need the config option, the charm can come with some sensible defaults and the opinionated deployment to use Loki in production

    @alexdlukens-canonical wdyt?

  7. alejdg commented on Sep 29, 2026

    @alejdg
    Author

    After internal discussion, we don't strictly need full logrotate configuration exposed through charm options.

    In enterprise environments with COS integration, local disk logs only serve as a temporary buffer while being scraped into Loki (e.g., via the OTEL collector). Charms should remain opinionated, so exposing user-facing configuration knobs isn't necessary here.

    Instead, the charm could implement a hardcoded, opinionated capping and retention policy. Enforcing a frequent size-based or short-interval rotation (calibrated to safely buffer log volume during peak traffic bursts between scrape intervals) will prevent disk exhaustion while keeping charm management simple.

  8. seb4stien commented on Sep 29, 2026

    @seb4stien
    Contributor

    We discussed the topic internally with @gregory-schiano.
    Our approach was to propose a "logrotate-configurator" charm with the corresponding relation and lib for any machine charm to be able to configure their log rotation.
    The lib would expose methods for the charm to init a proper logrotate, and the relation would let the configurator override the default settings when necessary.
    This way the charm can remain opinionated and still be adapted to specific cases.

    The alternative of adopting an opinionated capping and retention policy could also be a cheaper approach. We will discuss it.
    The main caveat I see is the risk of not fitting all use cases (e.g. not being a good fit for small deployments).

  9. beliaev-maksim commented on Sep 30, 2026

    @beliaev-maksim
    Member

    @seb4stien, do we expect the charm to be used without COS in an enterprise setting?

    If not, local development could use the same 10% of storage. For example, 1 GB would result in 100 MB of data, which is still a lot to store.

  10. seb4stien commented on Oct 6, 2026

    @seb4stien
    Contributor

    I don't know if we have a policy to develop charms assuming they would always be connected to CoS.
    To keep it open, we propose to change the default logrotate configuration to the following:

    /var/log/haproxy.log {
        daily
        rotate 7
        maxsize XXG
        missingok
        notifempty
        compress
        delaycompress
        postrotate
            [ ! -x /usr/lib/rsyslog/rsyslog-rotate ] || /usr/lib/rsyslog/rsyslog-rotate
        endscript
    }
    

    Where XXG will be 10% of the disk size at install time, and we will trigger logrotate hourly to rotate if we hit maxsize.
    This way we keep daily rotation for low volume sites, and rotation happens automatically for heavy workloads.

  11. beliaev-maksim commented on Oct 6, 2026

    @beliaev-maksim
    Member

    @alejdg @alexdlukens-canonical What do you think? Could we saturate within 1 hour? It is still non-deterministic

  12. seb4stien commented on Oct 6, 2026

    @seb4stien
    Contributor

    We can go for 1GB maxsize to make it a bit more deterministic.

  13. alexdlukens-canonical commented on Oct 6, 2026

    @alexdlukens-canonical

    1GB could be saturated in 1hr on high throughput services. From a check on an archive VM on PS7, we had a spike in excess of 49GB/day. I would be in favor of the 10% disk usage default

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions