PROM-67: Scrape Time Rule Evaluation - #67
roidelapluie wants to merge 1 commit into
Conversation
Signed-off-by: Julien Pivotto <291750+roidelapluie@users.noreply.github.com>
60964cb to
c21578d
Compare
|
Wow, I really like this and think it's well thought out and clear the way you put it, thanks @roidelapluie! The only obvious question is of course whether all those scrape-time rules should really live in the main config file, which is inconsistent with normal rules living in separate files. So I wonder if it should be |
There was a problem hiding this comment.
Amazing! It would be great addition 👍🏽
cc @fpetkovski who started initial experiment.
As discussed before, the main obvious challenge is a relatively limited potential for this feature vs what people actually want, so:
- Quick 5m rollups (rate/irate) and drop the rest.
- Instance label drop + aggregation across targets (or even scrapers!) and drop rest.
- Drop metrics conditionally using existing TSDB data.
This means that if we accept this we might need to manage user expectations, but more we will need to solve this problem anyway eventually, with likely a 3rd solution (e.g. cross process aggregation proxy like vmagent is doing or aggregation 5m stateful layer in Prometheus process). This means user will have likely 3 types of rules to choose from (unless we reuse one of existing config surfaces).
I think that's acceptable and it's good to iterate here, just want to make the consequences clear here (:
The only obvious question is of course whether all those scrape-time rules should really live in the main config file, which is inconsistent with normal rules living in separate files. So I wonder if it should be scrape_rule_files instead of scrape_rules (or optionally in addition). But maybe we don't expect people to use many scrape-time rules at all, so putting them into the main config file is more doable?
Good question. Given it's expected from users to DROP certain metrics based on scrape rule it would be more difficult to see what's happening if the scrape rules were in the remote place. Plus as mentioned before, scrape rules have relatively limited usability, so maybe there won't be that many? Previous small discussion.
Good point. The use case for aggregating across instances is probably going to be way more important than the one for aggregating within an instance. FWIW, I have no strong opinion on whether the kind of scrape-time rule evaluation as laid out in this proposal here should be added to Prometheus or not, I can find arguments either way. I just think the proposal is well done :)
Agreed with both, even though it's a bit odd :) |
|
Thanks for writing this out @roidelapluie. Applying aggregations before SD relabeling makes perfect sense and I think will solve the caching issue described in my POC. |
ideally we would only add the "relevant" metrics to the in memory cache, by pre-parsing rules and extracting matchers as descibed in this document, but that could be added in a second PR. |
|
Hello from the Prometheus bug-scrub! Are there any updates since November? Is this likely to move forward? |
No description provided.