Skip to content

document what exporters should do if it only fails partially to collect metrics #2977

Description

@calestyo

Proposal

Hey.

Could you possible document best practises for exporters, when collection of their metrics fails only partially?

There is this:

Failed scrapes

There are currently two patterns for failed scrapes where the application you’re talking to doesn’t respond or has other problems.

The first is to return a 5xx error.

The second is to have a myexporter_up, e.g. haproxy_up, variable that has a value of 0 or 1 depending on whether the scrape worked.

The latter is better where there’s still some useful metrics you can get even with a failed scrape, such as the HAProxy exporter providing process stats. The former is a tad easier for users to deal with, as up works in the usual way, although you can’t distinguish between the exporter being down and the application being down.

But it’s IMO rather incomplete.

I mean it's clear, that if everything fails one might return a 5xx. But for the partial case <exporter>_up alone doesn't really seem to be enough.

In my example I write an exporter which collects part of its metrics via some SSH interface and the other via some REST interface.
If only either of them fails, the other metrics are still valuable.

And <exporter>_up would only indicate the whole exporter being up/down. So one would need a more granular schema, e.g. in my case <exporter>_ssh_based_metrics_up and <exporter>_rest_based_metrics_up.
But of course that would also be rather useless for the user of that metrics, because they don’t necessarily know which metrics are SSH and which are REST based.

In some cases it might be possible to use a special value for the metrics, which indicates that it’s "invalid"/down, but in general that seems rather a bad idea to me.
Similarly, it doesn’t seem feasible to add an <foo>_up to every metric named <foo>.

Or is the recommendation to simply leave out those metrics that couldn’t be collected? (i.e. not print them at all)?

Well, I', not claiming that I know the best way to handle these cases, so it would be nice if some best current practises could be documented.

Thanks,
Chris.

Activity

  1. changed the title [-]documet what exporters should do if it only fails partially to collect metrics[/-] [+]document what exporters should do if it only fails partially to collect metrics[/+] on Feb 27, 2026
  2. roidelapluie commented on Feb 27, 2026

    @roidelapluie
    Member

    If you have clearly distinct collection paths (e.g. SSH vs REST), it can make sense to either:

    • Run them as two logical exporters, or
    • Expose them as separate modules (e.g. ?module=ssh and ?module=rest, similar to how the blackbox exporter works).

    That way, each scrape target/module has its own up semantics, and a failure in the SSH path doesn’t implicitly degrade the REST path. From the user’s perspective, this is often clearer: each scrape either succeeds or fails for a well-defined collection scope.

    Or, follow the node_exporter collector pattern

    Another well-established approach is what node_exporter does for collectors:

    • Keep a single scrape endpoint.
    • For each collector, expose metrics like: node_scrape_collector_success{collector="foo"} and node_scrape_collector_duration_seconds{collector="foo"}

    This gives you per-subsystem visibility without:

    • Overloading _up
    • Inventing synthetic “invalid” values
    • Adding _up next to every single metric
  3. calestyo commented on Feb 28, 2026

    @calestyo
    ContributorAuthor

    Hey :-)

    • Run them as two logical exporters, or
    • Expose them as separate modules (e.g. ?module=ssh and ?module=rest, similar to how the blackbox exporter works).

    In my case, I think, neither of them makes really a lot sense as SSH vs REST is really just a internal technical implementation detail, that even hopefully goes away (once the software I'm monitoring has migrated all information from it's special SSH interface to REST).
    Right now it's simply that not all information is present in REST, and unfortunately, in order to properly parse their SSH, one also needs information from the REST ;-)

    Or, follow the node_exporter collector pattern

    Which exactly is that?

    For each collector, expose metrics like: node_scrape_collector_success{collector="foo"} and node_scrape_collector_duration_seconds{collector="foo"}

    That is actually a nice approach. There’s still the issue that in my case, both it's rather difficult to do something like collector=ssh and collector=rest, because:

    • as I've said at least metrics from ssh also depend on (some) metrics from rest
    • depending on the version of the software (which btw is dCache), there may be different information present in the REST API,...

    ... so ultimately I'd probably need rather something like: myexporter_metric_success{name="name_of_the_metric"}.
    I’ve had previously already thought about adding simply a myexporter_<metric_name>_failed = 1 if, and only if, the metric collection failed (which I guess is what you meant above with “Adding _up next to every single metric”), right?
    But I guess an overall myexporter_metric_success would be better, or what do you think?

    If you think the myexporter_metric_success{name="name_of_the_metric"} is a good idea:

    1. Would you say a time series for a given metric should only appear with e.g. = 0 if it failed? Or should these always be present with either = 0 or =1?
    2. Would it make sense to document such a schema e.g. in https://prometheus.io/docs/instrumenting/writing_exporters/#failed-scrapes ? I mean other exporters may have the same situation, and it would probably be better if all follow the same naming conventions?
    3. Would maybe *_metric_collection_success be a better name than just *_metric_success? To indicate that it’s not about a success/failure state of the metric itself.

    And apart from that:

    1. What should one do with the actual metric itself, while it cannot be collected? An invalid value is, as you've said, not really nice (and often not possible)... but simply returning the last value where collection still worked is also ugly.
      So leave it out?
  4. krajorama commented on Apr 21, 2026

    @krajorama
    Member

    Hello from the bug scrub!

    Moving to the docs repo, since it's outside the scope of the Prometheus server.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions