Metrics testkit reference

The metrics testkit provides in-memory metric collection and partial-matching expectation APIs for OpenTelemetry Java MetricData.

Use this page as an API reference for MetricsTestkit, MetricExpectation, PointExpectation, PointSetExpectation, MetricExpectations, and NumberComparison.

For an end-to-end test setup, see Test metrics emitted by your code. For the overview of all signal testkits, see Testkit.

The examples below assume these imports:

import io.opentelemetry.sdk.metrics.data.MetricData
import org.typelevel.otel4s.{Attribute, Attributes}
import org.typelevel.otel4s.oteljava.testkit.AttributesExpectation
import org.typelevel.otel4s.oteljava.testkit.{
  InstrumentationScopeExpectation,
  TelemetryResourceExpectation
}
import org.typelevel.otel4s.oteljava.testkit.metrics._

MetricsTestkit

MetricsTestkit is the signal-specific in-memory backend for metrics.

Member Purpose
MetricsTestkit.inMemory[F]() Creates a Resource[F, MetricsTestkit[F]] backed by an in-memory metric reader.
MetricsTestkit.builder[F] Creates a builder for customizing the underlying SdkMeterProviderBuilder.
meterProvider The otel4s MeterProvider[F] used by code under test.
collectMetrics Collects and returns List[MetricData]. Metrics are recollected on each invocation.

OtelJavaTestkit also exposes meterProvider and collectMetrics when a test needs metrics together with traces or logs.

Partial matching

Metric expectations are partial.

This means:

At the metric level, partial matching is also non-consuming by default: repeating the same MetricExpectation in checkAll(...) does not require two distinct collected metrics.

For example:

MetricExpectation.name("service.counter")

matches any collected metric named service.counter, regardless of its type, points, scope, or resource.

MetricExpectation.sum[Long]("service.counter").value(1L)

matches a long sum metric named service.counter that has at least one point with value 1L.

Numeric metrics

Use MetricExpectation.gauge[A] and MetricExpectation.sum[A] for numeric metrics.

MetricExpectation.gauge[Double]("service.temperature")
MetricExpectation.sum[Long]("service.requests")

The measurement type remains generic. Long and Double use the same API.

For value-only checks, value(...) is the most compact form:

MetricExpectation.sum[Long]("service.requests").value(1L)

If you also care about point attributes, there is a shorthand for exact point matching:

MetricExpectation
  .sum[Long]("service.requests")
  .value(1L, Attribute("http.method", "GET"))

value(value, attributes) uses exact attribute matching. It is equivalent to:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesExact(Attribute("http.method", "GET"))
    )
  )

Use points(PointSetExpectation.exists(...)) directly when you need subset matching or a more detailed point expectation.

Point attributes

Point attributes can be matched in three ways:

Both attributesExact(...) and attributesSubset(...) support Attributes and varargs of Attribute[_].

PointExpectation
  .numeric(1L)
  .attributesExact(Attribute("http.method", "GET"))

PointExpectation
  .numeric(1L)
  .attributesSubset(
    Attribute("http.method", "GET"),
    Attribute("http.route", "/users")
  )

PointExpectation
  .numeric(1L)
  .attributes(
    AttributesExpectation.where("expected only one attribute") { attributes =>
      attributes == Attributes(Attribute("http.method", "GET"))
    }
  )

Empty attributes can be matched with:

PointExpectation.numeric(1L).attributesEmpty

Metric metadata and predicates

Metric expectations can assert metadata on the MetricData itself:

MetricExpectation
  .sum[Long]("service.requests")
  .description("Total requests")
  .unit("1")
  .scopeName("service")

MetricExpectation
  .name("service.requests")
  .where("metric should have at least one point") { metric =>
    !metric.getData.getPoints.isEmpty
  }

Use where(...) when the expectation API does not expose the field you need, or when a test needs a custom invariant over the raw MetricData.

Scope and resource expectations

Metric expectations can also assert the instrumentation scope and telemetry resource.

This is useful when you want tests to preserve the same level of detail as raw MetricData.

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesExact(Attribute("http.method", "GET"))
    )
  )
  .scope(
    InstrumentationScopeExpectation
      .name("service")
      .version("1.0")
      .schemaUrl("https://opentelemetry.io/schemas/1.24.0")
      .attributesEmpty
  )
  .resource(
    TelemetryResourceExpectation.any
      .attributesSubset(Attribute("service.name", "auth-service"))
      .schemaUrl(None)
  )

As with point attributes, scope and resource attribute helpers support both Attributes and varargs. They also provide attributesEmpty when you want to assert that no attributes are present.

Summaries and histograms

The testkit also supports non-numeric point kinds directly.

Summary

MetricExpectation
  .summary("rpc.duration")
  .points(
    PointSetExpectation.exists(
      PointExpectation.summary
        .count(1L)
        .sum(42.0)
    )
  )

Histogram

import org.typelevel.otel4s.metrics.BucketBoundaries

MetricExpectation
  .histogram("http.server.duration")
  .points(
    PointSetExpectation.exists(
      PointExpectation.histogram
        .count(3L)
        .sum(42.0)
        .boundaries(BucketBoundaries(0.1, 1.0, 10.0))
        .counts(1L, 1L, 1L, 0L)
    )
  )

Exponential histogram

MetricExpectation
  .exponentialHistogram("queue.depth")
  .points(
    PointSetExpectation.exists(
      PointExpectation.exponentialHistogram
        .scale(2)
        .count(10L)
        .sum(100.0)
        .zeroCount(0L)
    )
  )

Point-set expectations

Point matching is collection-based. A MetricExpectation does not check points one by one in isolation. Instead, it evaluates a PointSetExpectation against the full point collection, which lets you express presence, absence, cardinality, and collection-wide invariants.

All typed metric expectations expose:

exists

Use exists when at least one collected point must match:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesExact(Attribute("http.method", "GET"))
    )
  )

This is the default mental model for helpers like value(...): they require at least one matching point, not that every point looks the same.

forall

Use forall when every collected point must satisfy the same rule:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.forall(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("kind", "ok"))
    )
  )

forall is useful for invariants such as "every point has the same value shape" or "every point includes this attribute subset". Unlike plain universal quantification over a Scala collection, it fails on an empty point set.

contains and exactly

Use contains when the metric must contain several distinct matching points, but extra points are still allowed:

MetricExpectation
  .sum[Long]("service.requests")
  .containsPoints(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "eu")),
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "us"))
  )

contains enforces distinct matching. If you ask for the same expected point twice, the collected data must contain two matching points.

Use exactly when the metric must contain exactly the expected points and no extras:

MetricExpectation
  .sum[Long]("service.requests")
  .exactlyPoints(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "eu")),
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "us"))
  )

Cardinality combinators

Use pointCount(...) when only the total point count matters:

MetricExpectation
  .sum[Long]("service.requests")
  .pointCount(2)

For lower-level count constraints, use PointSetExpectation directly:

MetricExpectation
  .sum[Long]("service.requests")
  .points(PointSetExpectation.minCount[PointExpectation.NumericPointData[Long]](1))
  .points(PointSetExpectation.maxCount[PointExpectation.NumericPointData[Long]](2))

Use countWhere when the total number of matching points matters more than the exact full set:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.countWhere(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("region", "eu")),
      expected = 2
    )
  )

This is useful when each point still carries extra dimensions that you do not want to enumerate explicitly.

none

Use none or withoutPointsMatching(...) when a point shape must not appear:

MetricExpectation
  .sum[Long]("service.requests")
  .withoutPointsMatching(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "test"))
  )

predicate

Use predicate or pointsWhere(...) for collection-wide assertions that are awkward to express with the built-in combinators:

MetricExpectation
  .sum[Long]("service.requests")
  .pointsWhere("expected exactly EU and US points") { points =>
    points.map(_.attributes).toSet == Set(
      Attributes(Attribute("region", "eu")),
      Attributes(Attribute("region", "us"))
    )
  }

For numeric metrics, the points passed to pointsWhere are typed NumericPointData[A], so you can inspect .value, .attributes, and .underlying.

and and or

Point-set expectations can be combined explicitly:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation
      .contains(
        PointExpectation
          .numeric(1L)
          .attributesSubset(Attribute("region", "eu")),
        PointExpectation
          .numeric(1L)
          .attributesSubset(Attribute("region", "us"))
      )
      .and(PointSetExpectation.count[PointExpectation.NumericPointData[Long]](2))
  )

Use or when a metric may legitimately satisfy one of several point layouts:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation
      .count[PointExpectation.NumericPointData[Long]](1)
      .or(PointSetExpectation.count[PointExpectation.NumericPointData[Long]](2))
  )

Because point expectations accumulate, you can also layer multiple points(...) calls on the same metric:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("region", "eu"))
    )
  )
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("region", "us"))
      )
  )

Each points(...) clause is checked independently against the full collected point set. Chaining points(PointSetExpectation.exists(...)) does not reserve matched points for later clauses, so it should not be used to require multiple distinct points.

If you need distinct matching, use containsPoints(...) or exactlyPoints(...) instead:

MetricExpectation
  .sum[Long]("service.requests")
  .containsPoints(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "eu")),
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "us"))
  )

Failure reporting

The API is framework-agnostic, so it does not provide assertions directly. Instead, it returns structured mismatches and a formatter.

def assertExpected(metrics: List[MetricData], expected: MetricExpectation*): Unit =
  MetricExpectations.checkAll(metrics, expected: _*) match {
    case Right(_) =>
      ()
    case Left(mismatches) =>
      sys.error(MetricExpectations.format(mismatches))
  }

There are two main failure cases:

This is especially helpful when a metric name is correct but, for example, one scope attribute or one point attribute is wrong.

Top-level matching

Use MetricExpectations to match expectations against a collected list of exported metrics.

Available helpers:

checkAll(...) is non-consuming: the same exported metric may satisfy multiple expectations.

checkAllDistinct(...) enforces distinct assignment and is the safer default when repeated expectations should match different collected metrics.

Clues

Metric, point, and point-set expectations support optional clues.

MetricExpectation
  .sum[Long]("service.requests")
  .clue("the request counter must be emitted")
  .points(
    PointSetExpectation
      .exists(
        PointExpectation
          .numeric(1L)
          .clue("GET requests should increment the counter")
          .attributesExact(Attribute("http.method", "GET"))
      )
      .clue("the GET point must be present")
  )

Clues are included in mismatch messages to make failures easier to interpret.

Numeric comparison

Numeric expectations use NumberComparison[A].

This applies to:

If needed, you can override the default Double comparison implicitly in a test suite:

import org.typelevel.otel4s.oteljava.testkit.metrics.NumberComparison

implicit val cmp: NumberComparison[Double] =
  NumberComparison.within(1e-4)

After that, double-based point expectations built in that scope will use the custom comparison.