Testkit | Metrics

The metrics testkit provides a partial-matching expectation API for SDK MetricData.

This is useful in tests because metric data contains much more than the values you usually care about: instrumentation scope, telemetry resource, point attributes, collection windows, exemplars, histogram buckets, and more.

The expectation API lets you assert only the relevant parts of a metric while still preserving detail when you need it. Metric point matching is expressed at the collection level, so a single metric expectation can accumulate multiple point constraints.

Getting started

Add settings to the build.sbt:

libraryDependencies += "org.typelevel" %%% "otel4s-sdk-testkit" % "0.19.3" % Test

Add directives to the *.scala file:

//> using test.dep "org.typelevel::otel4s-sdk-testkit::0.19.3"

Basic flow

The usual flow is:

  1. Run your program against OpenTelemetrySdkTestkit
  2. Collect metrics as MetricData
  3. Build MetricExpectation values
  4. Check them with MetricExpectations.checkAll
import cats.effect.IO
import org.typelevel.otel4s.metrics.MeterProvider
import org.typelevel.otel4s.sdk.metrics.data.MetricData
import org.typelevel.otel4s.sdk.testkit.OpenTelemetrySdkTestkit
import org.typelevel.otel4s.sdk.testkit.metrics._

def program(meterProvider: MeterProvider[IO]): IO[Unit] =
  for {
    meter <- meterProvider.get("service")
    counter <- meter.counter[Long]("service.counter").create
    _ <- counter.inc()
    gauge <- meter.gauge[Long]("service.gauge").create
    _ <- gauge.record(42L)
  } yield ()

def assertExpected(metrics: List[MetricData], expected: MetricExpectation*): Unit =
  MetricExpectations.checkAll(metrics, expected.toList) match {
    case Right(_) =>
      ()
    case Left(mismatches) =>
      sys.error(MetricExpectations.format(mismatches))
  }

def test: IO[Unit] =
  OpenTelemetrySdkTestkit.inMemory[IO]().use { testkit =>
    for {
      _ <- program(testkit.meterProvider)
      metrics <- testkit.collectMetrics
    } yield assertExpected(
      metrics,
      MetricExpectation.sum[Long]("service.counter").value(1L),
      MetricExpectation.gauge[Long]("service.gauge").value(42L)
    )
  }

checkAll(...) is non-consuming: each expectation is checked independently against the full collected metric list. If you need to ensure that repeated expectations match different collected metrics, use MetricExpectations.checkAllDistinct(...) instead.

def checkAllDistinct(metrics: List[MetricData]) =
  MetricExpectations.checkAllDistinct(
    metrics,
    MetricExpectation.sum[Long]("service.counter").value(1L),
    MetricExpectation.sum[Long]("service.counter").value(1L)
  )

Partial matching

All expectations are partial.

This means:

At the metric level, partial matching is also non-consuming by default: repeating the same MetricExpectation in checkAll(...) does not require two distinct collected metrics.

For example:

MetricExpectation.name("service.counter")

matches any collected metric named service.counter, regardless of its type, points, scope, or resource.

MetricExpectation.sum[Long]("service.counter").value(1L)

matches a long sum metric named service.counter that has at least one point with value 1L.

Numeric metrics

Use MetricExpectation.gauge[A] and MetricExpectation.sum[A] for numeric metrics.

MetricExpectation.gauge[Double]("service.temperature")
MetricExpectation.sum[Long]("service.requests")

The measurement type remains generic. Long and Double use the same API.

For value-only checks, value(...) is the most compact form:

MetricExpectation.sum[Long]("service.requests").value(1L)

If you also care about point attributes, there is a shorthand for exact point matching:

import org.typelevel.otel4s.Attribute

MetricExpectation
  .sum[Long]("service.requests")
  .value(1L, Attribute("http.method", "GET"))

value(value, attributes) uses exact attribute matching. It is equivalent to:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesExact(Attribute("http.method", "GET"))
    )
  )

Use points(PointSetExpectation.exists(...)) directly when you need subset matching or a more detailed point expectation.

Point attributes

Point attributes can be matched in three ways:

Both attributesExact(...) and attributesSubset(...) support Attributes and varargs of Attribute[_].

import org.typelevel.otel4s.Attributes
import org.typelevel.otel4s.sdk.testkit.AttributesExpectation

PointExpectation
  .numeric(1L)
  .attributesExact(Attribute("http.method", "GET"))

PointExpectation
  .numeric(1L)
  .attributesSubset(
    Attribute("http.method", "GET"),
    Attribute("http.route", "/users")
  )

PointExpectation
  .numeric(1L)
  .attributes(
    AttributesExpectation.where("expected only one attribute") { attributes =>
      attributes == Attributes(Attribute("http.method", "GET"))
    }
  )

Empty attributes can be matched with:

PointExpectation.numeric(1L).attributesEmpty

Scope and resource expectations

Metric expectations can also assert the instrumentation scope and telemetry resource.

This is useful when you want tests to preserve the same level of detail as raw MetricData.

import org.typelevel.otel4s.sdk.testkit.{InstrumentationScopeExpectation, TelemetryResourceExpectation}

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesExact(Attribute("http.method", "GET"))
    )
  )
  .scope(
    InstrumentationScopeExpectation
      .name("service")
      .version("1.0")
      .schemaUrl("https://opentelemetry.io/schemas/1.24.0")
      .attributesEmpty
  )
  .resource(
    TelemetryResourceExpectation.any
      .attributesSubset(Attribute("service.name", "auth-service"))
      .schemaUrl(None)
  )

As with point attributes, scope and resource attribute helpers support both Attributes and varargs. They also provide attributesEmpty when you want to assert that no attributes are present.

Histograms

The testkit also supports histogram metrics directly.

import org.typelevel.otel4s.metrics.BucketBoundaries

MetricExpectation
  .histogram("http.server.duration")
  .points(
    PointSetExpectation.exists(
      PointExpectation.histogram
        .count(3L)
        .sum(42.0)
        .boundaries(BucketBoundaries(0.1, 1.0, 10.0))
        .counts(1L, 1L, 1L, 0L)
    )
  )

Exemplars

Numeric and histogram point expectations can also assert exemplars.

PointExpectation
  .numeric(1L)
  .exemplarCount(1)
  .containsExemplars(
    ExemplarExpectation
      .numeric(1L)
      .filteredAttributesSubset(Attribute("sampled", true))
  )

PointExpectation
  .histogram
  .exemplarCount(1)
  .containsExemplars(
    ExemplarExpectation
      .numeric(42.0)
      .withoutTraceContext
  )

As with points, exemplar expectations are partial. You can assert only the value, only filtered attributes, only trace context, or any combination of them.

Point-set expectations

Point matching is collection-based. A MetricExpectation does not check points one by one in isolation. Instead, it evaluates a PointSetExpectation against the full point collection, which lets you express presence, absence, cardinality, and collection-wide invariants.

All typed metric expectations expose:

exists

Use exists when at least one collected point must match:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesExact(Attribute("http.method", "GET"))
    )
  )

This is the default mental model for helpers like value(...): they require at least one matching point, not that every point looks the same.

forall

Use forall when every collected point must satisfy the same rule:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.forall(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("kind", "ok"))
    )
  )

forall is useful for invariants such as "every point has the same value shape" or "every point includes this attribute subset". Unlike plain universal quantification over a Scala collection, it fails on an empty point set.

contains and exactly

Use contains when the metric must contain several distinct matching points, but extra points are still allowed:

MetricExpectation
  .sum[Long]("service.requests")
  .containsPoints(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "eu")),
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "us"))
  )

contains enforces distinct matching. If you ask for the same expected point twice, the collected data must contain two matching points.

Use exactly when the metric must contain exactly the expected points and no extras:

MetricExpectation
  .sum[Long]("service.requests")
  .exactlyPoints(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "eu")),
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "us"))
  )

Cardinality combinators

Use pointCount(...) when only the total point count matters:

MetricExpectation
  .sum[Long]("service.requests")
  .pointCount(2)

For lower-level count constraints, use PointSetExpectation directly:

import org.typelevel.otel4s.sdk.metrics.data.PointData

MetricExpectation
  .sum[Long]("service.requests")
  .points(PointSetExpectation.minCount[PointData.NumberPoint.Aux[Long]](1))
  .points(PointSetExpectation.maxCount[PointData.NumberPoint.Aux[Long]](2))

Use countWhere when the total number of matching points matters more than the exact full set:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.countWhere(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("region", "eu")),
      expected = 2
    )
  )

This is useful when each point still carries extra dimensions that you do not want to enumerate explicitly.

none

Use none or withoutPointsMatching(...) when a point shape must not appear:

MetricExpectation
  .sum[Long]("service.requests")
  .withoutPointsMatching(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "test"))
  )

predicate

Use predicate or pointsWhere(...) for collection-wide assertions that are awkward to express with the built-in combinators:

MetricExpectation
  .sum[Long]("service.requests")
  .pointsWhere("expected exactly EU and US points") { points =>
    points.map(_.attributes).toSet == Set(
      Attributes(Attribute("region", "eu")),
      Attributes(Attribute("region", "us"))
    )
  }

For numeric metrics, the points passed to pointsWhere are typed numeric points, so you can inspect .value, .attributes, and exemplar data directly.

and and or

Point-set expectations can be combined explicitly:

import org.typelevel.otel4s.sdk.metrics.data.PointData

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation
      .contains(
        PointExpectation
          .numeric(1L)
          .attributesSubset(Attribute("region", "eu")),
        PointExpectation
          .numeric(1L)
          .attributesSubset(Attribute("region", "us"))
      )
      .and(PointSetExpectation.count[PointData.NumberPoint.Aux[Long]](2))
  )

Use or when a metric may legitimately satisfy one of several point layouts:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation
      .count[PointData.NumberPoint.Aux[Long]](1)
      .or(PointSetExpectation.count[PointData.NumberPoint.Aux[Long]](2))
  )

Because point expectations accumulate, you can also layer multiple points(...) calls on the same metric:

MetricExpectation
  .sum[Long]("service.requests")
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("region", "eu"))
    )
  )
  .points(
    PointSetExpectation.exists(
      PointExpectation
        .numeric(1L)
        .attributesSubset(Attribute("region", "us"))
    )
  )

Each points(...) clause is checked independently against the full collected point set. Chaining points(PointSetExpectation.exists(...)) does not reserve matched points for later clauses, so it should not be used to require multiple distinct points.

If you need distinct matching, use containsPoints(...) or exactlyPoints(...) instead:

MetricExpectation
  .sum[Long]("service.requests")
  .containsPoints(
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "eu")),
    PointExpectation
      .numeric(1L)
      .attributesSubset(Attribute("region", "us"))
  )

Failure reporting

The API is framework-agnostic, so it does not provide assertions directly. Instead, it returns structured mismatches and a formatter.

def assertExpectedFormatted(metrics: List[MetricData], expected: MetricExpectation*): Unit =
  MetricExpectations.checkAll(metrics, expected.toList) match {
    case Right(_) =>
      ()
    case Left(mismatches) =>
      sys.error(MetricExpectations.format(mismatches))
  }

There are three main failure cases:

This is especially helpful when a metric name is correct but, for example, one scope attribute or one point attribute is wrong.

Clues

Metric, point, point-set, and exemplar expectations support optional clues.

MetricExpectation
  .sum[Long]("service.requests")
  .clue("the request counter must be emitted")
  .points(
    PointSetExpectation
      .exists(
        PointExpectation
          .numeric(1L)
          .clue("GET requests should increment the counter")
          .attributesExact(Attribute("http.method", "GET"))
      )
      .clue("the GET point must be present")
  )

Clues are included in mismatch messages to make failures easier to interpret.

Numeric comparison

Numeric expectations use NumberComparison[A].

This applies to:

If needed, you can override the default Double comparison implicitly in a test suite:

locally {
  implicit val doubleComparison: NumberComparison[Double] =
    NumberComparison.within(1e-4)

  MetricExpectation.gauge[Double]("service.temperature")
}

After that, double-based metric, point, and exemplar expectations built in that scope will use the custom comparison.