Skip to content

Metrics

Ductwork Pro can report what it’s doing to a StatsD client. It emits counters when pipelines and jobs start and finish, and timers for how long they took. That’s usually enough to build a dashboard and a couple of alerts without instrumenting your steps yourself.

Hand Ductwork a lambda that builds your StatsD client. It’s called once, lazily, the first time a metric is emitted:

config/initializers/ductwork.rb
require "datadog/statsd"
Ductwork.statsd = -> { Datadog::Statsd.new("localhost", 8125) }

Setting Ductwork.statsd to anything other than a Proc raises an ArgumentError. If you never set it, metrics go to a null client and nothing is sent. Put this in an initializer so it’s in place before any Ductwork process starts doing work, since the client is memoized the first time it’s used.

Ductwork calls increment, timing, and time on the client and passes tags as an array of "key:value" strings via a tags: keyword argument. That’s the dogstatsd-ruby interface. Other clients work if they match it. Clients that don’t take a tags: keyword won’t work without a wrapper.

Every metric gets a service:ductwork tag plus the per-metric tags listed below. If you want a prefix on the metric names or extra tags on everything, set them on the client:

Ductwork.statsd = -> {
Datadog::Statsd.new(
"localhost",
8125,
namespace: "ductwork",
tags: ["env:#{Rails.env}"]
)
}

That turns pipeline.triggered into ductwork.pipeline.triggered and tags everything with the environment.

MetricTypeTagsWhen
pipeline.triggeredcountpipeline.trigger was called
pipeline.completedcountpipelineevery branch finished and the run completed
pipeline.haltedcountpipelinethe run halted, for any reason
pipeline.runtimetiming (ms)pipelinethe run completed. Measured from the run’s start time (trigger time plus any start delay) to completion. Not emitted for halted runs
pipeline.dampenedcountpipelinea dampen transition paused the run
pipeline.resumedcountpipelineresume! was called on a dampened run
job.enqueuedcountjoba job was created for a step
job.completedcountklass, resulta job execution finished, for any reason
job.runtimetiming (ms)jobwraps the step’s execute call
thread.undeadcountrole, reasona killed worker thread refused to die (see Step Timeout)

The pipeline tag is the pipeline class name. job and klass are both the step class name (yes, job.enqueued and job.completed use different tag keys for the same thing). result on job.completed is one of success, failure, timed_out, or crashed. The last one means the thread died before a result was recorded.

A few alerts worth setting up:

  • pipeline.halted above zero, or above whatever baseline you’re comfortable with. A halt means retries were exhausted and somebody needs to look.
  • job.enqueued pulling away from job.completed over a window. That’s your workers falling behind.
  • p95 of job.runtime by the job tag creeping up. Usually an external dependency getting slow.
  • Any thread.undead at all.