Metrics
Ductwork Pro can report what it’s doing to a StatsD client. It emits counters when pipelines and jobs start and finish, and timers for how long they took. That’s usually enough to build a dashboard and a couple of alerts without instrumenting your steps yourself.
Configuration
Section titled “Configuration”Hand Ductwork a lambda that builds your StatsD client. It’s called once, lazily, the first time a metric is emitted:
require "datadog/statsd"
Ductwork.statsd = -> { Datadog::Statsd.new("localhost", 8125) }Setting Ductwork.statsd to anything other than a Proc raises an ArgumentError. If you never set it, metrics go to a null client and nothing is sent. Put this in an initializer so it’s in place before any Ductwork process starts doing work, since the client is memoized the first time it’s used.
Ductwork calls increment, timing, and time on the client and passes tags as an array of "key:value" strings via a tags: keyword argument. That’s the dogstatsd-ruby interface. Other clients work if they match it. Clients that don’t take a tags: keyword won’t work without a wrapper.
Tagging and namespacing
Section titled “Tagging and namespacing”Every metric gets a service:ductwork tag plus the per-metric tags listed below. If you want a prefix on the metric names or extra tags on everything, set them on the client:
Ductwork.statsd = -> { Datadog::Statsd.new( "localhost", 8125, namespace: "ductwork", tags: ["env:#{Rails.env}"] )}That turns pipeline.triggered into ductwork.pipeline.triggered and tags everything with the environment.
Available metrics
Section titled “Available metrics”| Metric | Type | Tags | When |
|---|---|---|---|
pipeline.triggered | count | pipeline | .trigger was called |
pipeline.completed | count | pipeline | every branch finished and the run completed |
pipeline.halted | count | pipeline | the run halted, for any reason |
pipeline.runtime | timing (ms) | pipeline | the run completed. Measured from the run’s start time (trigger time plus any start delay) to completion. Not emitted for halted runs |
pipeline.dampened | count | pipeline | a dampen transition paused the run |
pipeline.resumed | count | pipeline | resume! was called on a dampened run |
job.enqueued | count | job | a job was created for a step |
job.completed | count | klass, result | a job execution finished, for any reason |
job.runtime | timing (ms) | job | wraps the step’s execute call |
thread.undead | count | role, reason | a killed worker thread refused to die (see Step Timeout) |
The pipeline tag is the pipeline class name. job and klass are both the step class name (yes, job.enqueued and job.completed use different tag keys for the same thing). result on job.completed is one of success, failure, timed_out, or crashed. The last one means the thread died before a result was recorded.
What to watch
Section titled “What to watch”A few alerts worth setting up:
pipeline.haltedabove zero, or above whatever baseline you’re comfortable with. A halt means retries were exhausted and somebody needs to look.job.enqueuedpulling away fromjob.completedover a window. That’s your workers falling behind.- p95 of
job.runtimeby thejobtag creeping up. Usually an external dependency getting slow. - Any
thread.undeadat all.