Bringing Datadog Observability to Our Lambda + ECS Stack

What is Datadog?

Datadog is a cloud monitoring/observability platform that unifies APM (distributed tracing), log management, and infrastructure metrics in one place. Instead of digging through raw CloudWatch logs to figure out why a request was slow or which service in a call chain failed, you get traces correlated with logs and metrics, searchable and dashboarded.

Why not just use CloudWatch?

We already had CloudWatch, every Lambda and ECS task ships logs and metrics there by default, for free, with zero extra setup. So it’s a fair question. A few concrete gaps pushed us past it:

  • No distributed tracing. CloudWatch (without layering AWS X-Ray on top, and even then only partially) doesn’t give you a single trace that follows a request across a Lambda call, into an ECS service, through Postgres, and back.
    Our stack spans both compute models - Lambda for the API layer, ECS Fargate for longer-running services, so a request can legitimately hop between them. CloudWatch shows you each hop’s logs separately; it doesn’t stitch them into one flame graph.
  • Logs and traces aren’t correlated automatically. With DD_LOGS_INJECTION=true, dd-trace stamps every log line with the active trace/span ID, so you can jump from a slow trace straight to the exact log lines it produced. Getting that out of plain CloudWatch Logs means hand-rolling correlation IDs everywhere and still not getting the visual trace waterfall.
  • CloudWatch Logs Insights doesn’t scale well as an investigation tool. It’s fine for a quick grep, but there’s no persistent dashboarding, no APM-style service map, and cross-log-group queries get slow and awkward once you’re spanning nine Lambda services plus three ECS services.
  • CloudWatch is AWS-native; Datadog isn’t tied to one environment. CloudWatch only really monitors what’s running inside your AWS account. Datadog’s Agent can run on-prem, on other clouds, or on AWS, and reports into the same org either way, so if we (or a client) ever need to bridge cloud and on-premises infrastructure, or go multi-cloud, we’re not stuck re-tooling observability from scratch. Not a gap we hit in this stack specifically (we’re all-AWS today), but part of why we picked something that isn’t AWS-only.
  • One pane of glass across two compute models. Lambda and ECS Fargate have different metrics, different log formats, different failure modes. Datadog gives us one service catalog and one alerting surface regardless of which one a given service runs on CloudWatch’s metrics/alarms are per-resource-type and don’t unify that view.
  • Datadog’s APM is the reason dd-trace auto-instruments express/pg/knex/http just by being imported first, we got tracing across the whole request path without manually instrumenting every route or query.

To be fair, this isn’t “CloudWatch is bad”, it’s still where the raw logs live and where the Datadog Forwarder Lambda actually reads from (see below). Datadog sits on top of it: CloudWatch is the source, Datadog is where we correlate, trace, and alert on what CloudWatch collected.

What we worked on

Our backend runs on two different compute models - Lambda and ECS Fargate, so instrumenting Datadog actually meant two separate integrations, plus a log pipeline to get everything there:

1. Lambda APM via serverless-plugin-datadog
Added the plugin across several Lambda-based services. Each service’s serverless.yml now has:

datadog:
    enabled: ${self:custom.datadogEnabledStages.${self:provider.stage}, false}
    apiKeySsmArn: arn:aws:ssm:${aws:region}:${aws:accountId}:parameter/DD_API_KEY
    site: datadoghq.eu
    env: ${self:provider.stage}
    service: ${self:service}

We also added explicit ssm:GetParameter + kms:Decrypt IAM permissions for DD_API_KEY, since we couldn’t confirm the plugin injects that permission automatically for us.

2. ECS Fargate APM via a Datadog Agent sidecar
For a few ECS-hosted services, each task definition now optionally runs a second datadog-agent container alongside the app container, sharing a Unix Domain Socket volume (dd-sockets) so dd-trace in the app can ship spans to the agent without going over the network stack. In each service’s entrypoint:

// Must run before anything else — dd-trace patches Node's module loader
import tracer from 'dd-trace'
tracer.init()

3. Log forwarding
Deployed Datadog’s officially hosted Forwarder Lambda (from their CloudFormation template, not hand-rolled) and wired up a CloudWatch Logs subscription filter per ECS service, so both the app container’s and the agent container’s log streams ship to Datadog automatically, this is exactly where CloudWatch still does the collecting, and Datadog just gets a copy.

All of this is gated to preprod/prod only — dev/qa run with datadog_enabled = false, which makes the whole thing a true no-op (no sidecar, no volume, no IAM, no plugin behavior) rather than a half-configured integration.

What we found out along the way

A few things that weren’t obvious going in:

  • Import order matters for tracing. dd-trace patches Node’s module loader to auto-instrument express/pg/knex/http. Anything already require()'d before tracer.init() runs — even transitively — won’t get traced. It has to be the very first import.

  • The Forwarder needs a full ARN, not a bare SSM parameter name. Its CloudFormation template’s DdApiKeySsmParameterName only accepts values starting with / or a full ARN — our parameter is named flatly (DD_API_KEY), so we had to pass the ARN explicitly.

  • Fargate needs platform version ≥ 1.4.0 for the ephemeral shared volume the two containers use to talk over UDS instead of the network.

  • Two different IAM roles need the secret permission. The ECS execution role (which launches the task and resolves secrets/valueFrom) needs ssm:GetParameters + kms:Decrypt for DD_API_KEY — separate from the task role the app itself runs as.

  • The agent container is marked non-essential. If the Datadog agent hiccups, it shouldn’t take the whole ECS task down with it — dd-trace buffers and retries on its own if the agent isn’t ready.

  • concat(), not a ternary, for the container list. Terraform’s conditional operator requires both branches of a ternary to have identical tuple types, which a 2-container vs 1-container list of differently-shaped objects violates.
    Ex: “Before / After”:

    Before (errors out):

    container_definitions_list = var.datadog_enabled ?
      [local.app_container, local.datadog_agent_container] :
      [local.app_container]
    
    

    After (works):

    container_definitions_list = concat(
      [local.app_container],
      var.datadog_enabled ? [local.datadog_agent_container] : []
    )