Site icon InstrumentalFx

Running Dapr on Kubernetes in Production: Upgrades, Observability and Day-2 Operations

Getting Dapr running on Kubernetes takes one command. Keeping it healthy across dozens of services, several clusters and a new minor release every few months is a different job. Day one is about installing it. Day two is about upgrades that don’t cause outages, certificates that don’t expire at 3 a.m., sidecars sized correctly, and telemetry that tells you something is wrong before your users do.

This guide collects the operational practices that matter most once Dapr is carrying production traffic.

Know what you’re operating

A Dapr installation on Kubernetes has two parts.

The control plane, usually in the dapr-system namespace:

The data plane: a daprd sidecar in every Dapr-enabled pod.

Most production incidents involving Dapr come from one of four places: a control plane that isn’t highly available, sidecars without resource limits, a version skew after an upgrade, or an expired root certificate. The sections below address each one.

1. Run the control plane in high-availability mode

By default, control-plane services run as a single replica, which is fine for development and a problem for production. Placement and scheduler rely on consensus among replicas, so a single instance means a single point of failure for actors and workflows.

Enable HA at install time with Helm:

helm upgrade –install dapr dapr/dapr \

  –namespace dapr-system –create-namespace \

  –set global.ha.enabled=true \

  –wait

This runs three replicas of each control-plane service. Also:

The official Dapr production guidelines have current resource recommendations for each control-plane component.

2. Size and tune your sidecars

Every Dapr-enabled pod gets a sidecar, so sidecar settings multiply across the cluster. Set requests and limits explicitly with annotations:

metadata:

  annotations:

    dapr.io/enabled: “true”

    dapr.io/app-id: “checkout”

    dapr.io/app-port: “8080”

    dapr.io/sidecar-cpu-request: “100m”

    dapr.io/sidecar-memory-request: “128Mi”

    dapr.io/sidecar-cpu-limit: “500m”

    dapr.io/sidecar-memory-limit: “512Mi”

    dapr.io/log-as-json: “true”

    dapr.io/enable-app-health-check: “true”

    dapr.io/app-health-check-path: “/healthz”

A few practical rules:

3. Upgrade without drama

Dapr ships minor releases regularly, and each release is supported for a limited window, so staying current isn’t optional. The safe sequence is always control plane first, then data plane:

  1. Read the release notes for breaking changes and deprecated component versions.
  2. Upgrade in a non-production cluster first and run your integration tests, especially for workflows and actors.
  3. Upgrade the control plane with Helm (or dapr upgrade -k –runtime-version <version>). Existing sidecars keep running and stay compatible with the new control plane within the supported skew.
  4. Roll the data plane by restarting deployments so they pick up the new sidecar version:

kubectl rollout restart deployment/checkout -n shop

  1. Verify with dapr status -k and by checking that sidecar versions match across pods.

Roll workloads gradually, starting with low-risk services, instead of restarting everything at once. Don’t skip minor versions in a single jump. The Kubernetes upgrade guide lists the supported paths.

4. Don’t let mTLS certificates expire

Dapr’s Sentry service issues short-lived workload certificates that rotate automatically. The root and issuer certificates that anchor the trust chain are long-lived, though, and if they expire, service-to-service communication fails across the cluster. This is one of the few Dapr failures that can take down everything at once.

Check expiry regularly:

dapr mtls expiry -k

Renew ahead of time:

dapr mtls renew-certificate -k –valid-until 365 –restart

Better still, turn this into an alert. Dapr exposes certificate expiry information, so fire a warning at 30 days before expiry and a critical alert at 7 days. If your organization has its own PKI, bring your own root certificates and put them on the same rotation process as the rest of your infrastructure.

5. Build real observability

Dapr gives you a lot of telemetry for free, but only if you collect it.

Metrics. Each sidecar and control-plane component exposes Prometheus metrics (by default on port 9090). Useful signals include:

Tracing. Configure a Configuration resource to export traces over OpenTelemetry (OTLP) to your tracing backend. Because Dapr propagates W3C trace context through service invocation and pub/sub, you get end-to-end traces across services without instrumenting each hop yourself.

apiVersion: dapr.io/v1alpha1

kind: Configuration

metadata:

  name: tracing

spec:

  tracing:

    samplingRate: “0.1”

    otel:

      endpointAddress: “otel-collector.observability:4317”

      isSecure: false

      protocol: grpc

Alerting. Start with a short list of high-signal alerts: control-plane pods not ready, sidecar injection failures, component init errors, certificate expiry, sustained pub/sub delivery failures, and workflow failure rate. Add more only when an incident shows you need them.

6. Manage configuration drift across clusters

Once you have more than one cluster, drift sets in quietly: one cluster is a minor version behind, another has a component pointing at a deprecated version, a third has HA disabled because someone reinstalled it in a hurry. Treat Dapr’s control-plane settings, components, resiliency policies and configurations as code in Git, deployed through Argo CD or Flux, and review changes like any other infrastructure change.

Tooling that helps

You can build all of this yourself with Helm, Prometheus, Grafana and a well-maintained runbook, and many teams do. If you’d rather not, dedicated Dapr operations tools exist. Diagrid‘s Dapr Ops Dashboard, for example, is a free SaaS tool that automates control-plane upgrades and certificate rotation, checks clusters against a set of best-practice advisories, and provides prebuilt metrics dashboards and multi-cluster views. Teams running business-critical workloads may also want commercial Dapr support with defined response times for incidents and upgrade planning.

A day-2 checklist

Conclusion

Dapr removes a lot of distributed-systems complexity from application code, but that complexity doesn’t disappear. Some of it moves to the platform. The teams that succeed with Dapr in production treat it like any other critical infrastructure: highly available by default, upgraded on a regular schedule, monitored with real alerts, and managed as code. Get those basics right and the sidecar fades into the background, which is where good infrastructure belongs.

Exit mobile version