Technology

Running Dapr on Kubernetes in Production: Upgrades, Observability and Day-2 Operations

Getting Dapr running on Kubernetes takes one command. Keeping it healthy across dozens of services, several clusters and a new minor release every few months is a different job. Day one is about installing it. Day two is about upgrades that don’t cause outages, certificates that don’t expire at 3 a.m., sidecars sized correctly, and telemetry that tells you something is wrong before your users do.

This guide collects the operational practices that matter most once Dapr is carrying production traffic.

Know what you’re operating

A Dapr installation on Kubernetes has two parts.

The control plane, usually in the dapr-system namespace:

  • dapr-operator: manages component and configuration updates.
  • dapr-sidecar-injector: injects the daprd sidecar into annotated pods.
  • dapr-sentry: the certificate authority that issues workload identities for mTLS.
  • dapr-placement-server: tracks actor placement across instances, and underpins workflows.
  • dapr-scheduler-server: stores and triggers jobs, reminders and workflow timers.

The data plane: a daprd sidecar in every Dapr-enabled pod.

Most production incidents involving Dapr come from one of four places: a control plane that isn’t highly available, sidecars without resource limits, a version skew after an upgrade, or an expired root certificate. The sections below address each one.

1. Run the control plane in high-availability mode

By default, control-plane services run as a single replica, which is fine for development and a problem for production. Placement and scheduler rely on consensus among replicas, so a single instance means a single point of failure for actors and workflows.

Enable HA at install time with Helm:

helm upgrade –install dapr dapr/dapr \

  –namespace dapr-system –create-namespace \

  –set global.ha.enabled=true \

  –wait

This runs three replicas of each control-plane service. Also:

  • Spread replicas across availability zones with pod anti-affinity or topology spread constraints.
  • Give the scheduler reliable persistent storage, since it stores durable state for jobs and workflow timers.
  • Pin the Dapr version explicitly with –version and manage it through GitOps rather than ad-hoc CLI installs.

The official Dapr production guidelines have current resource recommendations for each control-plane component.

2. Size and tune your sidecars

Every Dapr-enabled pod gets a sidecar, so sidecar settings multiply across the cluster. Set requests and limits explicitly with annotations:

metadata:

  annotations:

    dapr.io/enabled: “true”

    dapr.io/app-id: “checkout”

    dapr.io/app-port: “8080”

    dapr.io/sidecar-cpu-request: “100m”

    dapr.io/sidecar-memory-request: “128Mi”

    dapr.io/sidecar-cpu-limit: “500m”

    dapr.io/sidecar-memory-limit: “512Mi”

    dapr.io/log-as-json: “true”

    dapr.io/enable-app-health-check: “true”

    dapr.io/app-health-check-path: “/healthz”

A few practical rules:

  • Start small and measure. Sidecar usage depends heavily on traffic and on which building blocks are in use. Review real usage after a week and adjust.
  • Enable app health checks so Dapr stops delivering messages to an app that is up but unhealthy.
  • Scope components to the app IDs that need them, with the scopes field on each component. It limits the blast radius and reduces the work each sidecar does at startup.
  • Log as JSON so your log pipeline can parse sidecar logs alongside application logs.

3. Upgrade without drama

Dapr ships minor releases regularly, and each release is supported for a limited window, so staying current isn’t optional. The safe sequence is always control plane first, then data plane:

  1. Read the release notes for breaking changes and deprecated component versions.
  2. Upgrade in a non-production cluster first and run your integration tests, especially for workflows and actors.
  3. Upgrade the control plane with Helm (or dapr upgrade -k –runtime-version <version>). Existing sidecars keep running and stay compatible with the new control plane within the supported skew.
  4. Roll the data plane by restarting deployments so they pick up the new sidecar version:

kubectl rollout restart deployment/checkout -n shop

  1. Verify with dapr status -k and by checking that sidecar versions match across pods.

Roll workloads gradually, starting with low-risk services, instead of restarting everything at once. Don’t skip minor versions in a single jump. The Kubernetes upgrade guide lists the supported paths.

4. Don’t let mTLS certificates expire

Dapr’s Sentry service issues short-lived workload certificates that rotate automatically. The root and issuer certificates that anchor the trust chain are long-lived, though, and if they expire, service-to-service communication fails across the cluster. This is one of the few Dapr failures that can take down everything at once.

Check expiry regularly:

dapr mtls expiry -k

Renew ahead of time:

dapr mtls renew-certificate -k –valid-until 365 –restart

Better still, turn this into an alert. Dapr exposes certificate expiry information, so fire a warning at 30 days before expiry and a critical alert at 7 days. If your organization has its own PKI, bring your own root certificates and put them on the same rotation process as the rest of your infrastructure.

5. Build real observability

Dapr gives you a lot of telemetry for free, but only if you collect it.

Metrics. Each sidecar and control-plane component exposes Prometheus metrics (by default on port 9090). Useful signals include:

  • HTTP and gRPC request latency and error rates for service invocation, per app ID
  • Pub/sub publish and delivery failures, and message processing latency
  • Component initialization failures, which often indicate a misconfigured secret or connection string
  • Workflow and activity execution counts, durations and failures
  • Sidecar CPU and memory against their limits

Tracing. Configure a Configuration resource to export traces over OpenTelemetry (OTLP) to your tracing backend. Because Dapr propagates W3C trace context through service invocation and pub/sub, you get end-to-end traces across services without instrumenting each hop yourself.

apiVersion: dapr.io/v1alpha1

kind: Configuration

metadata:

  name: tracing

spec:

  tracing:

    samplingRate: “0.1”

    otel:

      endpointAddress: “otel-collector.observability:4317”

      isSecure: false

      protocol: grpc

Alerting. Start with a short list of high-signal alerts: control-plane pods not ready, sidecar injection failures, component init errors, certificate expiry, sustained pub/sub delivery failures, and workflow failure rate. Add more only when an incident shows you need them.

6. Manage configuration drift across clusters

Once you have more than one cluster, drift sets in quietly: one cluster is a minor version behind, another has a component pointing at a deprecated version, a third has HA disabled because someone reinstalled it in a hurry. Treat Dapr’s control-plane settings, components, resiliency policies and configurations as code in Git, deployed through Argo CD or Flux, and review changes like any other infrastructure change.

Tooling that helps

You can build all of this yourself with Helm, Prometheus, Grafana and a well-maintained runbook, and many teams do. If you’d rather not, dedicated Dapr operations tools exist. Diagrid‘s Dapr Ops Dashboard, for example, is a free SaaS tool that automates control-plane upgrades and certificate rotation, checks clusters against a set of best-practice advisories, and provides prebuilt metrics dashboards and multi-cluster views. Teams running business-critical workloads may also want commercial Dapr support with defined response times for incidents and upgrade planning.

A day-2 checklist

  • ☐  Control plane runs in HA mode, spread across zones
  • ☐  Dapr version is pinned and managed through GitOps
  • ☐  Every Dapr-enabled workload has sidecar resource requests and limits
  • ☐  App health checks are enabled for services that consume messages
  • ☐  Components are scoped to the app IDs that need them
  • ☐  Upgrade runbook exists: non-prod first, control plane, then rolling data plane
  • ☐  Root certificate expiry is monitored and alerted on
  • ☐  Prometheus is scraping sidecar and control-plane metrics
  • ☐  Traces are exported through OpenTelemetry with a sensible sampling rate
  • ☐  Resiliency policies are defined for critical app-to-app calls

Conclusion

Dapr removes a lot of distributed-systems complexity from application code, but that complexity doesn’t disappear. Some of it moves to the platform. The teams that succeed with Dapr in production treat it like any other critical infrastructure: highly available by default, upgraded on a regular schedule, monitored with real alerts, and managed as code. Get those basics right and the sidecar fades into the background, which is where good infrastructure belongs.

Related posts
Technology

What Does a Factory Sound Like? How Engineers Listen to Modular Process Skids

Technology

A Beat Breakdown Needs Listening Time, Not More Narration

Technology

A Beat Breakdown Needs Listening Time, Not More Narration

Technology

How Smart Water Purification Technology Is Changing Drinking Water at Home

Leave a Reply