Self-hosted Plane on Kubernetes: A production guide

Learn how to run self-hosted Plane on Kubernetes in production, from Helm configuration and ingress to scaling, security, backups, monitoring, and upgrades.

Manish Gupta and Akshat Jain
●
9 Oct, 2026
Deploy Plane on Kubernetes

If your team already runs Kubernetes, you can self-host Plane Commercial Edition with the official plane-enterprise Helm chart. One Helm release installs the web apps, API, live service, workers, database migrations, and any optional dependencies you turn on.

A first install can look healthy: Helm completes, the pods turn green, and Plane opens in the browser. Then someone’s editor disconnects after sixty seconds. You check the ingress timeout. A beat setting has no effect, so you check the chart’s value names. During a restore, the database comes back, but Plane needs the original SECRET_KEY to read its saved configuration. Kubernetes can report healthy pods through all of this while people struggle to use the application.

This guide is for the platform engineers and SREs who run Plane after Helm finishes. It covers how the chart behaves: value names that fail silently and defaults that affect security and reliability, and the checks a readiness probe will never make for you.

TL;DR

Plane’s official Helm chart is a practical option if your team already operates Kubernetes. Before running it in production, work through what happens after installation:

  • Platform ownership- Your team supplies ingress, certificates, storage, secrets, monitoring, and backups. An existing Kubernetes operating model helps.
  • Availability - Bundled data stores are a starting point for evaluation. In production, choose database, cache, broker, and object storage arrangements that meet your uptime and recovery requirements.
  • Capacity - You can scale API and worker workloads independently. Use database connection counts, queue depth, and workload tests to decide when to add capacity.
  • User experience - Test sign-in, collaborative editing, and uploads through the full traffic path before opening the instance to your team.
  • Recovery and upgrades - Your recovery plan needs the database, attachments, application keys, and deployment configuration. Test a restore, and take a fresh recovery point before each upgrade.

The sections below cover the chart settings and checks involved.

Before you install

If you have read the Docker production guide, the shape will be familiar. Plane still depends on PostgreSQL, Valkey, RabbitMQ, and object storage, and a recoverable instance still needs the database, the bucket, the application keys, and the configuration that ties them together.

What Kubernetes changes is where those decisions hide. A Compose file puts the whole topology in front of you; Helm compresses the same choices into one command and a values file.

What the cluster must provide

The chart installs Plane, and nothing underneath it. You bring the ingress controller, storage classes, certificates, secret delivery, image access, observability, and backups.

  • Storage. Every bundled data-store PVC is ReadWriteOnce, so on zonal storage each volume pins its pod to one availability zone. Use a default StorageClass with WaitForFirstConsumer so the scheduler picks the node before the disk is provisioned in the wrong zone.
  • Ingress. Run more than one ingress-controller replica if the install should survive a node failure, and decide who issues and renews certificates.
  • Images. Every node that may run Plane must be able to pull from artifacts.plane.so, or you mirror the images into a registry you operate.
  • Secrets. The chart can reference a Secret by name, but moving values out of Vault, AWS Secrets Manager, Google Secret Manager, or Azure Key Vault is your platform's job.

A small internal deployment can reasonably live in one zone with bundled stores. Write that trade-off down in the runbook.

Render the chart before you install or upgrade

Helm accepts unknown keys. A value copied from another chart, or a key with a typo, renders successfully and is silently ignored by the plane-enterprise templates.

Before the first install and before every upgrade, run helm template and inspect the actual rendered field you meant to change: environment variables, replica counts, service URLs, Secret references. Valid YAML is not proof that your override was consumed.

helm template plane plane/plane-enterprise --version <chart-version> \
-n plane -f values.yaml > rendered.yaml
# Confirm the overrides actually landed
grep -n "replicas:" rendered.yaml
grep -n "kind: Deployment" -A 3 rendered.yaml | grep "name:"

Check Artifact Hub for the current chart version, and read the README and values file that match it. All workload and dependency settings live under services.*. The commands in this guide assume a release named plane in namespaceplane; adjust to yours.

Names that don't match

A default install renders the web, space, admin, API, live, worker, beatworker, and migration workloads, plus optional search, automation, email, integration, observability, and AI components. Two structural details trip people up:

  • There is no Caddy pod. The chart renders an Ingress or a Traefik IngressRoute; the controller serving it belongs to your cluster.
  • Services are headless (clusterIP: None), so a mesh, policy, or monitoring integration that expects a normal ClusterIP will miss them. Deployments and StatefulSets also carry a -wl suffix that their Services drop.

Several names differ between the values file, the rendered objects, and the pod environment:

In the cluster
In values.yaml
Note

Deployment beat-worker

services.beatworker

services.beat-worker is accepted and ignored.

StatefulSet pgdb

services.postgres

services.pgdb.local_setup=false leaves bundled PostgreSQL running.

Env var AWS_S3_BUCKET_NAME

env.docstore_bucket

Set the bucket in values, then check this variable in the API pod to confirm which bucket Plane is using

Valkey pods

services.redis

The key still uses the Redis name.

Env var AMQP_URL

services.rabbitmq.external_rabbitmq_url

This is the broker URL. The chart sets no Celery result backend.

Where your data should live

The bundled data stores are great for evaluation. They are single pods on small ReadWriteOnce volumes, with none of the replication or backup controls production usually expects.

Store
Disable bundled
External setting
Production concern

PostgreSQL

services.postgres.local_setup: false

env.pgdb_remote_url

Bundled is single-node. Use a multi-AZ service with PITR.

Valkey

services.redis.local_setup: false

env.remote_redis_url

One pod, no replica group, no failover.

RabbitMQ

services.rabbitmq.local_setup: false

services.rabbitmq.external_rabbitmq_url

Note the URL stays under services.rabbitmq. A single-node broker is not HA.

Object storage

services.minio.local_setup: false

env.aws_* and env.docstore_bucket

Needs independent durability and a deliberate access policy.

OpenSearch

services.opensearch.local_setup: false

env.opensearch_remote_url + credentials

Use multi-AZ with three master-eligible nodes.

local_setup: false only removes the bundled service. The external one still needs the failover, replication, backups, and capacity your recovery plan assumes. A managed database can snapshot and fail over, but you still choose retention, test the restore, and watch connection pressure.

⚠ Workload identity and empty keys - If pods get storage access through workload identity, leave env.aws_access_key and env.aws_secret_access_key out of your values file. Setting them to "" counts as a credential. The AWS SDK tries that empty key first and never falls back to the pod's identity, so uploads fail.

Here is the smallest useful values skeleton for external services:

services:
postgres:
local_setup: false
redis:
local_setup: false
minio:
local_setup: false
rabbitmq:
local_setup: false
external_rabbitmq_url: "amqp://<user>:<password>@<host>:5672/<vhost>"
opensearch:
local_setup: false

env:
pgdb_remote_url: "postgresql://<user>:<password>@<host>:5432/<database>"
remote_redis_url: "redis://<host>:6379/0"
opensearch_remote_url: "https://<opensearch-host>:9200"
opensearch_remote_username: "<username>"
opensearch_remote_password: "<password>"
aws_region: "<region>"
aws_s3_endpoint_url: "https://s3.<region>.amazonaws.com"
docstore_bucket: "<bucket>"
requireExplicitSecrets: true
# Omit aws_access_key and aws_secret_access_key when using workload identity.

external_secrets:
app_keys_existingSecret: "plane-app-keys"

For a highly available deployment, pair this guide with the Commercial Edition high-availability reference architecture, which covers workload tiers, managed services, zone spreading, and the supported failure model.

Ingress and TLS

Ingress failures are rarely total. The UI loads but collaborative editing drops. Small uploads work while large ones return 413. Password login succeeds and OAuth fails. Debug the whole path from browser to pod.

Commercial Edition takes the application host from license.licenseDomain. If it is empty, the chart renders no application ingress at all.

Pick the controller explicitly

ingress.controller selects the template: traefik renders an IngressRoute, openshift renders Routes, and nginx renders a standard Ingress. ingress.ingressClass supplies the class name. An empty controller defaults to Traefik, so set both explicitly if you run anything else, and confirm the right controller admitted the resource.

The main routes are / to web, /api and /auth to API, /live/ to live, /spaces to space, and /god-mode to admin, plus paths for any enabled Commercial services. If the instance is public, protect /god-mode at the edge.

How TLS settings affect WEB_URL

Plane builds WEB_URL from the configured host and TLS settings, and both OAuth redirects and bundled object-storage URLs depend on it. If the outside world speaks HTTPS but Plane thinks it's on HTTP, sign-in and attachments can break even though every service is healthy.

Where TLS terminates
Main value
Operational note

Existing Kubernetes Secret

ssl.tls_secret_name

The chart references it; your platform renews it.

cert-manager

ssl.createIssuer and ssl.generateCerts

Both are required. Creates a namespaced Issuer, not a ClusterIssuer.

Load balancer or CDN

ssl.externalTermination: true

No tls: block is rendered, but Plane should still build an HTTPS WEB_URL.

Traefik entrypoint

ingress.traefik.entryPoints

A wrong entrypoint shows up as a routing error.

Pick one owner for certificate renewal and monitor expiry there. None of these options proves that HTTP redirects to HTTPS, so test that at whichever component owns it.

⚠ Verify secure cookies directly. The browser's lock icon tells you nothing about SESSION_COOKIE_SECURE and CSRF_COOKIE_SECURE. On a standard TLS install these can still end up false, because the derived CORS_ALLOWED_ORIGINS includes an HTTP origin. Read both settings from a running API pod.

kubectl exec -n plane deploy/plane-api-wl -- python manage.py shell -c \
"from django.conf import settings; print(settings.SESSION_COOKIE_SECURE, settings.CSRF_COOKIE_SECURE)"
# Expect: True True

Timeouts and upload limits

The collaborative editor holds a WebSocket open under /live/. ingress-nginx defaults proxy-read-timeout and proxy-send-timeout to 60 seconds, so a quiet editing session dies at about the one-minute mark. A cloud load balancer with a 60-second idle timeout causes the same symptom one hop earlier.

Uploads cross three limits: the load balancer or CDN, the ingress proxy, and env.doc_upload_size_limit. Set them from largest to smallest, so that Plane rejects an oversized file with a clear error before a proxy does it with a bare 413.

ingress:
enabled: true
controller: nginx
ingressClass: nginx
ingress_annotations:
nginx.ingress.kubernetes.io/proxy-body-size: "20m"
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"

ssl:
externalTermination: true

license:
licenseDomain: plane.example.com

env:
doc_upload_size_limit: "20971520"

Four checks before real traffic: keep an editor open past the shortest idle timeout, upload a file near your limit, complete an OAuth round trip, and confirm the secure-cookie flags. Together they cover far more of the ingress path than a 200 from /.

Scaling and capacity

Every application workload starts at one replica. The chart can render PodDisruptionBudgets and topology-spread constraints, but with a single replica there is nothing for them to protect.

API: database connections set the limit

Plane uses Redis/Valkey for caching and user-session storage, as described in its architecture documentation. Point every API replica at the same shared services, and watch PostgreSQL connections as you add replicas. Concurrent requests and background tasks both use database capacity. Evaluate connection pooling with PgBouncer or RDS Proxy before raising concurrency substantially.

Workers: set concurrency yourself

Workers compete for messages from RabbitMQ. Plane starts Celery without an explicit concurrency setting, so the prefork pool sizes itself from os.cpu_count(), which can report the node's cores rather than the container's CPU limit. A pod limited to two CPUs can fork as if it owned a 32-core node.

Set worker concurrency in your deployment overlay, or place workers on nodes whose core count matches the pool you want. Watch RabbitMQ backlog (messages_ready) when scaling replicas. It can show workers falling behind on I/O-heavy tasks even when CPU use stays low.

# Queue backlog on the bundled broker
kubectl exec -n plane plane-rabbitmq-wl-0 -- \
rabbitmqctl list_queues name messages_ready messages_unacknowledged consumers

Beat: run exactly one replica

⚠ Keep services.beatworker.replicas: 1. Plane uses a database-backed Celery beat scheduler with no leader election, so a second beat pod schedules every periodic job twice. The key has no hyphen; --set services.beat-worker.replicas=1 is a silent no-op.

services:
beatworker: # no hyphen
replicas: 1

# Verify on the cluster: should print 1
kubectl get deploy -n plane plane-beat-worker-wl -o jsonpath='{.spec.replicas}'

Web, space, admin, API, live, and worker can all run multiple replicas. The bundled PostgreSQL, Valkey, RabbitMQ, and MinIO cannot be made highly available by raising a replica count; move them to an HA service instead.

Autoscaling

The chart does not render HPAs, so you write them. An HPA computes utilization from resource requests, so set realistic requests first, then tune targets against measured load:

  • Web and API: scale on CPU and latency.
  • Workers: scale on queue depth, with a scale-down window long enough to drain a backlog.
  • Never autoscale beatworker, pi_beat_worker, monitor, or migration Jobs.

Plane's Kubernetes best-practices guide documents the chart's topology-spread and PDB values, including which workloads are deliberately excluded.

For node sizing, start from the requests in the manifests you will actually deploy (optional services change the total considerably) and leave headroom for a node drain and for a rollout where old and new pods coexist. The HA topology spreads workers across at least three zones. Then load-test the paths your teams use: imports, uploads, collaborative editing, API automation, search, and notification bursts.

Backup and restore

A Plane restore needs four things, and they only work together:

  • The PostgreSQL backup.
  • The object-storage bucket.
  • The Secret containing SECRET_KEY and LIVE_SERVER_SECRET_KEY.
  • The Helm values and overlays that describe the installation.

Record the Plane image version next to every database backup. Migrations only run forward, so a dump restored into a newer release may be altered before you know whether the backup itself was good.

Back up SECRET_KEY with the database

SECRET_KEY does more than sign sessions. Plane derives a Fernet key from it and uses that to encrypt instance-configuration rows, including SMTP passwords and OAuth client secrets. Change the key and those rows silently stop decrypting: Plane shows blank configuration rather than a clear cryptographic error.

⚠ Treat SECRET_KEY as part of the database backup. Store it in a versioned secret manager, copy the matching version into every restore drill, and keep it out of automatic rotation.

LIVE_SERVER_SECRET_KEY has a different job but must match between API and live. Keep both in one deliberately managed Secret referenced by external_secrets.app_keys_existingSecret.

Valkey holds only cache and short-lived coordination (sessions are in PostgreSQL), so it has nothing worth restoring. RabbitMQ carries work in flight; workers re-declare queues after a restore, and some in-flight work may be lost. Back up broker definitions only if you customized them.

Run the drill

⚠ Never test a restore against the live bucket. A drill release can run cleanup code and delete real objects.

Restore into a scratch database and a scratch bucket, deliver the original application keys, and install the same Plane version into a fresh namespace:

kubectl create namespace plane-dr

createdb -h "$DR_PGHOST" -U postgres plane_dr

pg_restore -h "$DR_PGHOST" -U postgres -d plane_dr \
--no-owner --no-privileges --clean --if-exists plane.dump

aws s3 sync s3://plane-prod-uploads s3://plane-dr-uploads

helm install plane-dr plane/plane-enterprise -n plane-dr --version <chart-version> \
--set planeVersion=<version-from-backup> -f values.dr.yaml \
--wait --wait-for-jobs --timeout 15m

Then prove it works like a user would: open a work item that predates the dump, download one of its attachments, sign in through your real identity provider, and trigger a queue-backed action such as a notification. An API 200 tests none of those relationships.

Record how long it took and every manual step. The drill shows whether someone other than the author can follow the runbook.

Observability

Running pods show process state. They can't tell you whether queues are draining, whether migrations finished cleanly, or whether a user can complete a real workflow.

Use each probe for its job

Commercial exposes /api/live/, /api/ready/, and /api/health/. Readiness and health check the database and cache; liveness only proves the process responds.

Use /api/ready/ for readiness and keep liveness free of external dependencies. A dependency-aware liveness probe restart healthy API processes during a database outage and make the incident noisier. Pair the built-in probes with a synthetic check that signs in or performs a low-impact, database-backed action. Note that /api/ready/ says nothing about RabbitMQ or end-to-end workflows.

Logs and telemetry

kubectl logs is fine for first response, but the evidence leaves with the pod. Ship logs to the platform your operators already use, and keep namespace, pod, container, release, node, and Helm revision as fields, so you can line up an API error with an ingress event, a migration Job, or a worker restart from the same minute.

For metrics and traces, observability.otel configures Plane's OpenTelemetry exporters; you provide the collector and backend. Start with backend telemetry. Browser telemetry needs a public receiver, rate limiting, CORS, and headers you're comfortable exposing to clients.

Alert on what users notice

Set alerts for:

  • A sustained rise in RabbitMQ messages_ready, which may mean workers are falling behind or have stopped consuming.
  • Unacknowledged messages that remain outstanding longer than expected.
  • A missing beat heartbeat or expected periodic log entry.
  • A failed or incomplete api-migrate Job after a release.
  • PostgreSQL connection saturation.
  • Certificates approaching expiry.
  • High PVC usage or volume-attachment errors for bundled stores.
  • A failed external synthetic check that exercises ingress, API, PostgreSQL, and authentication.

Confirm label selectors against the rendered manifests. The -wl suffix and the different labels on Jobs make name-based dashboards surprisingly easy to get wrong.

Upgrades and rollback

planeVersion sets the image tag for the release. The migration workload is an ordinary Kubernetes Job deployed alongside the updated application objects, not a Helm hook, so --wait --wait-for-jobs keeps the Helm client waiting without holding back the new pods. API, worker, and beat wait for pending Django migrations at startup, while old pods keep serving as the schema changes.

Treat the chart version, Plane version, values, overlays, database recovery point, and migration result as a single release. Render and diff that exact combination in staging first, paying attention to Jobs, selectors, PVC templates, ingress resources, and Secret references.

helm get values <release> -n <namespace> -o yaml > current-values.yaml

helm template <release> plane/plane-enterprise --version <target-chart> \
-n <namespace> -f values.yaml > rendered.yaml

helm upgrade <release> plane/plane-enterprise --version <target-chart> \
-n <namespace> -f values.yaml --wait --wait-for-jobs --timeout 15m

kubectl logs -n <namespace> job/<migration-job-name>

helm history <release> -n <namespace>

When the Job completes, sign in, edit a work item, open an attachment, and trigger a queue-backed side effect. That catches what a completed Job and Running pods miss.

⚠ Rollback does not reverse a migration. helm rollback restores Kubernetes objects and leaves the migrated database in place. Don't improvise a Django reverse migration mid-incident. If the old version can't run on the new schema, restore the matching database, application version, bucket state, and keys together.

Take a recovery point immediately before each upgrade. A nightly backup can be many hours old.

Security and access

Lock down three things before the first production login: the signing keys, the network boundary, and the admin UI.

Replace the signing keys

The chart ships public example keys so a first install can start. Swap them for your own Secret before anyone logs in:

env:
requireExplicitSecrets: true # fail the render if a key is missing

external_secrets:
app_keys_existingSecret: plane-app-keys # your Secret replaces the chart's keys

⚠ requireExplicitSecrets: true only rejects empty keys. The chart's values.yaml already fills in example keys, so the check passes with them in place. Point external_secrets.app_keys_existingSecret at your own Secret to replace them. If you set neither, the install notes warn that the public example keys are in use.

What goes in that Secret:

  • SECRET_KEY: long and random. Never rotate it for the life of the database.
  • LIVE_SERVER_SECRET_KEY: a separate value, identical in API and live.

If you keep any bundled service, also replace its default PostgreSQL, RabbitMQ, or MinIO credentials. For attachments, prefer external object storage with a private bucket and workload identity.

Add a NetworkPolicy

The chart doesn't create one. A reasonable starting point:

  • Allow the ingress controller to reach Plane's routed services.
  • Allow Plane to reach PostgreSQL, Valkey, RabbitMQ, object storage, and DNS.
  • Allow Plane to reach your identity provider, SMTP, webhooks, integrations, telemetry, and the license service.
  • Deny all other cross-namespace traffic.

Build the policy from observed traffic. An incomplete deny-all usually shows up as broken OAuth or email.

Scope the ServiceAccount

All Plane workloads share one ServiceAccount, so a cloud identity attached to it reaches every pod, bundled data stores included. If only API and worker need bucket access, give them their own ServiceAccount through an overlay.

Protect /god-mode/

Part of authentication is configured in the /god-mode/ admin UI, not your values file. It shares the public hostname by default, so restrict it at the ingress or an identity-aware proxy.

Use SKIP_ENV_VAR to pick one source of truth for auth settings:

  • Database owns them: change and rotate them in /god-mode/.
  • Environment owns them: change them in values; the UI just reflects what's deployed.

Either way, record which login methods are enabled, who owns each provider app, and how client secrets are delivered.

Need SSO, OpenTelemetry, or an air-gapped deployment? Review the self-hosting documentation, then talk to Plane if your topology goes beyond a standard Commercial deployment.

Incident playbook

Start from the user's symptom, check the narrowest Plane-specific cause first, and only then widen the search to Kubernetes.

Symptom
Check first
How to confirm

Pages stop syncing or cursors disappear

LIVE_SERVER_SECRET_KEY differs between API and live

Compare a hash of the value in both pods (command below).

Uploads fail after moving from MinIO to workload identity

A stale explicit AWS key in the rendered Secret

Inspect the pod env and doc-store Secret; empty or stale keys override the SDK chain.

Attachments upload but return 404

Endpoint, credentials, bucket policy, or public URL mismatch

Inspect the rendered storage env, identity binding, bucket policy, and generated asset URL.

Editor disconnects at about one minute

A 60-second timeout at ingress-nginx or the load balancer

Check ingress annotations and the WebSocket close time in the browser.

Imports and notifications stop, pods still Running

Workers not consuming RabbitMQ

Check messages_ready, unacked-message age, and broker-reconnect logs.

SMTP and OAuth settings blank after a restart

SECRET_KEY changed

Compare the running key with the version stored alongside the backup.

Beat override has no effect

services.beat-worker used instead of services.beatworker

Render the chart and check the Deployment's replica count.

The first row is the most common, and the quickest to confirm:

# The two hashes must match
for d in api live; do
kubectl exec -n plane deploy/plane-$d-wl -- \
sh -c 'printf %s "$LIVE_SERVER_SECRET_KEY" | sha256sum'
done

For anything else, classify the failure domain: Pending pods and attachment events point to capacity or storage; a failing /api/ready/ points to a dependency; a failed migration Job points to release state; and a correct Ingress with no route points to controller admission, the load balancer, DNS, or certificates.

Ownership and the runbook

The first ten minutes of an outage shouldn't go to finding the release name or the person who can restore PostgreSQL. Name the team that owns the Helm release, and record owners and escalation paths for the database, object storage, RabbitMQ, Valkey, ingress, DNS, certificates, and observability.

Keep the runbook outside Plane, and have it point to the systems that hold the answers rather than copying every command. At minimum it should record:

  • Chart version, Plane version, Helm revision, namespace, and hostname.
  • The exact values file plus every post-render or Kustomize patch.
  • External-service endpoints and the teams that own them.
  • Secret names, and where the non-rotating SECRET_KEY is stored.
  • The Plane version attached to each database backup.
  • The date, duration, and result of the latest restore drill.
  • Authentication methods configured outside the values file.
  • The ingress controller, certificate owner, timeouts, and upload limits.

The record should tell an operator what changed before the failure, which recovery point matches the running version, and who can act on the dependency that broke.

Production hardening checklist for Plane on Kubernetes

1. Secrets and storage

The chart ships with example keys and default credentials so a first install can start. Replace them before the first production login, and keep SECRET_KEY with the database it encrypts.

  • Supply your own application-key Secret and enable requireExplicitSecrets.
  • Store SECRET_KEY with the database backup and exclude it from rotation.
  • Replace bundled database, broker, and object-storage credentials.
  • Use external object storage with a reviewed bucket policy and a tested upload/download path.
  • Externalize every data service that needs independent availability and backups.

2. Traffic

Test the editor, uploads, and sign-in through the same path your users take.

  • Set ingress.controller and ingress.ingressClass explicitly and confirm admission.
  • Keep a /live/ WebSocket open beyond every idle timeout on the path.
  • Order upload limits: load balancer > proxy > Plane.
  • Verify SESSION_COOKIE_SECURE and CSRF_COOKIE_SECURE on a running API pod.

3. Scale and health

API replicas are limited by database connections, and worker replicas should follow queue depth.

  • Keep services.beatworker.replicas at one.
  • Set worker concurrency explicitly and scale workers on RabbitMQ backlog.
  • Set realistic resource requests before adding HPAs.
  • Use /api/ready/ for readiness, plus a synthetic check on a real user path.

4. Recovery and change

A restore needs the database, bucket, application keys, and values from the same point in time. Migrations only run forward, so take a backup you can return to before every upgrade.

  • Back up PostgreSQL, the bucket, application keys, and values as one set.
  • Restore into a fresh namespace and open an old work item and its attachment.
  • Pin chart and Plane versions; render and diff every upgrade.
  • Upgrade with --wait --wait-for-jobs, then read the migration log.
  • Take a recovery point before each upgrade, because rollback won't undo migrations.

5. Operations

Keep logs, alerts, and the runbook somewhere that still works when Plane is down.

  • Ship logs off-cluster with Kubernetes and Helm metadata.
  • Alert on queue backlog, migration failure, dependency health, cert expiry, and storage pressure.
  • Keep the runbook outside Plane and assign an owner to every external dependency.

When Kubernetes is the right choice

Kubernetes fits teams that already operate clusters, ingress, secrets, managed data services, observability, and recovery drills. For them, Plane's stateless workloads spread cleanly across nodes and zones, and every release is something you can render, review, and rebuild.

But you also inherit a control plane. A small team without that operating model will often get a safer result from Plane Cloud or the Docker deployment path.

Planning to run Plane on Kubernetes?

Talk through your architecture, availability requirements, and recovery plan with the Plane team before you deploy. Schedule a call with our team.

Frequently asked questions

Q1. Is self-hosted Plane on Kubernetes suitable for production?

Yes, if you make the decisions above deliberately. The quick install is a starting point; production readiness comes from external stateful services, multi-zone topology, a tested ingress path, restore-tested backups, monitoring, and clear ownership.

Q2. How much CPU and memory does it need?

It depends on concurrency, imports, integrations, optional search and AI services, and whether data stores run in-cluster. Start from the chart's requests, leave headroom for drains and rollouts, then load-test and tune. See "Scaling and capacity" above.

Q3. Can it use our existing PostgreSQL, Redis, RabbitMQ, and object storage?

Yes. Each has a switch and a remote connection key, listed in Configure external services. For HA, the external service must be HA too: a single external node is still a single point of failure.

Q4. Can we run it with no internet access?

Yes, via the supported air-gapped path: mirror the images, transfer the chart, enable air-gapped mode, and add your internal CA. Either way, list your outbound destinations before go-live: licensing, telemetry, registry, identity provider, SMTP, and every integration you enable.

Q5. Can we migrate an existing Docker or Community Edition deployment to Commercial Edition on Kubernetes without losing data?

Yes. You move the PostgreSQL database, the object-storage bucket and the original SECRET_KEY into a new plane-enterprise release that runs the same Plane version. It's the same process as a restore drill. Run it in a separate namespace first, and confirm that old work items, attachments, sign-in, and notifications all work before you schedule the cutover. For large or regulated environments, talk to Plane before you set a date.

Recommended for you

View all blogs
Plane

Every team, every use case, the right momentum

Hundreds of Jira, Linear, Asana, and ClickUp customers have rediscovered the joy of work. We’d love to help you do that, too.
Plane
Nacelle