Self-Hosted n8n Production Readiness Checklist
A practical checklist for verifying ownership, persistence, encryption, HTTPS, webhook routing, updates, monitoring, and recovery before self-hosting n8n.

Checked against the cited sources on .
Confirm ownership and deployment prerequisites
Start by separating workflow readiness from operational readiness. A workflow that behaves correctly in practice is not automatically ready for a self-hosted production environment. Self-hosting requires technical knowledge, and the team operating it must make deliberate choices about infrastructure, state, maintenance, and support.
Check that the host, database, encryption key, certificates, upgrades, monitoring, incident response, and recovery each have a named owner. This is an editorial acceptance standard, not a requirement established by the supplied sources. Consider this check complete when every responsibility has an accountable owner, a documented procedure, and a backup contact or escalation path.
Document the intended topology and workload assumptions. Suggested questions include: Which processes will run? Where will state live? What concurrency is expected? What resource ceiling is acceptable? Who owns infrastructure cost and planned downtime? There is no organization-independent sizing threshold in the supplied evidence. One referenced AWS module shows why limits matter—minimum capacity can create continuous cost, while a hard maximum can leave work pending—but its settings cannot be generalized to every deployment.
Record the exact workflow version, dependencies, integrations, environment variables, and production endpoints included in the approval. Completion means reviewers can identify what is being promoted, who will operate it, and which assumptions still require validation.
Protect persistent state and credential encryption

Check that persistent storage retains the n8n data directory across container restarts. The supplied Docker guidance mounts persistent storage at the n8n data directory, although that page is marked outdated and recommends Docker Compose. Treat the path as evidence that persistence matters, not as a complete production architecture.
Document every state store used by the chosen topology, including the database and persistent instance data. Perform a controlled restart and confirm that expected workflows, credentials, and required configuration remain available. This completion test is an editorial suggestion; the source does not prove that a particular environment will preserve every required component.
Set an instance encryption key and retain it through an approved secret-handling process. In queue mode, every worker must receive the same key so stored credentials remain usable. Suggested checks include confirming that the key is absent from source control and test records, access is restricted, and an authorized recovery path exists. Those handling practices are editorial safeguards because the supplied evidence does not prescribe secret-manager integration, escrow, access controls, or routine rotation.
If the team plans to rotate the encryption key, create a full database backup first and verify that the main process and all workers use consistent keys. The supplied rotation guidance describes this as a one-way operation, so record the decision and recovery plan before proceeding. Completion means credentials still decrypt after controlled restarts and no key material appears in the evidence package.
Configure HTTPS and public webhook routing
Place a supported TLS termination layer in front of n8n. The supplied guidance recommends a reverse proxy or network load balancer that handles TLS and certificate renewal. It does not establish required cipher suites, protocol versions, certificate authorities, or network-access rules, so your security team must define those separately.
Assign ownership for certificate issuance and renewal, then test the externally visible editor and webhook endpoints over HTTPS. Completion means clients reach the intended endpoint securely and the responsible owner can explain how renewal failures will be detected and handled. The operational test and ownership requirement are editorial acceptance checks.
Behind a reverse proxy, configure the external webhook URL and the proxy hop count for the real routing path. The documented hop count of one applies to a single-proxy arrangement and should not be copied blindly into a more complex chain. Verify that the editor displays the intended public webhook URL and that an external test request reaches the correct production-like workflow.
Use a non-destructive test payload and avoid pointing a readiness drill at live downstream actions. This is a suggested testing precaution, not a validated n8n test standard. Record the requested URL, observed response, workflow execution, and any proxy headers needed for troubleshooting.
Establish a staged update procedure

Pin the version used in production and define a repeatable promotion path. Before rollout, review release notes for breaking changes and test the update on a separate instance. The supplied guidance supports staged testing but does not define a universal maintenance window, rollback method, or compatibility suite.
Create an editorial test checklist for the workflows that matter to your team. Suggested checks include triggers, credentials, data transformations, error paths, downstream API calls, webhook registration, and representative execution volume. These are proposed coverage areas, not a validated instrument or mandatory standard.
Define rollback before deployment rather than during an incident. Record the current version, target version, relevant release-note findings, backup reference, test results, decision maker, and rollback trigger. Completion means the selected version has passed the team’s documented checks and the recovery decision can be executed by the assigned operator.
Sources: S2
Set up health monitoring and workflow-failure alerts
Monitor both service health and workflow outcomes. A supplied Kubernetes chart enables liveness and readiness probes, showing that these endpoints can serve as health signals in that chart. The source is chart-specific and truncated, so enabled probes alone do not prove adequate monitoring, alert delivery, retention, or operational coverage.
Connect health signals to a notification route owned by a person or team. Separately, configure an error workflow beginning with the Error Trigger so failed executions can notify operators. The documentation gives email and Slack as examples, but it does not show that any alert was delivered or acted on in a real deployment.
Run two controlled checks: deliberately fail a safe test workflow, and simulate an unhealthy instance in a non-production environment. Consider monitoring ready only when both conditions create understandable, actionable notifications received by the responsible team. This is an editorial completion criterion; organizations must choose their own alert latency, escalation rules, retention, and acceptable downtime.
Record who acknowledges alerts, where diagnostic information can be found, and what happens outside normal working hours. Avoid treating a green dashboard or configured notification node as proof that the response process works.
Back up state and complete a recovery drill

Create protected backups that cover the database, workflows, credentials, persistent instance data, and encryption key. The self-hosted CLI supports workflow and credential export and import, which can contribute to a logical backup process. Exports alone do not demonstrate recoverability, and decrypted credential exports expose sensitive material.
Plan imports carefully because matching identifiers may be overwritten. Keep restore drills isolated from production endpoints and downstream systems. The isolation requirement is an editorial safety measure, while the overwrite and credential-exposure risks come from the supplied CLI guidance.
Perform an isolated recovery drill rather than stopping after backup creation. Suggested completion checks are that the restored instance opens, expected workflows appear, credentials decrypt, and a representative workflow executes successfully without contacting production endpoints. No firsthand recovery test was supplied, so these checks are proposed acceptance criteria rather than evidence that any deployment has passed them.
Record backup scope, storage ownership, access permissions, restore steps, observed results, and unresolved gaps. Your organization must choose backup frequency, retention, recovery objectives, and acceptable data loss because the supplied sources provide no universal thresholds.
Record approval evidence and remaining risks
Assemble one approval record linking the owners, topology, persistence test, encryption-key handling, HTTPS and webhook checks, staged update results, alert simulations, backup reference, and recovery-drill outcome. Mark each item passed, conditionally accepted, or blocked, and name the person accepting any residual risk.
Treat this checklist as a focused operational review, not a complete compliance or security audit. It covers the requested production-readiness areas but does not establish measured reliability, security, cost savings, acceptable downtime, or recovery performance. Requirements drawn from queue mode, Kubernetes, AWS modules, CLI operations, or key rotation should be applied only when the selected topology uses them.
Before approval, repeat the workflow itself in a safe environment and confirm that its behavior matches the version being promoted. Keep this application-level practice distinct from infrastructure approval: completing an automation exercise does not demonstrate that hosting, monitoring, or recovery controls are ready.


