Migrating from n8n to Bentho: scaling agentic work in production
A phased migration you can reverse at every step: what to extract first, how to run both stacks side by side, and the checklist for the day you turn n8n off.
You are not here because n8n is bad. You are here because the flows that move data still work fine and the ones with a model in them have started to wake people up at night.
This guide moves the second kind and leaves the first alone. Every step is reversible until the last one, and the last one is a checklist, not a leap.
1. Why n8n runs short with LLM agents in production
n8n runs a workflow as one execution inside one process, holding the data of every node it has run so far. That is the right design for moving records between services, and it is the thing that bends when a step is an agent: the call is long, the payload is large, and the output is not predictable.
Three documented defaults describe the shape of that ceiling. They are defaults, not laws — check yours before you quote them back to anyone:
EXECUTIONS_MODEships asregular: without changing it, workflows run inside the main process itself.EXECUTIONS_TIMEOUTships as-1— no timeout at all. The ceiling a user can set on an individual workflow isEXECUTIONS_TIMEOUT_MAX, 3600 seconds.- Scaling out means queue mode, and queue mode means Redis, a persistent database (PostgreSQL recommended; SQLite is not supported) and several instances sharing
N8N_ENCRYPTION_KEY.
Read that list together and the problem states itself: adding workers multiplies how many executions you can hold, not how long one may take or how much it may hold in memory. A worker running an agent is still one process carrying the whole conversation.
2. Migration map: n8n nodes to typed tools
The unit of migration is not the workflow. It is the authority: the thing that knows the real answer — prices, stock, shipping rates, order state. In n8n that knowledge is spread across nodes and expressions; in Bentho it becomes a tool with a signature.
| What you have in n8n | What it becomes | Why |
|---|---|---|
| A chain of HTTP nodes that reads a catalogue and builds a price | One call that asks for a price and gets one | One authority, one answer, and something to verify it against. |
| An AI Agent node with tools attached | The orchestrator, with transitions in code | The model extracts; the code decides. The graph stops being the decision. |
| A Set/Code node normalising a payload | Typed input validation at the boundary | A drifting schema fails where it arrives, not three nodes later. |
| Credentials per connection | Per-tenant claims checked by each service | A service refuses a call whose token tenant isn't the request's. |
The tools are mounted next to the REST routes that already exist, so a consumer can move one at a time. Nothing has to be switched off for the first one to work.
3. Getting the work out of the execution
Two things have to leave the execution process, and they are different problems:
- The waiting. A step that blocks for minutes holds a slot and everything it has accumulated. Once the work is a request to a service that answers on its own time, the flow stops being what has to stay alive.
- The state. Conversation memory inside an execution dies with it and grows with it. Out of the process, it belongs to the session and survives a restart.
There is a third, quieter one: what n8n keeps about past runs. EXECUTIONS_DATA_PRUNE ships enabled with a 336-hour window and 10000 executions retained — fine for records, heavy when every run carries model payloads. Check it before you conclude the database is the problem.
4. Pre-validation and dual-run, before anything is shut down
This is the part people skip, and it is the only part that makes the rest safe. Do not cut over. Run both, compare, and let the comparison decide.
- Freeze the contract first. Write down what each extracted tool takes and returns, with types. If you cannot write it down, it is not ready to move.
- Replay real traffic. Feed the same inputs to both stacks. Not synthetic ones: the payloads that already broke something are the valuable ones.
- Compare outputs, not status codes. Two green runs that disagree on a number is the failure this whole migration exists to catch.
- Shadow first, then split. Bentho answers alongside n8n without anyone seeing it, until the diffs are boring. Only then does a share of real traffic move.
- Keep n8n warm. Reversible means the old path still runs, not that you could rebuild it.
5. Cutover checklist
- Every extracted tool has its signature written and validated at the boundary.
- Dual-run has been on long enough to cover a full business cycle, month-end included.
- The output diff is explained — not just small. A pattern you cannot explain is a bug you have not found.
- Rollback is one switch, and someone who is not you has used it in a drill.
- Alerting points at the new path, and at the diff between the two.
- The n8n workflow is disabled, not deleted, and its credentials are rotated on a date you wrote down.
6. Common mistakes when replacing the AI Agent node
- Recreating the graph one-to-one. If the new system has the same branches drawn in code, you moved the problem and paid for the move.
- Letting the model keep deciding. The point is that the code decides and the model extracts, at temperature 0. A prompt that returns 'the next step' is the AI Agent node again with extra steps.
- Retrying a bad answer. A blind retry on a non-deterministic step is the same gamble, paid twice. An unbacked figure should be discarded, not re-rolled.
- Migrating the flows that already work. The ones without a model in them are not the problem, and moving them spends the goodwill you will need.
- Shutting down before month-end. Whatever is going to disagree will disagree on the busiest day.
Where the n8n numbers come from
The n8n defaults quoted above are from n8n's own documentation. They change between versions, so treat these as a starting point and check what your own install actually says.
- executions —
EXECUTIONS_TIMEOUTships as-1, which disables the cap: a run can take as long as it takes.EXECUTIONS_TIMEOUT_MAX— the ceiling a user may set on an individual workflow — ships as 3600 seconds. - executions —
EXECUTIONS_MODEacceptsregularorqueueand ships asregular: untouched, workflows run inside the main process itself. - enable-queue-mode — Scaling out requires queue mode, and with it Redis as the broker, a persistent database (PostgreSQL recommended, SQLite NOT supported) and several instances sharing
N8N_ENCRYPTION_KEY. Each worker's concurrency defaults to 10, and n8n recommends not going below 5. - executions —
EXECUTIONS_DATA_PRUNEships enabled, withEXECUTIONS_DATA_MAX_AGEat 336 hours andEXECUTIONS_DATA_PRUNE_MAX_COUNTat 10000 executions retained.
Checked in n8n's documentation on 2026-09-24. It may have changed since: if you are about to make a decision on one of these, check your own.
Frequently asked questions
How do I scale n8n with large LLMs?
Queue mode adds workers, which raises how many executions run at once. It does not shorten a single long call or shrink what one execution holds in memory, so for agent work the fix is to move the call out of the execution rather than to add workers.
Can I replace only the AI Agent node and keep the rest?
Yes, and that is the recommended first step. The node becomes a call to a typed tool; the surrounding workflow, its triggers and its connectors stay exactly as they are.
How long should dual-run last?
At least one full business cycle, month-end included. The threshold is not time, it is that you can explain every difference between the two outputs.
What if we need to roll back after shutting n8n down?
Disable the workflows instead of deleting them, and rotate credentials on a date you set in advance rather than on cutover day. That keeps the old path one switch away for as long as you need it.