n8n crashes with JavaScript heap out of memory on AI agent nodes
Your container dies mid-execution and comes back with no explanation beyond one line in the log. That line tells you which of two completely different problems you have — and they have opposite fixes.
A self-hosted n8n instance that had been stable for months starts restarting. The workflow is the same one as always, except now it has an AI agent node in it, and sometimes a loop around that node. The logs end with a stack trace you did not write, and the container comes back up as if nothing happened — until the next execution.
The first job is not to fix it. It is to find out which of two different failures you are looking at, because the standard advice for one makes the other worse.
Reading the FATAL heap error in n8n logs
There are two ways an n8n container dies for memory reasons, and they look similar from the outside and are nothing alike underneath.
<--- Last few GCs --->
[1:0x7f8b2c0] 412039 ms: Mark-sweep 3969.1 (4066.5) -> 3968.3 (4067.2) MB
<--- JS stacktrace --->
FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory$ docker inspect n8n --format '{{.State.ExitCode}} {{.State.OOMKilled}}'
137 trueThe distinction decides your next move. In case A the process had heap headroom left according to the OS, but V8 refused to grow past its configured old-space limit. In case B the container hit its cgroup memory limit and was terminated — no warning, no trace, exit 137.
Why AI agent nodes retain memory
n8n moves data between nodes as an in-memory array of items. That design is what makes the editor feel immediate — you can inspect any node's output at any time — and it is also why the memory profile of a workflow is roughly the sum of everything every node has produced, not the size of what the current node is working on.
An agent node stacks several multipliers on top of that:
- Conversation memory. A buffer window keeps the last N exchanges alive so the agent has context. N is a count of turns, not a count of bytes — a window of 10 is small with short messages and enormous once a tool starts returning documents.
- Tool outputs go into the context. Everything a tool returns is appended to the conversation so the model can reason about it. A tool that returns a full API response puts that whole response in memory, and keeps it there for the rest of the run.
- Sub-node output is retained. Each sub-node under the agent keeps its own output for inspection in the editor, so the intermediate steps do not disappear when the agent moves on.
- Binary data defaults to memory. Unless you set the binary data mode to filesystem or S3, attachments and downloads live in the heap alongside everything else.
- Loops multiply all of the above. Run the agent once per row of a 500-row spreadsheet and you are not running 500 small executions — you are running one execution that accumulates 500 agent contexts.
This is why the crash correlates with data volume rather than workflow complexity, and why it shows up weeks after the workflow was built: nothing changed in the logic, the inputs just got bigger.
Self-hosted crash loops and container restarts
The restart is what turns a bad execution into a bad afternoon. With a restart policy in place, the container comes back, n8n picks up where it can, and if the triggering data is still queued it runs the same doomed execution again. Three or four cycles in, the picture is a service that is technically up and getting nothing done.
Things worth setting before you touch anything architectural — none of these are fixes, they are how you stop losing information about the failure:
| Setting | What it does | Why it matters here |
|---|---|---|
| N8N_DEFAULT_BINARY_DATA_MODE=filesystem | Attachments off the heap | Removes the largest single contributor in file-heavy flows |
| EXECUTIONS_DATA_SAVE_ON_SUCCESS=none | Stops storing successful run payloads | Shrinks both the database and what is held while running |
| EXECUTIONS_DATA_PRUNE=true | Ages out old execution data | Stops a slow leak that looks like a memory leak |
| EXECUTIONS_MODE=queue | Executions run in separate workers | A crash kills one worker, not the editor and every other flow |
| Container memory limit ≥ heap limit + overhead | Aligns the two ceilings | Turns silent exit 137 kills back into readable JS errors |
That last row is the one people skip. If the container is limited to 2 GB and V8 is configured for 4 GB, V8 will never reach its own limit — the kernel gets there first, and you never see the FATAL error that would have told you what was going on.
External state and per-agent context
Every mitigation above is about surviving the accumulation. The structural fix is to not accumulate: the agent's context stops being a variable inside a process and becomes a record outside it.
That changes what a crash means. When conversation history, intermediate results and tool outputs live in an external store, the process handling a step holds only that step. It can be restarted, moved to another machine, or run at the same time as a hundred others, because there is nothing in it worth preserving.
It also changes what a loop means. Five hundred rows become five hundred independent units of work with their own isolated context, instead of one execution whose memory footprint grows with every row. Failure stops being all-or-nothing: row 340 fails, rows 1 to 339 are already done, and row 340 retries on its own.
This is how Bentho runs agents — durable per-agent context in external state, executed by workers that hold one unit of work at a time. It is not a tuning parameter, which is the point: there is no value of --max-old-space-size that makes an in-process design stop accumulating.
Migration path for critical flows
Nobody should migrate a working n8n instance wholesale, and you do not need to. The flows that crash are a small subset — typically the ones with an agent inside a loop — and they are the ones worth moving first.
- Confirm which failure you have. Check the exit code and whether OOMKilled is true. Everything downstream depends on this answer.
- Buy headroom the cheap way. Binary data to filesystem, execution data pruned, queue mode on. This often stops the restarts entirely and always gives you room to work.
- Find the accumulating node. It is almost always an agent, a loop, or an HTTP node returning something much larger than anyone assumed. Log item counts and payload sizes between nodes.
- Move that one flow. The agent runs as a Bentho worker with external state; n8n keeps the triggers and the integrations it is good at, and calls the worker instead of hosting the agent.
- Verify under the load that broke it. Replay the real volume, not a test row. A flow that survives one row was never the problem.
- Leave the rest alone. The flows that never crashed do not need a new platform.
Step one takes two minutes and rules out half the advice on the internet. Start there, even if you do nothing else today.
Frequently asked questions
What does FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory mean in n8n?
It means the Node.js process exhausted the heap V8 was configured to use and stopped itself. In n8n it usually means one execution accumulated more data in memory than the limit allows — commonly an AI agent inside a loop, or binary data being held in the heap. It is not a bug in your workflow logic; it is the amount of data the execution holds at once.
Is exit code 137 the same problem?
No, and telling them apart matters. Exit 137 with OOMKilled=true means the kernel killed the container for exceeding its memory limit — there is no stack trace because the process was never asked. The heap FATAL ERROR is V8 stopping itself. Raising --max-old-space-size helps the second and makes the first worse.
Will more RAM fix it?
It postpones it. More RAM raises the ceiling, and an execution whose memory grows with the size of its input will reach any ceiling eventually. It is a reasonable emergency measure and a poor plan.
Does queue mode solve the agent memory problem?
It contains it. Executions run in separate workers, so one bad run kills a worker instead of taking down the editor and every other workflow. The execution that accumulates too much still dies — it just stops taking everything else with it.
Why did this start happening without changing the workflow?
Because memory use tracks data volume, not workflow complexity. A flow built against 20 rows and a short conversation history behaves completely differently against 2,000 rows and tool outputs that grew over time.