Why Make.com scenarios time out on large LLM calls — and how to fix it
A blocking HTTP request cannot wait for a model that thinks for four minutes. Here is how to read the error, where the real ceiling is, and how to decouple the call without rewriting the scenario.
The scenario worked for months. Then you swapped the model for one that reasons before answering, or you let the prompt grow past a few thousand tokens, and now the run dies partway through with a timeout — sometimes at 40 seconds, sometimes at five minutes, and almost never at the same step twice.
This is not a flaky integration. It is the predictable result of putting an operation whose duration you do not control behind a connection that has to stay open the whole time. Below: how to confirm that is what is happening, why raising the timeout only moves the wall, and what the shape of the fix looks like.
Causes of Make scenario timeouts with LLMs
Make executes a scenario as a chain of synchronous steps. Each module opens a request, blocks until it gets a response, hands its output to the next module, and only then releases. There is no place in that model for a step that answers in its own time, which is precisely what a reasoning model is.
Three things push a model call past the ceiling, and they compound:
- Reasoning tokens. A model that deliberates before answering can spend most of its wall-clock time producing output you never see. The response is short; the call is not.
- Prompt size. Every extra thousand tokens of context is more time to first token. Scenarios that re-inject an entire document, or a full JSON payload from a previous step, pay this on every single call.
- Tool loops. If the call goes to an agent rather than a bare completion, one request can hide several model round-trips plus whatever the tools themselves take.
None of these is visible from the Make side. From there it is one module, one line in the execution log, and one duration that occasionally exceeds the limit.
The hard limits: HTTP modules vs native agents
It helps to know which ceiling you are hitting, because they are not the same ceiling and they do not have the same escape hatch. At the time of writing, Make documents roughly this shape — check the numbers against your own plan before you build on them:
| Where the limit lives | Typical value | Can you raise it? |
|---|---|---|
| HTTP module timeout | 40 s default | Yes, up to a documented maximum (300 s) |
| Native app modules (OpenAI, Anthropic…) | Fixed by the connector | No |
| Agent / multi-step steps | Minutes, per step | Partially |
| Whole scenario execution | ~40 min | No |
| Webhook response window | Seconds | No — the caller is waiting |
So the first fix everyone tries — raise the timeout — works exactly once. You move from 40 s to 300 s, the failures stop for a few weeks, and then a slightly longer document or a slightly chattier agent puts you back where you started, except now every failure costs five minutes of a scenario slot instead of forty seconds.
Blocking calls vs multi-step reasoning
The mismatch is structural. A blocking call assumes the work is bounded and fast: you ask, it answers, the connection closes. Multi-step reasoning is neither. It is bounded only by the problem, and its duration is a distribution with a long tail, not a number.
Putting the second inside the first has costs beyond the failure itself:
- You pay for work you throw away. The model finished thinking; the provider billed those tokens. Make hung up before the answer arrived, so you paid and got nothing.
- Retries multiply the bill. An automatic retry on a timeout does not resume anything — it starts the whole call over, at full price, with the same odds of being cut off.
- Partial state is unrecoverable. Step four succeeded, step five timed out. Whatever step four wrote is already in your CRM, and the scenario has no way to know how far it got.
- The failure is not reproducible. The same input succeeds on Tuesday and fails on Thursday, which makes it nearly impossible to fix by reading the scenario.
If the duration of a step is not something you control, the connection to that step should not be something you hold open.
The fix: async queues and durable state
The shape of the solution is the same whether you build it yourself or adopt a platform: the request that starts the work and the request that collects the result stop being the same request.
Concretely, that means four pieces. An endpoint that accepts the job and returns immediately with an id. A queue that holds it. A worker that runs the model call with no connection to keep alive, so its only real limit is the provider's. And a durable record of where the job got to, so a crash resumes instead of restarting.
POST /v1/jobs → 202 Accepted { "id": "job_8fA2", "state": "queued" }
# minutes later, from anywhere — the scenario, a webhook, a retry:
GET /v1/jobs/job_8fA2 → 200 OK { "state": "done", "output": { … } }This is what Bentho runs agents on: work is accepted as an event, executed by a worker with persistent state, and reported back when it is done — by callback if you want to be told, by polling if you would rather ask. The step that used to be a four-minute blocking call becomes a request that returns instantly, and the scenario stops being the thing that has to stay alive.
The part that matters for cost: because state is durable, a failure in step five does not throw away steps one through four. It resumes.
Migration checklist
You do not have to move the whole scenario. In most cases one or two steps are responsible for every timeout you have seen. Work in this order:
- Measure before you change anything. Export the last 30 days of executions and get the duration distribution per module. You are looking for the p95, not the average — the average is fine, and that is why the problem is confusing.
- Separate the genuinely slow steps from the merely large ones. A step that is slow because the prompt carries a 40-page document is fixed by not carrying the document, not by going async.
- Make the slow step asynchronous first. Replace the blocking module with the accept-and-return call. Keep everything else exactly as it is.
- Decide how the result comes back. A callback into a second scenario is usually cleaner than polling, and costs fewer operations.
- Make the whole thing idempotent. Pass your own job key so a retry cannot produce two of whatever the scenario creates.
- Only then look at what else can move. Once the timeouts stop, the remaining reasons to migrate are cost and state — which is a decision you can make calmly.
If your scenarios are failing in production right now, the measurement step is the one worth doing today. It takes an afternoon and it tells you whether you have a timeout problem or a prompt-size problem, and those have completely different fixes.
Frequently asked questions
Why does the same scenario succeed sometimes and fail others?
Because model latency is a distribution, not a constant. The same prompt can take 30 s or 200 s depending on how much the model reasons, current provider load, and output length. A fixed timeout cuts off the slow tail of that distribution, which is why failures look random.
Can I just raise the Make timeout to its maximum?
You can, and it buys time rather than solving the problem. The maximum is still a fixed number in front of an unbounded operation, and every failure now costs five minutes instead of forty seconds. Use it to stop the bleeding while you decouple the call.
Do I get billed for a call that timed out?
Generally yes. The timeout happens on your side of the connection; the provider still ran the model and still counted the tokens. This is why retrying on timeout is an expensive strategy.
Does this apply to Make's native AI agents too?
Yes, with a larger budget. Agent steps allow minutes rather than seconds, which pushes the wall further out but does not remove it — and an agent hides several model round-trips behind one step, so it reaches the wall in ways that are harder to predict.