Skip to content
← Back to News

Architecture10 min read

Token cost: linear workflows vs agent orchestration

Most of an automation bill is not reasoning. It is the same instructions, the same schemas and the same JSON, paid for again at every node. Here is where it hides and what it costs.

The invoice grew faster than the traffic. Volume went up maybe 30% and API spend doubled, and nobody can point at the feature that did it. That gap has a structural explanation, and it is not the reasoning — it is what the graph carries between nodes.

Where token spend hides in Make and n8n graphs

A linear workflow has no memory between steps. Each node that calls a model is a stateless call: it knows nothing about the previous one, so everything it needs has to travel with it. In practice, that means four passengers ride along on every single call.

  • The system prompt. Written once, sent every time. Three model nodes in a flow means paying for the same instructions three times per run.
  • Tool and schema definitions. Most integrations send the complete set of available tools on every call, whether or not this step could possibly use them.
  • The previous step's full output. The easiest way to wire two nodes is to pass the whole JSON. The model needs two fields; it receives forty.
  • Accumulated conversation. In agent nodes, the whole exchange rides along again on each turn, and it only grows.

None of this is visible in the editor. The canvas shows a clean arrow between two boxes; the arrow is carrying several thousand tokens, and you pay for it at input rates on every run.

Schema re-injection and prompt bloat

Schema re-injection is the largest of these and the least obvious. A tool definition is not free text: it is a JSON schema with names, types, descriptions and nesting, and a realistic set of six or seven tools runs to the better part of a thousand tokens. Sending all of them to a step whose only job is to classify a ticket into three categories is pure waste — and it is the default behaviour almost everywhere.

Prompt bloat is the same failure applied to data. A node returns an API response with forty fields; the next node passes it through whole, because passing it whole is one click and picking fields is ten. The model reads all forty, reasons about two, and you paid for forty.

The cheapest token is the one you never send. The second cheapest is the one you send once.

Deterministic flows vs goal-oriented agents

It is worth being fair here, because this is not a case where one architecture wins everything.

Linear flowAgent orchestration
Path through the workFixed, identical every runChosen per case
Calls per runConstant and predictableVariable — sometimes fewer, sometimes more
Context carriedEverything, every callWhat this step needs
Cost of a simple caseThe same as a hard oneLow — it stops when it is done
Cost of a hard caseFails or needs another branchHigher, and it finishes
AuditabilityExcellent — you can read the graphRequires tracing per run
A deterministic flow is the right answer when every case really does go the same way. The waste appears when a fixed path is used to handle cases that do not.

The key asymmetry: a linear flow pays its worst-case cost on every run, because the path does not adapt. An agent pays roughly what each case is worth. If 80% of your cases are simple, that difference is most of your bill.

Cost controls: pruning and strict schemas

Three controls do most of the work, and they are worth applying whatever platform you are on:

  • Semantic context pruning. What travels to the model is what this step needs to decide, not everything that happened before. Older context gets compacted into a summary rather than carried verbatim.
  • Conditional tool selection. Only the tools that are plausible for the current step get their schema sent. The other six exist and cost nothing until they are relevant.
  • Strict output schemas. additionalProperties: false and a closed field set. The model returns the fields and nothing else, which shortens the output and removes a whole class of parsing failures.

These are defaults in Bentho rather than options, which is the only reason they hold up over time: a cost control you have to remember to apply on each new flow is one you will stop applying in month three.

A simple ROI model for API budgets

Take a support flow: classify an incoming ticket, look up account context, draft a reply. Three model calls. Assume a 400-token system prompt, 900 tokens of tool definitions, a 2,500-token JSON payload carried between steps, and a 600-token ticket.

Per runLinear flowPruned agent
System prompt400 × 3 = 1,200400 × 1 = 400
Tool / schema definitions900 × 3 = 2,700~250 × 3 = 750
Carried payload2,500 × 3 = 7,500~700 × 3 = 2,100
The ticket itself600 × 3 = 1,800600 × 1 = 600
Prior-turn context—~300 × 2 = 600
Input tokens per run13,2004,450
Output tokens per run~900~900
Same three decisions, same business outcome. The difference is passengers, not reasoning.

At 10,000 runs a month, that is 132M input tokens against 45M — a 66% reduction in input spend. At an input price of $3 per million tokens, roughly $396 a month becomes roughly $134. The absolute figure is small at this volume and the ratio is what travels: it holds at 100,000 runs, where the same 66% is the difference between a line item and a conversation with finance.

Two caveats worth stating plainly. Agent orchestration has variance — a hard case can take more calls than the fixed flow would have, so the saving shows up in the monthly total and not in any individual run. And prompt caching, where your provider offers it, cuts the repeated-prefix portion of the linear column substantially; it helps, and it does not touch the carried payload, which is the biggest row in the table.

The useful exercise is not to trust this table. It is to run the same breakdown against one of your own flows: count the tokens each node actually sends, and see how many of them are the same tokens you already paid for at the previous node.

Frequently asked questions

Are agents always cheaper than linear workflows?

No. An agent is cheaper when cases vary in difficulty, because it stops when it is done rather than running a fixed path every time. For genuinely uniform work with no branching, a deterministic flow is simpler, more auditable and costs about the same.

Does prompt caching solve this?

Partly. Caching helps a lot with the repeated system prompt and tool definitions, which is a real portion of the waste. It does nothing about the carried payload — the full JSON passed between nodes — because that changes every run, and in the model above that is the single largest row.

Why does a strict JSON schema reduce cost?

Because a loose schema invites the model to add fields, restate the input and explain itself, and all of that is output tokens, usually billed at several times the input rate. A closed schema with additionalProperties: false shortens responses and removes a class of retry-on-parse-failure, which is a second saving.

How do I measure this on my own workflows?

Take one flow and log the input token count of every model call for a week. Then, for each call, estimate what fraction is the system prompt, tool schemas and passthrough payload. The share that is not the actual task is your ceiling for improvement.

Keep reading