Domain-Specific AI Agents: Why Composition Beats One Big General Agent

Table of contents
Every Boomi and n8n workflow I’ve built that tried to do everything eventually became the workflow nobody wanted to touch. The one that synced products, chased stock levels, matched carriers and handled exceptions all in one flow: a schema change three steps upstream, and three unrelated things broke at once.
I’m watching the same pattern show up in AI agents, just faster. A single agent wired to every tool, every internal doc and every system a company owns hits the same ceiling a mega-workflow does, not because the model behind it got worse, but because nothing that overloaded stays reliable for long.
What is happening
An agent, in the useful sense, is deterministic software wrapped around a non-deterministic model: code that constrains what an LLM can do, so its output turns into a reliable action instead of a guess. MCP servers and markdown-based “skills” are the current default way to hand an agent its tools and knowledge, and the default habit has become handing it all of them. One agent, hundreds of tools, every internal doc pasted into its context, on the theory that more access means more capability.
It doesn’t. It means more ways for the model’s attention to fragment across things that don’t matter for the question in front of it. I built Nexus deliberately narrow for this reason: its MCP tools are search, search_slack, search_code, who_knows, each one LLM-free and single-purpose. An orchestrating agent decides which to call; no single tool tries to be everything.
The alternative to stuffing one agent with everything is composing several agents that each do one thing well:
flowchart TB
accTitle: Inheritance versus composition
accDescr: Inheritance wires one general-purpose agent to every tool and every document, so its attention fragments across an overloaded context. Composition splits the same work across a coordinator and several narrow, domain-specific agents, each with its own small, task-specific context.
subgraph INH["Inheritance: one overloaded agent"]
direction TB
A1[General-purpose agent]
T1[CRM tools]
T2[File system]
T3[Support docs]
T4[Billing API]
T5[Code repos]
T6[...dozens more]
A1 --- T1
A1 --- T2
A1 --- T3
A1 --- T4
A1 --- T5
A1 --- T6
end
subgraph COMP["Composition: a coordinator and specialists"]
direction TB
C[Coordinator agent]
S1[Billing agent<br/>billing API only]
S2[Support agent<br/>docs + tickets only]
S3[Code agent<br/>repos only]
C --> S1
C --> S2
C --> S3
end
INH ~~~ COMP
The top half is one agent trying to hold everything in its head at once. The bottom half is closer to how a delivery team actually works: a coordinator who knows who to ask, and specialists who only ever see what their job requires.
| Inheritance (general-purpose) | Composition (domain-specific) | |
|---|---|---|
| Context | Bloated: every tool and doc, all the time | Narrow: only what this task needs |
| Failure mode | Attention fragments; the model guesses at tool parameters | An agent stays inside an explicit, small boundary |
| Scaling | Adding tools makes everything less reliable | Adding agents is additive, not multiplicative |
| Permissions | One agent that can technically do almost anything | Each agent’s blast radius is only its own domain |
Why it matters for integration work
Picture an “operations agent” wired to sync product feeds, watch stock levels, reconcile carrier data and draft replies to customer exceptions: the AI equivalent of the one Boomi process that does everything. It works fine in the demo. Then a marketplace changes an attribute schema, and the same agent that used to draft calm exception replies starts hallucinating a stock number because the model’s attention is split seven ways.
Split that into a feed agent, a stock agent, a carrier agent and a support-reply agent, each with only the tools its job needs, and the schema change breaks exactly one of them, the same way a well-scoped Boomi process or n8n workflow fails in one place instead of everywhere at once. This isn’t a new idea for anyone who has designed integrations for a living. It’s separation of concerns, applied to agents instead of workflows.
The anatomy of a domain-specific agent
A narrow agent isn’t just a shorter prompt. It’s four layers, each doing one job:
flowchart TB
accTitle: The domain-specific agent stack
accDescr: A domain-specific agent has four layers. The model and system prompt define its role and boundaries. The tool layer holds deterministic functions, sub-prompts for narrow cognitive tasks, and delegated sub-agents. The hook layer manages side effects like history mutation and time injection. The sandbox layer gives the agent an isolated file system and code execution environment.
M["Model + system prompt<br/>role and boundaries, nothing else"]
T["Tool layer<br/>functions · sub-prompts · sub-agents"]
H["Hook layer<br/>history mutation · time injection"]
FS[Filesystem]
CE[Code Execution]
M --> T --> H --> FS
H --> CE
classDef solid fill:#111111,stroke:#111111,color:#ffffff;
classDef outline fill:#ffffff,stroke:#111111,color:#111111,stroke-dasharray: 3 3;
classDef modelNode fill:#111111,stroke:#e8386d,stroke-width:2px,color:#ffffff;
class T,H solid;
class FS,CE outline;
class M modelNode;
- Model + system prompt. The “life objective”: a role that’s specific enough to say what the agent should refuse, not just what it should do.
- Tool layer. Deterministic functions for real actions (an API call, a file write), narrow sub-prompts for small cognitive sub-tasks, and the ability to delegate an entire piece of work to another domain-specific agent instead of trying to do it itself.
- Hook layer. The plumbing that keeps an agent deterministic over a long session: mutating message history, and time-injection, which quietly adds a message with the current date/time so the model has state without the user having to say “by the way, it’s Tuesday.”
- Sandbox. An isolated file system and a sandboxed place to execute code. ChatGPT and Claude already do this under the hood for things like generating a PDF. It’s worth treating as a first-class requirement for any agent you build yourself, not an implementation detail.
The economic case
Efficiency stops being optional once “cheap intelligence” stops being the default assumption:
- Token efficiency. A narrow agent isn’t re-reading an entire project’s context on every call. It only holds what its one job needs, which is where most of the token savings actually come from.
- Cost. A small, well-scoped model handling one narrow task is dramatically cheaper per call than routing that same task through a flagship model. The exact multiplier depends on the task, but the direction holds consistently across every narrow-vs-flagship comparison I’ve seen.
- Capability control. A general-purpose agent with broad tool access is a hard sell to your own security team, because it can technically be prompted into almost anything. A domain-specific agent’s tool list is its permission boundary: there’s no broad access to accidentally misuse.
- Scaling. Narrow, isolated agents parallelize cleanly. A dozen instances of a single-purpose feed agent running at once is a much smaller operational problem than a dozen instances of one agent that can touch everything.
For anything customer-facing, this isn’t a nice-to-have. Unless the customer relationship is worth a lot per interaction, routing every message through a flagship, broad-context model is not a cost structure that survives contact with real volume.
Where this is heading
Vercel’s open-sourced eve framework is the clearest public signal I’ve seen that this shift is already underway: agents defined with a specific role, domain instructions kept as separate “skills” rather than baked into one always-on prompt, and tools scoped per agent rather than shared across everything. That’s composition, shipped as a framework rather than argued as a theory.
The interesting engineering problem stops being “how good is this one agent” and becomes “how well do these agents coordinate”: a coordinator that knows which specialist to call, rather than one agent that tries to be every specialist at once.
Delegation isn’t a separate mechanism bolted on top. It’s just one more entry in the tool layer: an Agent tool that hands the whole task to another domain-specific agent’s own full stack, rules, hooks, tools, system prompt, model and sandbox included. That agent can have an Agent tool of its own, and the chain continues:
flowchart LR
accTitle: How one agent calls another
accDescr: A coordinator agent's tool layer includes an Agent entry, a delegation mechanism that hands a task to another domain-specific agent's own full stack of rules, hooks, tools, system prompt, model and sandbox. That agent's own Agent tool can delegate further, forming a chain instead of one agent holding every capability itself.
subgraph CARD1["Coordinator agent"]
direction TB
R1[Agent Rules]
H1[Hooks]
T1["Tools: functions + prompts"]
AG1[Agent]
S1[System Prompt]
M1[(Model)]
SB1[Sandbox]
R1 --> H1 --> T1 --> AG1 --> S1 --> M1 --> SB1
end
subgraph CARD2["Feed agent"]
direction TB
R2[Agent Rules]
H2[Hooks]
T2["Tools: functions + prompts"]
AG2[Agent]
S2[System Prompt]
M2[(Model)]
SB2[Sandbox]
R2 --> H2 --> T2 --> AG2 --> S2 --> M2 --> SB2
end
subgraph CARD3["Compliance agent"]
direction TB
R3[Agent Rules]
H3[Hooks]
T3["Tools: functions + prompts"]
S3[System Prompt]
M3[(Model)]
SB3[Sandbox]
R3 --> H3 --> T3 --> S3 --> M3 --> SB3
end
CARD1 -.delegates to.-> CARD2
CARD2 -.delegates to.-> CARD3
classDef solid fill:#111111,stroke:#111111,color:#ffffff;
classDef outline fill:#ffffff,stroke:#111111,color:#111111,stroke-dasharray: 3 3;
classDef modelNode fill:#111111,stroke:#e8386d,stroke-width:2px,color:#ffffff;
class R1,H1,T1,S1,R2,H2,T2,S2,R3,H3,T3,S3 solid;
class AG1,AG2,SB1,SB2,SB3 outline;
class M1,M2,M3 modelNode;
The compliance agent in this chain is a leaf: it has no Agent tool of its own because it doesn’t need to delegate further. A pricing question never reaches it, and it never sees the feed agent’s context, because nothing routes that way.
What to do about it
- Scope agents the way you’d scope an integration. One job, one small tool list, one clear boundary: the same discipline you’d already apply to a Boomi process or an n8n workflow.
- Design the tool layer on purpose. Treat an agent’s MCP tools as a contract, not a junk drawer. If a tool doesn’t serve this agent’s one job, it doesn’t belong on this agent.
- Sandbox before you connect real data. An isolated file system and execution environment should exist before an agent touches production systems, not after something goes wrong.
- Budget for orchestration, not just for a bigger model. Plan for a coordinator plus several small, cheap calls to replace what used to be one expensive call to a flagship model.
Key takeaways
- One agent wired to everything fails the same way one workflow wired to everything fails: quietly, and everywhere at once.
- A domain-specific agent is four layers (model, tools, hooks, sandbox), not just a shorter prompt.
- Narrow agents are usually cheaper, not just safer, because they skip the cost of re-processing context they never needed.
- The near-term shift to plan for isn’t a smarter single agent. It’s more agents, coordinating.


