Each model works for a different situation. The expensive mistake is buying one while expecting another, which is most of what goes wrong in this category. Our own engagement shape is described in full if you want something to compare against.
The three models
- The consultancy
- Buys you a map. Senior people assess your operations and produce a prioritized roadmap, a business case, and often a vendor recommendation. Genuinely valuable when the organization needs internal alignment before it can act, or when the decision is going to a board. The failure mode is a deck nobody builds, because the people who wrote it left when the engagement closed.
- The dev shop
- Buys you hands. You bring the specification, they deliver against it, usually per hour or per sprint. Efficient when you already know exactly what to build and have someone internal who can own the spec and the acceptance criteria. The failure mode is a system that matches the spec and misses the problem, which nobody is contractually wrong about.
- The AI-native agency
- Buys you an outcome. The same team that finds the bottleneck builds the system and operates it after launch. Fits companies with budget allocated and no in-house AI expertise, where the diagnosis is the hard part. The failure mode is scope creep and dependency, so the contract needs to say what running-it-for-you means and how you would take it back.
A quick test for which one you need: if you can write the specification yourself today, hire hands. If you cannot, and someone would have to leave your team to write it, buying a diagnosis and a build together is usually cheaper than buying them separately.
Questions that separate them
Ask all six. The answers will sort the room faster than any capability deck.
- "Who operates this in month 6, and what does that cost?"
- Agents are not set-and-forget. Models get deprecated, your workflows change, edge cases surface in production that no audit would have found. A firm that has no answer here is quoting you a build and leaving the operating cost off the invoice.
- "What runs on our stack versus yours?"
- An agent that lives on a vendor platform means a migration to leave. An agent that runs against your CRM, ERP and helpdesk through their own APIs means the work stays yours. Ask where the data sits and what happens to the system if you stop paying.
- "Show me the approval checkpoints."
- Every serious deployment has places where a human signs off. A firm that describes full autonomy from day one has either not deployed into a business with real consequences, or is not telling you where the risk sits.
- "What did you decide not to automate on your last engagement?"
- The most useful answer in the whole conversation. Teams that have shipped can name a workflow they recommended against, and why. Teams that cannot are describing a sales process.
- "Who exactly is doing the work?"
- Ask whether the people in the room are the people building. Ask whether any of it is subcontracted, and where. This is a fair question and the answer is easy to verify later.
- "How is this priced, and what changes the number?"
- Hourly billing rewards the slow path. Outcome pricing needs a defined outcome. Either can be fine, and both need to be legible to you before you sign. How we price and what moves the number is written out.
Red flags
- A fixed quote before anyone has seen your systems
- Scope in this work is driven by how many systems the agent has to touch and how much of your policy is written down. Neither is knowable from a discovery call. A number that arrives before the audit is a number that will move.
- A demo on their data
- Every agent demo works on curated inputs. Ask them to run it against a sample of your real tickets, records or emails, with the messy ones left in. The performance difference is the whole story. What a support agent does and does not handle is a worked example of the distinction.
- Percentage savings with no baseline
- A claim of 60% cost reduction is meaningless without knowing what was measured, over what period, against what starting number. Ask for the denominator. Frequently there is not one.
- A platform sale wearing services clothing
- If the engagement requires their proprietary orchestration layer, you are buying software with implementation attached. That can be the right purchase. It should be a decision you made knowingly.
- No named client references you can call
- Logos on a website are not references. A logo can mean a pilot that never shipped. Ask to speak to the operator who used the thing daily. We will put you in touch on request: team@agentintegrator.io.
What a good engagement looks like
Concretely, on our engagements. Use it as a comparison shape for whoever else you are talking to.
An audit that stands alone
Three weeks, ending in a ranked list of workflows with expected return on each and an honest note on which ones we would skip. The report is yours to keep and act on with anyone, including a cheaper builder. A firm unwilling to sell the diagnosis separately is protecting the build revenue. Ours is step one and buyable alone.
A build you see weekly
Six to twelve weeks, with a working version you can try every week. The integrations, the approval checkpoints and the monitoring are part of the build, so there is no phase 2 where the system finally becomes safe to use.
A deployment that starts supervised
Two to four weeks of shadow mode and supervised operation before anything runs alone, with autonomy granted per workflow against a bar your team sets.
An operating relationship with a number on it
Monitoring, model changes, edge cases, and the monthly review of where the approval boundaries sit. Priced and stated up front, so it is a decision and not a surprise.
Common questions
What is the difference between an AI automation agency and an AI consultancy?
A consultancy sells analysis and a roadmap, then hands it to someone else to build. An agency of our kind sells the diagnosis and the working system together, and stays responsible for it in production. If the hard part of your problem is knowing what to build, buying both from one team removes the handoff where most of these projects die.
Should we hire an agency or build an internal AI team?
Build internally when agent development is going to be a permanent capability and you can hire for it, which is a real commitment at current market rates. Hire an agency when you have a specific set of workflows to fix, want them working this quarter, and would rather not carry the headcount. Many companies do both, using an agency for the first two deployments and hiring against the pattern once it is proven.
How much involvement does an engagement need from our team?
A few hours a week during the audit, mostly from whoever knows how exceptions are handled today, and a weekly demo during the build. The decisions that cannot be outsourced are where approval boundaries sit and what counts as good enough to go autonomous.
What size company does this work for?
We work with mid-market companies, typically 10 to 500 employees, that have budget allocated but no in-house AI expertise. Below that, the coordination cost an agent removes is often small enough that better software solves it. Above it, you likely have an internal platform team and a different set of options.
Keep reading
Buying AI
How AI agent pricing works
Why AI agent work is scoped after an audit, the 6 things that move the number, what you can do to lower it, and what to budget for after launch.
Agents by function
AI agents for customer support
What a support agent can take over, where it breaks, and how to roll one out without damaging the queue. From the team that builds and runs them.
Last updated August 12, 2026. Questions this page did not answer go to team@agentintegrator.io, or start with the audit.

