Custom AI Agent Development for Business
An AI agent is software where a language model decides its own next steps and calls the tools you give it. It can read a request, look something up in your systems, draft a reply, and hand off to a person when it is unsure.
We build one agent for one task, as a pilot, and measure it against real examples from your business.
Starting at $8,900 for a pilot.
We have not published results from a production agent for a client. That is why this is sold as a pilot, with a written test and a stop point.
What you get
| Part | What it covers |
|---|---|
| Task choice | One task, picked with you, with a clear "done" |
| Written spec | The task, the tools the agent may use, the data it may read, and what it may never do |
| Test set | Real examples from your business with the right answers marked by you |
| The agent | Built and connected to the tools in the spec |
| Human handoff | Rules for when the agent stops and asks a person |
| Logs | A record of what it did, so you can check any action |
| Pilot report | Results against the test set, with a keep, change, or stop recommendation |
Timeline, the size of the test set, and running costs (model fees, hosting) are set in your written quote.
The process
- Pick the task. Narrow beats broad. "Draft replies to quote requests" is a task. "Run customer service" is not.
- Write the spec. You approve it before we build.
- Build the test set. We use your real cases, not invented ones.
- Build and test. We run the agent against the test set and fix what fails.
- Shadow run. The agent drafts; a person approves each action.
- Report. You decide: keep, change, or stop.
Do you need an agent at all?
Often a fixed workflow is enough. Anthropic, a company that builds language models, makes the same point in its guide to building agents.
| Question | Answer | Source |
|---|---|---|
| What is the difference? | A workflow follows predefined code paths. In an agent, the model directs its own process and tool use | https://www.anthropic.com/engineering/building-effective-agents |
| When is simple enough? | Anthropic says to find the simplest solution and add complexity only when simpler solutions fall short | same page |
| What does an agent cost you? | Agentic systems often trade latency and cost for better task performance | same page |
| How should it be tested? | Extensive testing in sandboxed environments, along with appropriate guardrails | same page |
If your steps are known in advance, see AI workflow automation. If you only need answers on your site, see AI chatbot for your website. We say so when a cheaper option fits.
What we measure in the pilot
| Measure | How we count it |
|---|---|
| Correct on the test set | Share of test cases the agent handled correctly, as judged by you |
| Handoff rate | Share of cases passed to a person |
| Wrong actions | Actions that someone had to undo |
| Cost per task | Model fees divided by tasks completed |
| Time per task | Agent time against a person's time on the same sample |
We have not yet measured these on a paying client's production agent. Your pilot report is the first place they appear for your task.
Worked example: what the pilot has to beat
This example uses made-up inputs. A person spends 12 minutes on a task, and their time costs $25 per hour.
Cost per task = 12 ÷ 60 × $25 = $5.00.
Pilot fee ÷ cost per task = $8,900 ÷ $5.00 = 1,780 tasks before the pilot fee alone is covered. Running costs come on top.
Swap in your own minutes and hourly cost. If you do the task a few hundred times a year, an agent will not pay back soon.
What we do not promise
- Savings, revenue, or lead numbers. We have no client results to quote.
- An agent that never errs. Models make mistakes. The handoff rules and logs exist for that reason.
- Free rein. The agent only gets the tools and access named in the spec.
- A finished product. A pilot proves or disproves one task. A wider rollout is a separate quote.
FAQ
What happens if the pilot fails?
The report says so, and you stop. The fee pays for the work, not for a result.
How is this different from a chatbot?
A chatbot answers questions. An agent can also take steps in other tools, such as creating a task. That is why it needs a stricter spec and testing.
Which AI model do you use?
We name the model in the spec and explain why. Model fees are billed as stated in your quote.
What data can the agent see?
Only what the spec lists. We ask for the least access that lets it do the task.
Can it send messages to customers on its own?
Not during the pilot. A person approves outgoing messages in the shadow run.
Pricing
Custom AI agent pilot: starting at $8,900.
The quote lists the task, the test set size, the timeline, and running costs. Book a 15-minute scoping call
Why us
- Pilot first. One task and one report, then you decide.
- Plain test results. You see the cases, not a score alone.
- Cheaper options named. If a workflow does the job, we tell you.
Where our information comes from
- Anthropic, "Building effective agents": https://www.anthropic.com/engineering/building-effective-agents (accessed 2026-10)
- Prices: our service price list of 2026-10-06 (starting prices).
- Worked example: our arithmetic on made-up inputs (12 minutes, $25 per hour, $8,900).
+1 570 955 3538