AI agent development, with limits you control
We build AI agents, multi-agent systems and LLM features that do real work. Permissions, spending limits and an audit trail are enforced in code, outside the model, where it can't argue with them.
Talk to an engineer, not a salesperson.
Founder's certification: Multi AI Agent Systems with CrewAI (DeepLearning.AI, 2023)
AI-01Limits: tools · spend · steps
IllustrationThe problem
Where agents go wrong
An AI agent reads, decides and acts on its own. It calls your tools, writes to your systems and sometimes spends money. That is the value, and it is also the risk.
An agent that can call tools can call them in a loop, spend money nobody approved, or take an action nobody authorised. Prompt instructions are not a security boundary. The controls have to live in the orchestrator, where the model cannot argue with them.
Language models also fail softly. They return something plausible rather than an error, so the system around the model has to catch a wrong answer before it reaches a customer or a database.
Prompts ask. Code enforces.
What we build
What goes into an agent system
Multi-agent systems (CrewAI)
A workflow split across specialised agents, each with one job, its own tools and a clear stop condition. Routing, retries and hand-offs between agents are designed, not left to the model.
Tool permissions
Each agent gets an explicit list of tools it may call, enforced by the runtime rather than described in a prompt. Write access to production systems stays behind a confirmation step until the behaviour is proven.
Spending and step limits
Hard ceilings on tokens, API spend and steps stop the run when they are reached. The check sits outside the agent loop, so a confused agent can't talk its way past it.
Memory and state
Task and conversation state is stored durably. An interrupted agent resumes where it stopped instead of starting again.
Audit trail
Every decision, tool call and payment is logged with its inputs. You can reconstruct what the agent did, and why, after the fact.
Human hand-off
When confidence drops below a threshold, the work goes to a person with the context attached. The escalation path is part of the design, not an admission of failure.
Evidence
What we can show you today
We don't have a published AI-agent client case study yet. This is the evidence we do have.
Not an AI project. We show it for how we work: the agreed scope, and an existing codebase understood quickly.
“Excellent freelancer to work with. The implementation was handled professionally, communication was clear, and the work followed the requested MVP scope properly. The freelancer understood the existing setup quickly…”
Multi AI Agent Systems with CrewAI · DeepLearning.AI · July 2023
Ratings, reviews and engagement dates come from client contracts delivered through Upwork, checked 24 Sep 2026. '…' marks where a quote was shortened. Client names stay private.
LLM applications
LLM application development
Not every product needs an agent. Many need one careful model call inside a normal app: pull fields from a document, sort a support ticket, draft a reply, answer from your own data. We build these as engineering projects, not a prompt behind a text box.
- Structured outputs: validated against a schema, so the next system gets typed data, not prose.
- An evaluation set: so a prompt or model change is measured instead of argued about.
- Retrieval over your data: designed around the questions the system has to answer.
- Cost control: token budgets per customer and per workflow, enforced on the server.
- Observability: traces, spend and quality drift, because error monitoring won't catch a wrong answer.
Detail
Which model? Whichever fits the task, the latency budget and the cost ceiling. We build behind a provider abstraction, so switching models is a configuration change, not a rewrite.
Automation
AI automation pipelines
Automation fails at the edges: a duplicate trigger, a source that changes its format without warning, an output that is confidently wrong and lands in your system of record anyway. We keep the predictable steps as plain code. The model handles only the step code can't, and anything uncertain goes to a person.
- Safe writes into ERP, CRM and databases you don't control, so a retry doesn't create a duplicate record.
- Validation and confidence gates that send low-confidence work to a person, not downstream.
- Fact-checking workflows with source attribution and an explicit "uncertain" state.
- Monitoring of throughput, error rate, escalations and cost. A system that runs unattended has to report on itself.
Agent payments
Agent payments with x402
x402 lets an API charge per request. The API answers "402 Payment Required" with a price. The agent pays, then retries with proof of payment. That way an agent can pay for what it uses without an account or a card.
Autonomous payment is a spend-control problem before it is a protocol problem.
Detail
The hard parts are replay protection and idempotency across retries. A retried request must never pay twice, and a call that fails after payment needs a refund path.
Agent reputation
Agent reputation with ERC-8004
Before one agent hires another, it needs to know who it is dealing with. ERC-8004, a draft Ethereum standard called "Trustless Agents", defines on-chain registries for agent identity, reputation and validation.
Process
How an agent project runs
Four steps. You get something you can read at the end of each.
Discover
We map the workflow, the decisions the agent makes, and what it must never do.
You get
A written brief.
Architect
We set tools, permissions, spending limits and stop conditions, and choose a model per task.
You get
An architecture doc, the key decisions (ADRs) and a milestone plan with estimates.
Build
We build the limits first, then the agent behaviour inside them.
You get
Regular demos and written updates.
Validate and hand over
We test loops, exhausted budgets, failing tools and bad inputs, and score quality against the evaluation set.
You get
UAT sign-off and handover notes.
At handover you also get the evaluation set and a written description of what the agent is allowed to do.
Pricing
How pricing works
We start with a paid discovery sprint, then fixed-price milestones. Ongoing work can run on a dedicated engineer by the month, and small jobs can be hourly.
How many tools and systems the agent touches, whether labelled evaluation data already exists, and whether the agent moves money. Model usage is billed by your provider, on your account.
- Discovery sprint
- Fixed-price milestones
- Dedicated engineer, monthly
- Hourly
Start here
Not ready to talk?
Scope your project
9 short steps. See your result before you share an email.
Project assistant (guided)
Answer guided questions. Get a technical brief an engineer reads.
Questions
Questions we get about agents
How do you stop an agent looping?
Termination conditions, step ceilings and budget limits enforced by the orchestrator. When a limit is reached, the agent stops and escalates. It doesn't degrade quietly.
How do you stop the model making things up?
You constrain rather than eliminate. Schema-enforced outputs, retrieval grounding, confidence thresholds, and a human escalation path for anything below threshold. Any claim beyond that would be dishonest.
Can agents act on our production systems?
Only inside an explicit permission set. We normally keep write access behind a confirmation step until the behaviour is proven against real traffic.
Can you work with our existing AI prototype?
Yes. Taking a prototype that works in a notebook and giving it what it needs to run unattended is a good first project: schema enforcement, evaluation, cost control and observability.
Three ways to start
Planning an AI agent project?
Pick whichever suits you. A person reads every message.