Choosing an AI automation agency, and what the first ninety days should produce
Wobble is an AI automation agency in Pakistan, based in Karachi and working with businesses here and abroad. We install AI employees, automations and dashboards inside your own accounts, then keep them running month to month. The rest of this page covers how to choose any agency, us included.
You have decided to hire someone. What is left is telling a good engagement from an expensive one before the money is committed, and knowing what should exist by the time the second invoice arrives.
A retainer comes in two shapes, and mixing them is the usual mistake
An AI automation agency in Pakistan is normally hired on a retainer that covers two separate things, building new systems and keeping the existing ones running. Mixing those two into one number is the usual mistake. Ask what should exist by day 30 and by day 90, in objects you can point at, before the second invoice arrives.
Where this has actually run. Wobble has run 25 client engagements across six countries, and every one is published with its outcome rather than described in general terms. See the work, with the numbers.
There is a build retainer and there is a run retainer, and they are priced differently because they are different work. A build retainer buys a stream of new capability: this month a lead response flow, next month the quoting process, the month after the reporting layer. A run retainer buys the system continuing to work: monitoring, fixes, small changes, accuracy checks and a monthly account of what happened.
Trouble starts when one fee covers both without saying so. Six months in, the building slows because most of the obvious work is done, and the client is paying a build rate for what has become maintenance. Or the opposite: a low monthly fee that was really a maintenance number gets asked to absorb a significant new build, and the agency starts avoiding your emails.
The fix is to separate them on paper from the start. State what the monthly fee covers as running, state how new build is scoped and priced, and agree what happens when the build queue empties. An agency that has run engagements past the first year will already have a view on this. One that has not will tell you it is all included, which is the answer that causes the argument later.
- Running: monitoring, error handling, small changes, accuracy sampling, the monthly report
- Building: new workflows, new integrations, anything that changes what the system does
- Stated in the agreement: how a new build is requested, estimated and approved
- Stated in the agreement: what happens to the fee when there is nothing left to build
Day 30, day 90, day 180
Automation engagements go wrong slowly, and by the time it is obvious a lot of money has been spent. Milestones you can check against are the cheapest protection available, and any agency worth hiring will agree to them before starting.
By day 30, nothing impressive should exist and several boring things should. A written map of how work actually moves today, produced by somebody who watched people do it rather than by interviewing a manager about it. A shortlist of what will be automated, each with a reason. An explicit list of what will not be, which is the item that tells you whether they were thinking. Access resolved, which sounds trivial and eats more of week two than anyone plans for. And your baseline numbers recorded.
By day 90, one workflow should be live, carrying real volume, with a human checking a sample of its output. Not a demonstration. Live, in the hands of the staff who will use it, with the awkward edge cases found. If day 90 arrives and the only thing to look at is a prototype, that is the moment to ask hard questions, not month six.
By day 180, the first workflow should be boring and a second or third should be running. Boring is the goal. The monthly report should be showing the same measure you baselined at the start, and somebody should be able to say what broke since launch and what changed as a result.
The day 90 question
Ask exactly this on the first call: what will be live and carrying real volume ninety days after we start? Write the answer down. An agency that answers with a demonstration rather than live usage has told you how the engagement will go.
Where automation scope goes wrong
Scope here fails differently from scope on a website project. It fails because the ground moves under it. The process changes mid-build when a client changes a form or a supplier changes an export. A platform changes a rule and something legal on Monday needs redesigning by Friday. And the process turns out to have three undocumented exceptions that only the person doing the job knew about.
WhatsApp is the clearest example in this market, because the assumptions people bring to it are consistently wrong. Messaging through the WhatsApp Business API runs on opt-in, on message templates approved in advance, and on a messaging window that closes after the customer's last reply.
A scope line that says send customers a reminder is not a small task, it is a template that has to be written, submitted and approved, plus a consent position, plus a design for what happens when the window has closed. Agencies that have not shipped on that platform routinely price it as though it were an email.
Written scope survives this better when it describes outcomes and names the constraints, rather than listing features. State what the workflow must achieve, what it must never do without a human, which systems it touches, and what the agreed process is when reality turns out to differ from the map. A change process everyone agreed to in advance costs an hour. One invented mid-project costs a relationship.
- Describe the outcome and the exceptions, not a list of features
- Name the human approval points explicitly, especially anything touching money or pricing
- Agree in advance how a discovered exception gets priced and scheduled
- Get platform constraints such as template approval into the timeline, not the assumptions
Questions to ask before signing, and what a good answer sounds like
Most buyers ask about tools and price. Neither predicts how the engagement goes. These do, and the quality of the answer matters more than its content.
Who operates this after launch, and are they on this call? A good answer names a person and their hours. A weak answer describes a team. Ask what happens when that person is on leave.
What will you refuse to build? A good answer is a specific example, told with slight irritation, because it actually happened. A weak answer is a policy statement about ethics.
Show me something that broke and what you changed afterwards. This is the most useful question on the list. Anybody who has operated systems for real has a story here. Anybody who has only built demonstrations will change the subject to a success.
What do you need from us, and how many hours a week? A good answer is uncomfortable and specific: access, a decision maker who replies inside a day, and somebody who knows the process free for a few hours in the first fortnight. An agency that needs almost nothing from you is describing a project that will drift.
- Who operates it after launch, and what hours do they cover
- What will you refuse to build, with an example
- What broke on another engagement, and what changed afterwards
- What do you need from us, in hours per week
- What is in the monthly fee, and what is billed separately
- What do we hold if this ends badly in month four
What an honest monthly report looks like
The monthly report is where a retainer earns trust or quietly loses it. A report containing only good news is marketing, and experienced buyers discount it entirely.
The useful version has four parts. What the system did, in volume. How well it did it, measured against the baseline recorded at the start rather than against last month. What went wrong, including the things nobody outside the agency would have noticed. And what is planned next, with the reason.
The section that matters most is the second one, because it is the one that can embarrass the agency. If nothing in a monthly report has ever been uncomfortable to write, the measurement is not real. Ask for the accuracy sample: how many outputs were checked, how many were wrong, and what kind of wrong they were.
When to wait rather than start
If nobody on your side can spend a few hours a week for the first month, do not start. The map of how work moves cannot be produced without your people, and an agency that agrees to build without it will build against an imagined process and discover the real one during launch, at your cost.
If your busiest season starts in six weeks, wait. Installing a new system into the middle of peak trading means training staff while they are under pressure, and the first inevitable defect lands on the worst possible day. Build in the quiet months, run in the busy ones.
If a single existing product would solve the problem, buy the product. Plenty of what gets scoped as custom automation is a booking tool, a shared inbox, or turning on a feature already included in a system you pay for. An agency that never recommends this has an incentive problem, and a buyer who never asks has an expensive habit.
Common questions
What does an AI automation agency do month to month?
In a running month it monitors what the system is doing, handles failures, makes small changes as the business shifts, samples AI output for accuracy, and reports on volume and results. In a building month it also scopes and ships new workflows. Separate the two in the agreement so neither quietly subsidises the other.
What should be live ninety days after we start?
One workflow carrying real volume, used by the staff it was built for, with a person sampling its output. Not a prototype and not a demonstration. Before that, by day 30, expect a written map of how the work moves today, a list of what will and will not be automated, access resolved, and your baseline numbers recorded.
Should we pay a fixed project fee or a monthly retainer?
Fixed pricing suits a scope both sides genuinely understand, which usually means a second or third workflow rather than the first. Monthly suits discovery and operation. The common arrangement is a paid diagnosis, a phased build, then a monthly operating fee, each priced separately so you can stop one without losing the others.
How do we know the retainer is still earning its keep in month six?
Compare the monthly report against the baseline recorded before the build, not against last month. Ask for the accuracy sample: how many outputs were checked by a person, how many were wrong, what kind of wrong. Ask what broke and what changed as a result. A retainer where nothing has ever been uncomfortable to report is not being measured.
Who on our side needs to be involved?
A decision maker who can answer within a day, and at least one person who actually does the work being automated, available for a few hours a week in the first month. After launch you need one internal owner who notices when the system misbehaves. Engagements without that internal owner decay quietly, whatever the agency does.
What should we prepare before the first meeting?
Two weeks of simple counting: enquiries received, median time to a first reply, how many got no reply, and hours per week spent moving information between systems by hand. Also a list of the tools you pay for and who holds the admin logins. That preparation shortens the diagnosis and stops you paying somebody else to count for you.
See where this applies to your business
The AI Readiness Call is a short, free conversation about where automation would actually pay back in your business. The call is free. The diagnosis is not.
Book AI Readiness Call ↗