K
KowalAI

GuidesDiagnose

Rank AI work that pays before you buy more tools

An AI ROI audit, in this house, is not a compliance review. It is not NIST, SOC 2, or a policy binder. It is an operational ranking: which recurring work is worth changing, which can wait, and which should stay human. You do that with a short inventory, a baseline for each item, and a score across frequency, pain, data readiness, and blast radius. Then you split the list into DIY now, build later, and do not automate. The $997 assessment is a live consultation and a written analysis that does this ranking with you. It does not include implementation.

Who this is for

This playbook is for owner-operated companies where the owner still sits in leads, inbox, reporting, or delivery. It is also for the operations lead who has been asked to “look at AI” and now has a pile of tool demos, no shared list of the actual work, and no way to say no. If you already know the leak — weekly reports, follow-up, documents, or scattered knowledge — start there. Use this guide when the problem is the ranking itself: too many asks, no order, and a budget about to buy another seat.

It is not for teams that need a security attestation, a model-risk program, or a vendor questionnaire answered for a regulated buyer. Those are real jobs. They are a different product. Mixing them with “which weekly report should we stop rebuilding” produces a document nobody can act on.

Symptoms

If several of these are true in the same month, this is a live operational leak — not a tooling preference.

  • Leadership has approved an AI budget, but nobody can name the first workflow it is supposed to change.
  • The same demo gets booked three times because each department brought a different pain.
  • A coordinator still rebuilds a weekly pack from exports while a unused dashboard sits in the same stack.
  • Two tools already draft replies, and neither has an owner, a review rule, or a log.
  • The CRM, the inbox, and the spreadsheet disagree on the same pipeline number.
  • Someone proposed an agent that would email customers, move money, or update the system of record without a human in the loop.
  • The owner can list five leaks in a hallway conversation and cannot say which one to touch this month.

What done looks like

Done is a working operating change, not a purchased seat or a dashboard nobody opens.

  • A one-page inventory exists: work, owner, frequency, source systems, and who is hurt when it slips.
  • Each item has a baseline in plain language — not a fabricated hour count, a described current path.
  • Every candidate has a score for frequency, pain, data readiness, and blast radius.
  • The list is split into DIY now, build later, and do not automate, with a reason on each row.
  • Anything that touches customers, money, contracts, or systems of record has a named reviewer and a fallback.
  • The next purchase, if any, is tied to one named workflow instead of a generic “AI stack.”

This is ranking, not a compliance audit

Search results will try to sell you frameworks, maturity models, and policy templates. Those can be useful later if you sell into enterprises that ask for them. They do not tell an owner whether to stop rebuilding the Monday pack, put a next-step field on every opportunity, or leave invoicing alone. Operational ranking answers a narrower question: given the work we already do, what is cheap to change, what is expensive if it fails, and what is not ready because the data or the process is a mess.

Keep the language boring. “Audit” here means look at the work and write down what you see. It does not mean grade the company. It does not mean produce a score that pretends to be science. If a number would have to be invented — hours saved, percent lift, payback in weeks — leave the number out and describe the path instead.

Start with a baseline, not a tool shortlist

Before you score anything, write the current path. Who starts the work. What they open. What they copy. Who they wait on. What “done” means today. What breaks when that person is out. You are not estimating hours you never measured. You are making the path visible enough that two people in the business would recognize it.

A usable baseline for one item looks like this: “Every Monday, ops exports four CSVs, pastes them into a workbook, and emails a two-page summary the owner reads before the 10:00 standup. If ads or the CRM are late, the pack waits. If the owner is traveling, the pack still goes out and nobody uses it.” That is enough to rank. You do not need a time-and-motion study.

Do the same for pipeline follow-up, weekly reporting, and operational knowledge. Those three show up in almost every owner-operated company that asks for “AI.” They are also the ones people try to skip by buying a chatbot. A chatbot on top of an unread report, a CRM with no next step, and a folder of stale PDFs does not create a return. It creates another place to look.

Score frequency, pain, data readiness, and blast radius

Use four questions. Keep the scale small — high, medium, low is enough. If you need a spreadsheet, use one column per question and a fifth column for the decision. Do not average them into a fake composite that hides a deadly low score.

Frequency

How often does this work happen, and does it happen on a clock the business already respects? Daily inbox triage and a Monday report beat a quarterly planning packet. Frequency without a decision attached is just busywork. Ask who waits on the output.

Pain

Who feels it when the work is late, wrong, or skipped? Owner time, a stuck customer, a missed follow-up, and a report leaders do not trust are different kinds of pain. Write the name of the person or role, not “the business.” If nobody can name the person, the pain is probably theatrical.

Data readiness

Can a competent person already assemble the answer from systems that mostly agree? If the CRM, billing, and delivery tools tell three stories, you do not have an automation candidate. You have a definition problem. Data readiness is also about access: if the only copy lives in one person’s laptop, you are not ready.

Blast radius

What happens if the new path is wrong? A wrong internal brief is recoverable. A wrong customer email, a wrong invoice, or a wrong CRM overwrite is not a DIY afternoon. High blast radius does not mean “never.” It means the item cannot sit in DIY now unless the output is a draft and a human sends it.

A row that is high frequency, high pain, high readiness, and low blast radius is a DIY candidate. A row that is high pain and low readiness is a later build or a cleanup, not a purchase. A row with high blast radius and no reviewer is do not automate.

Split the list: DIY now, build later, do not automate

Three buckets. Not five. Not a matrix poster.

DIY now

Changes a person on the team can make with tools you already pay for: a report definition, a next-step field, a shared inbox rule, a template, a prompt that produces a draft, a checklist. The $997 written analysis is built to name two to four of these. You run them. Nobody installs a new system inside that fee.

Build later

Work that needs integration, an agent, a portal, a dashboard, or a data join. It may be the right work. It needs its own scope and its own quote if you want someone else to build it. Put it on a roadmap with the dependency named: “cannot build the client portal until the onboarding stages are written down.”

Do not automate

Judgment calls, one-off exceptions, anything where the source of truth is still an argument, and anything that would send money terms, legal language, or system-of-record changes without a person. Leaving work alone is a result. Write why so the next vendor demo cannot reopen it with a slide.

A worked ranking, without invented numbers

Imagine an owner-operated service firm with four loud asks: a Monday pack rebuilt from exports, inbound that sits in a shared inbox, proposals started from a blank page, and a drive full of stale decks. Do not assign hours or a return percentage. Score the four questions you already have.

The Monday pack is high frequency, high owner pain, mixed data readiness (two systems disagree on cash), low blast radius if the first version is an internal brief. Decision: DIY the definitions and one push brief; do not buy a dashboard until the two cash numbers are one sentence.

Inbound is high frequency, high pain, decent readiness if the form already writes somewhere, high blast radius the moment a draft can send itself. Decision: DIY capture, next-step field, and reviewed drafts. Do not automate selling.

Proposals are medium frequency, high pain on the owner’s calendar, readiness only if a current rate card exists, high blast radius on money terms. Decision: DIY one template and a review rule. Build later if several sellers need the same locked shell.

The drive is high pain for new hires, low readiness, medium blast radius if a chat layer can see client files. Decision: DIY a top-20 inventory and archive. Do not buy a chatbot this month. That is ranking. It is also how you stop a vendor from selling you the chatbot first.

What people usually get wrong

They start from the tool shortlist. Then every meeting becomes a bake-off and the work never gets written down. Reverse it. If a vendor cannot speak to one row on your inventory, they are selling a seat, not a change.

They invent a baseline so the slide looks quantitative. A guessed “eight hours a week” becomes a strategy. When the number was never measured, it cannot be used to justify a purchase. Describe the path. If you later measure a real cycle time, write the method next to it.

They treat “do not automate” as failure. It is often the most expensive-looking row and the cheapest decision. Pricing judgment, exception-heavy collections, and one-off custom work stay human until the exceptions are rare enough to name.

They skip blast radius because the demo was a draft. Drafts become sends. Writes become overwrites. If you cannot name the reviewer and the undo, the item is not DIY now.

DIY vs hire

DIY the ranking. One owner or ops lead can run this in a week of calendar time if they already sit in the work. Interview the people who assemble the report, chase the lead, or hunt for the latest proposal language. Write the inventory. Score in a room with the people who will have to live with the decision. The output is a ranked list, not a software project.

Hire help when the company cannot agree on the list, when the owner is too inside the work to see it, or when the asks span sales, delivery, and finance and nobody owns the cross-cut. The $997 assessment is the paid version of that ranking: a live call and a written analysis. It is not an install. Later implementation — integrations, agents, dashboards, document workflows — is custom-quoted only if you want us on it.

If you want the shape of that written analysis, the Willow sample on /examples is a format sample, not a client story and not a result you should expect. Use it to see how a ranked list and a short report look on the page.

Controls before automation

Recommendations that touch customers, money, contracts, private data, or systems of record need human review, limited permissions, a log, and a fallback. Do not automate a broken path because the tool is ready.

  • No production send of customer, money, or contract language without a named reviewer.
  • No write-back to the CRM, billing system, or file-of-record until the field map is written and a person can undo it.
  • Drafts need a visible source: which note, which row, which document the text came from.
  • New tools inherit the same permission the human already had — not a wider key “so the agent can finish.”
  • Keep a fallback path that still works when the model, the zap, or the vendor is down.
  • If two systems disagree, stop. Fix the definition before you automate the number.

How this sits next to the other playbooks

Ranking is the front door. The rest of the fourteen mid-market AI asks are the rooms. If reporting is the leak, read stop rebuilding weekly reports by hand. If revenue is going quiet after the first touch, read follow up before good prospects go cold. If the team cannot find the current rule, read build an operational knowledge base your team can trust. Those guides assume you already decided the work is worth touching. This one is how you decide.

If you want that ranking done with you, on a live call, and written down so the team can run the first two to four items themselves, book the assessment. Bring the messy list. Leave with an order.

Soft next step

If ranking this work would help, start with the assessment.

$997 is a live consultation and a written analysis. You leave with 2–4 improvements you can run yourself, plus a later roadmap quoted only if you want us on it. Implementation is not included.