Skip to main content

Command Palette

Search for a command to run...

Your first AI agent should be boring

A scoring framework for picking the workflow to automate first

Updated
•5 min read•View as Markdown

Most teams pick their first AI agent the wrong way. They pick the most impressive workflow.

"Let's have an agent run our whole outbound motion." "Let's have it write and publish all our content." These make great demos and terrible first projects. They're high-stakes, hard to evaluate, and when they go wrong, they go wrong in public.

The teams that get real value out of AI automation almost always start somewhere much less exciting. Here's a framework for finding that boring, high-value first workflow.

Score every candidate on five things

Grab a list of the recurring tasks your team does. Then score each one from 1 to 3 on these five dimensions.

1. Frequency

How often does this happen?

  • 3 — daily or more

  • 2 — weekly

  • 1 — monthly or less Agents compound. A task that runs every morning pays back the setup time within days. A quarterly task might never pay it back.

2. Clarity of "done"

Could you explain what a good result looks like to a new hire in two minutes?

  • 3 — yes, and you could check their work at a glance

  • 2 — mostly, with a few judgment calls

  • 1 — "you know it when you see it" Fuzzy success criteria are the number one reason agent projects stall. If you can't evaluate the output, you can't improve it, and you can't trust it.

3. Cost of a mistake

If the agent gets it wrong, what happens?

  • 3 — nothing much; it's internal or easily fixed

  • 2 — mildly embarrassing, caught before it spreads

  • 1 — a customer, a payment, or a legal commitment is affected Note: you can often raise this score by putting the risky step behind a human approval. An agent that drafts customer emails but waits for your OK before sending is a 3, not a 1.

4. Data availability

Does the information the agent needs already live in your tools?

  • 3 — it's all in apps you use (CRM, inbox, helpdesk, billing)

  • 2 — mostly, with some tribal knowledge to write down

  • 1 — it lives in people's heads

5. Pain

How much does your team dislike doing this?

  • 3 — people actively avoid it

  • 2 — mildly tedious

  • 1 — people kind of enjoy it This one sounds soft, but it matters. Automating a task people hate earns immediate goodwill. Automating a task someone takes pride in earns resistance.

Add it up

Anything scoring 12 or above out of 15 is a strong first candidate. Here's how a few common workflows tend to score:

Workflow Freq Clarity Mistake cost Data Pain Total
Morning lead triage from CRM to Slack 3 3 3 3 3 15
Chasing overdue invoices (with approval) 3 3 3 3 3 15
Tier-1 support ticket triage 3 2 2 3 3 13
Weekly status report to the team 2 3 3 2 3 13
Fully autonomous outbound sales 3 1 1 2 2 9
Publishing blog posts with no review 2 1 1 2 2 8

Notice the pattern. The winners are unglamorous: sorting, chasing, triaging, summarizing. They happen constantly, they're easy to check, and the downside of a mistake is small.

Then write the instruction like you'd brief a person

Once you've picked the workflow, the next mistake is under-specifying it. "Handle our leads" is not an instruction. Compare it with:

Every weekday at 8am, pull new leads from HubSpot created in the last 24 hours. For each one, look up the company, estimate headcount, and check whether they match our ideal customer profile (B2B SaaS, 20–500 employees). Post the ones that match in #sales-hot with a one-line reason. Don't contact anyone.

That's a job a new hire could do on day one. It's also a job an agent can do on day one. Trigger, inputs, decision rule, output, and a boundary.

Run it in the open for two weeks

Don't flip it on and walk away. For the first couple of weeks:

  • Keep approvals on for anything that leaves your team.

  • Read the activity log every day or two.

  • Note every time you edit or reject an output, and why. Those notes become the fixes to your instruction. By week three you'll usually know whether to loosen the approvals, expand the scope, or pick a different workflow.

Then, and only then, get ambitious

Once one boring agent is quietly saving someone an hour a day, you've earned the right to try something bigger. You'll also know what good looks like, which makes the next one much faster.

If you want more candidates to score, DeskFerry published a list of 25 real AI agent use cases across sales, support, marketing, finance and ops, each with its trigger and a worked example. It's a good source for filling in the table above. And if budget is part of the decision, their breakdown of what AI agents actually cost in 2026 covers pricing models and the hidden costs people tend to miss.

Start boring. It's the fastest way to get somewhere interesting.


Disclosure: I work on DeskFerry, a no-code platform for building AI agents that run across the apps you already use. The framework above works regardless of which tool you pick.