Merlino AI

Running an AI-Native Agency, Day to Day

2026-07-024 min readagent-fleettooling-disciplineagency-operationshiring-economics

The discipline problem nobody wants to admit

Every agency owner I know has the same weakness right now. A new AI tool launches, it looks shiny, and half the team wants to drop what is working to go try it. That instinct feels like innovation. Most of the time it is just distraction wearing a better outfit. Running a fleet of specialist agents only works if the humans running it have more discipline than the tools chasing their attention.

I have said this to my own team directly

This is not a talking point I workshopped for a stage. I pushed a member of my own team to stay focused on Claude Code and agent workflows specifically for agency and build work, and to avoid chasing every new AI tool that showed up, except in the narrow cases where it actually mattered, like image or video generation. The motivation behind that push was plain frustration: I was watching competing agency owners get more out of Claude and team-based agent workflows than we were, simply because they stayed disciplined about the stack while everyone else was still shopping.

That is the real cost of tool-hopping. It is not just wasted subscription money. It is losing ground to competitors who picked a stack, learned it deeply, and kept building on it while everyone else restarted their learning curve every six weeks.

A fleet is only as good as what it is not chasing

Running specialist agents at scale means every agent has a fixed domain: engineering, content, SEO, video, security, and so on. That structure only holds if the tools underneath each domain stay stable long enough to actually master. If Spielberg's video stack changes every time a new video AI launches, nobody ever gets past the beginner phase with any of it. Depth beats breadth when the fleet has to run without me micromanaging every call.

The narrow exceptions matter too. Image and video are named exceptions to the stay-focused rule, on purpose, because those categories move fast enough and differ enough tool to tool that chasing the better output is worth the switching cost. Everything else earns discipline, not novelty.

Hiring is the same discipline problem in a different outfit

I have made this same argument in public, on my own SEO Rockstars conference stage, about a much older question: whether a role is worth hiring for at all. The honest version of that question is not "can we afford this hire." It is whether it is worth spending hours on a deliverable to send to a client who probably does not know and does not care about it. That is an uncomfortable thing to say out loud in front of an audience of agency owners, but it is the same discipline as the tool question. Spend your effort, human or agent, where the client actually feels it. Do not spend it where it just makes the team feel busy.

What this looks like day to day

Running an AI-native agency is not glamorous once you strip the marketing language off it. It looks like:

  • A named agent owns a domain, and that ownership does not shift because a new tool launched this week.
  • New tools get evaluated against a narrow list of categories where switching is worth it, not adopted by default.
  • Every hiring or build decision gets the same test the stage talk described: does the client actually feel this, or are we just spending hours to feel productive.

None of that is complicated. It is just hard to do consistently, because staying disciplined is boring and chasing the next tool feels like progress. The agencies that are actually pulling ahead right now are not the ones with the most tools. They are the ones who picked a stack, gave every specialist a real domain, and stopped restarting.

The Build Log

One email.The whole build.

One email when something ships: the dashboard, the agent, the GMB play, and the prompts and configs that made it go. If it fell over on the first try, I say that too.