All articles
AI

How We Run Our Agency on AI Agents

Inside the production AI agents that run our agency: a 20-tool operations assistant, sandboxed design and code agents, and the guardrails that make them trustworthy.

Bhaskar Roy
Bhaskar Roy
Founder, cazywebJuly 17, 2026 9 min read
Share
How We Run Our Agency on AI Agents

Most agencies "use AI" the way most people do: someone drafts copy in a chatbot, pastes it into a doc, and calls it a workflow. We went a different way. Over the past year we rebuilt our own operations platform - the system that runs our projects, clients, reports and invoices - so that AI agents are part of the workforce inside it. Not demos. Production systems that our team and our clients touch every day.

This post is the honest tour: what the agents actually do, how they are wired, and the guardrails that turned out to matter more than the models. If you are evaluating whether agents belong inside your own business, this is the field report we wish we had a year ago.

The shape of the thing

Our platform is a full agency operating system: a project-management tool, client portals, a comms app, analytics, and invoicing. The agents are not bolted on top; they live inside it, with the same permission system as the humans. There are several in production today:

  • An operations assistant that reads and writes project data on request.
  • A design agent that produces A/B test variation mockups against a client's real site.
  • A code agent that turns an approved design into ship-ready variation code.
  • An email agent that designs on-brand lifecycle emails and pushes them to the client's ESP as drafts.
  • A report agent that writes post-test narratives around numbers computed by code.
  • A finance assistant that turns plain English into draft invoices.

Each one follows the same philosophy Anthropic describes in their engineering guidance: simple, composable loops with tightly scoped tools beat sprawling autonomous frameworks1. Here is what that looks like in practice.

The assistant that does real work

The operations assistant is a function-calling agent with exactly 20 tools: six that read (list projects, search tasks, fetch a task), nine that write (update status, assign, comment, create tasks, schedule meetings), and five that query our analytics warehouse and save dashboard widgets. It runs a plan-act-observe loop with a hard cap on steps, so it can chain "find the overdue tasks on this brand, bump their priority, and comment why" into one request - and it cannot spiral.

Three design decisions did most of the heavy lifting:

  • Permissions live below the agent, not in the prompt. Every write tool re-checks the signed-in user's permissions at the data layer before mutating anything. If the model decides to edit something the user cannot edit, the database boundary says no. We never rely on the prompt to enforce access.
  • The agent has zero filesystem or network access. Its world is the 20 declared tools, nothing else. No shell, no web, no files.
  • Every mutation is audited. Writes emit the same real-time events and activity-log entries a human edit would, so anyone can see exactly what the agent changed and when. There is also an org-level kill switch so any client can turn AI off entirely.

Agents that design and ship experiments

The most interesting pipeline is the one that produces A/B test variations for our CRO program. It is a relay of specialists rather than one clever agent:

  1. A scraper (no AI, just a headless browser) captures the client's live page: full-page screenshots at desktop and mobile sizes, the design tokens, the stylesheets, and a static snapshot of the control page.
  2. The design agent works inside a sandboxed session folder containing that context. It has exactly four tools - read, search, and write files in that folder - and its job is one artifact: a side-by-side mockup where the control is the real page, cloned faithfully, and the variation differs only by the change we are testing.
  3. A human reviews. Versions iterate. Nothing proceeds until a person approves the design.
  4. The code agent takes the approved design and produces the CSS, the JavaScript, and a setup guide for the testing tool. It too is sandboxed to its session folder. Its output goes through an explicit approval step before the link is attached to the task.

The sandbox is the point. These agents cannot browse the web, run commands, or touch anything outside their staged folder. Everything they need is pre-fetched into the folder by deterministic code. That inversion - move the risky work into plain code, keep the agent's world small - is the single biggest reliability win we found.

The email agent writes structure, not HTML

Email is where brand fidelity and deliverability collide, so we split the job. The email agent designs the campaign as a structured content model - sections, copy, which product images go where - and a deterministic template engine renders that model into ESP-safe HTML. The agent literally cannot emit broken markup, and it can only reference images from the brand kit we scraped; any other image URL is rejected.

When the client is on Klaviyo, the finished flow is pushed over as drafts only. The integration has no ability to send or schedule anything. A human activates the flow inside Klaviyo. That one property - draft-only outputs - is what let clients say yes to it.

Reports: AI writes prose, never numbers

Our post-test reports and weekly program reports are part-generated by AI, with a strict division of labor: all metrics, significance calculations and revenue figures are computed by code, frozen, and handed to the model, which is instructed to narrate only the numbers it was given. The report agent runs with no tools at all. The same rule applies to our finance assistant: it drafts an invoice as structured data for a human to review; it never issues or sends anything.

This split matters because the failure mode of a fluent model is not silence, it is confident invention. Take the numbers away from it and that failure mode disappears.

The unglamorous parts

A few things nobody puts in launch posts, but which made the system dependable:

  • Injection fencing. Client briefs and scraped content are wrapped as data with explicit instructions that nothing inside them is a command. Agent inputs are untrusted by default.
  • Monitoring. Every agent run is wrapped in failure alerting, and a daily probe verifies our AI credentials so a dead token surfaces within a day instead of at a client deadline.
  • Provider independence. The whole stack routes through one switch: Claude by default, with a Gemini fallback we can flip on with a single config change. No rewrite required to change providers.

What we learned

Benchmarks and demos overstate what agents can do unsupervised; the research community has been blunt that agent evaluations rarely reflect production constraints like cost and reliability2. Our experience agrees. The agents that earn their keep share a profile: one narrow job, a small set of constrained tools, deterministic code doing the risky work, and a human gate wherever a mistake would be expensive.

If you want the deeper, cited version of this - the tool-loop architecture, the security model, the failure modes - we wrote it up properly in our research paper on tool-using LLM agents in production workflows.

And if you are wondering what an agent like this would look like inside your own business - reading your systems, doing your repetitive work, gated by your rules - that is exactly what we build for clients. Talk to us and bring the workflow that eats the most hours.

References

  1. 1.Schluntz, E., & Zhang, B. (2024). Building Effective AI Agents. Anthropic. https://www.anthropic.com/engineering/building-effective-agents
  2. 2.Kapoor, S., Stroebl, B., Siegel, Z. S., et al. (2024). AI Agents That Matter. arXiv preprint. https://doi.org/10.48550/arXiv.2407.01502

Related research

Want a team to run this for you?

See how we help

Get the free AI-CRO Implementation Guide

The playbook we use to turn ecommerce traffic into revenue. Straight to your inbox.

Work email only. No spam. Unsubscribe any time. Privacy.