I gave two AI agents my outbound job
What it took to get Argus and Iris running, why I still review their leads, and why I might stop next month.

Recently, Iris, one of the two AI sales agents I built, proposed two leads at an American insurance business and asked me to confirm which automated sequence/campaign I wanted to include them in. I wasn't sure the company was a good fit for Nettle, the startup where I work, so I asked her to explain the reasoning behind the recommendation.
Iris shares a database with Argus, the agent responsible for researching accounts and finding leads. She pulled up his research: the company employed risk engineers who worked with underwriters to set renewal and safety guidelines. That connected the account to the work Nettle helps with. She also showed me which contacts were already in their shared lead book and where they were in the process.
This was the kind of answer I wanted when I started building these agents: enough context to understand a recommendation, challenge it, and make a decision without starting the research again myself.

Outbound kept competing with everything else
Cold outbound is part of my remit at Nettle, where I'm the Chief of Staff. The problem was finding the capacity to do it consistently alongside all my other workstreams, with tasks that often arrive at the last minute. Thirty percent of my week is just firefighting.
Doing outbound properly takes time. You need to understand the company, find the people who might care, work out why they would care, check whether someone has already spoken with them, put them into an appropriate sequence, and ideally craft a message unique to that individual. Then you need to do it again for the next account and its leads.
I wanted that work to happen consistently, to a high standard, with minimal involvement from me. I also wanted a challenging project to learn from. Building something beyond my capabilities at the time, but based on workflows I understood, seemed like a good way to learn.
So I built Argus and Iris.
Giving each agent a job
Argus researches accounts and finds the people worth approaching. Nettle's product is an AI workspace for loss control, so the research needs to establish who owns that work inside an organization. It also needs to identify underwriting leaders who might have an interest in improving it: underwriting benefits from the outputs of loss control and usually owns the budget.
He uses web research, Apollo, and LinkedIn Sales Navigator information extracted through a tool called Evaboot. Researching an account and finding leads takes around eight minutes. I can give him a hypothesis or a list of companies and leave him to work through it over the course of a day, or even several days.
Before a lead reaches Iris, an independent reviewer examines the proposal against our qualification criteria. The reviewer gets fresh context, including the evidence supporting the recommendation, so it can challenge Argus’s conclusion. Ultimately, the reviewer owns that decision, which I can revisit if I want to.
Iris takes the approved leads, checks their records and relationship history, and proposes an audience in La Growth Machine, our outbound tool. During this first month, I've been reviewing each lead at that point before she adds it. The audience determines the campaign/sequence the person enters.
All my interaction with both agents takes place through Slack. Iris notifies me when a lead needs my review or confirmation, as in the screenshot below. This means I can unblock outbound work anytime, from anywhere, as long as I have access to Slack.

The first version needed rethinking
I built two versions of this project. The second was built from scratch and came together much faster, using the lessons from v1. The first attempt had become too rigid. It reached ninety-eight tools built around particular workflows, and it struggled with anything outside those paths.
Even a qualified approval could become a problem. An answer like “yes, but put them into the property audience” contains both a decision and an instruction. I wanted the agent to understand that, without needing me to learn a special way of talking to it.
For the rebuild, I separated the actions the agents could take from the knowledge they needed to decide what to do. Tools handled things like searching, reading a record, or writing an update. Skills described our qualification criteria and how we normally worked. Those were plain documents I could read and change.
This made it easier to inspect the reasoning I was asking the agents to apply. If a lead was being rejected for the wrong reason, I could look at the rule behind that decision rather than treat the answer as something mysterious the model had decided.
I was also building with very limited observability – I was given a server to host the agents, but had no direct access to its production logs or the database. Debugging meant observing agent behavior in Slack, forming a hypothesis, changing the code, and checking whether the next result supported it. I started the v2 plan by establishing what those constraints ruled out.
The reviewer was rejecting almost everyone
In its first week of operation, Argus’s independent reviewer approved almost nobody. That wasn't particularly useful for an outbound system.
One problem was the definition of an underwriting leader. The reviewer was interpreting it as a strict Chief Underwriting Officer test, so people with other relevant leadership titles were being rejected.
We needed a system that recognized the influence a Head of Underwriting or a regional underwriting director could have over the process we wanted to improve, while distinguishing them from someone in underwriting without that responsibility. Where to draw that line depends on the organization’s size and how decision-making power is distributed, so I built that context into the system.
The title was only part of the problem. The reviewer also wasn't consistently receiving the evidence it needed.
The researching agent might have found something useful in a LinkedIn profile, but a fresh reviewer couldn't see that research unless the relevant information was passed to it. Referring to a source wasn't enough when the reviewer only had the message in front of it.
I changed the handoff so the proposal carried the actual excerpts, along with their sources and dates. I also made sure the research and enrichment stages happened before the lead was submitted for review. Missing research should lead to more research or a clear statement of uncertainty, rather than an unsupported conclusion that the person wasn't a fit.
This context challenge took more thought than just telling the reviewer to be less strict. I wanted to build a system that could run almost autonomously, so the reviewer needed to maintain high standards while having enough evidence to make a useful decision.
A good lead can still be the wrong person to contact
Qualification is only one part of the decision. There's also the relationship to consider.
In one case, Iris found that a proposed lead’s company already existed in Attio, our CRM, with prior email history. She couldn't see an open deal, but she paused before proposing an audience and asked whether I still wanted to put the person into outreach.

That was the behavior I wanted. An empty deal pipeline doesn't tell you everything about a relationship. There might be context behind the previous exchange that changes how we should approach someone, or whether we should approach them at all.
Iris gives me the relevant history and waits for my decision. That boundary lets the rest of the work progress while keeping me involved where the information available to the system may be insufficient.
What has changed so far
Argus and Iris now handle the research, qualification, and preparation that feed our outbound. I give them a hypothesis or account list, and they get to work.
We're midway through the first month, and I've reviewed every lead at Iris’s approval stage and rejected about 5%. Argus’s reviewer does the first check; my review has been the last line of defense before enrollment. I haven't had capacity to review the leads Argus’s reviewer rejected yet, but there might be good targets there.
Most of our outreach is through LinkedIn. So far, 6.4% of unique prospects contacted have replied and 3.2% have booked a meeting. I handle those conversations.
The biggest challenge is message customization. I've successfully tested having Iris help with it, but our outbound tool doesn't support a fully automated handoff for this. That leaves another manual step I simply don't have capacity for.
Allowing for thirty days of Gemini-only model usage, model costs plus Apollo, Evaboot, and La Growth Machine come to around US$700–1,000 a month. That excludes the server and my boss's Sales Navigator subscription.
I currently spend about an hour and a half a day on outbound: reviewing leads, handling replies to turn conversations into booked meetings, creating campaigns/sequences for new segments (e.g., “Risk Engineering Leaders in the US”), and occasionally debugging the system. If at least 95% of the leads continue to meet my qualification standard, my plan is to stop reviewing every lead next month and focus that time on messaging personalization.
Let's see what happens, but these are exciting times!