How We Built Barney, Superhuman’s Threat Hunter AI Agent
By Thijn Bukkems, Tech Lead Manager, Superhuman.
For a long time, security teams (like Superhuman’s) were constrained by bandwidth to investigate incoming alerts. This created a concerning cycle in which early warning signs (which often showed up as low-priority alerts) were ignored until they led to full-blown incidents.
With AI, we can finally reverse this dynamic. We can use an AI agent to examine all alerts (regardless of severity) with the same rigor, allowing us to close any gaps before attackers can exploit them. This vision inspired our Security Engineering team to build Barney, a homegrown threat-hunting agent that can tackle any investigative problem, whether it’s digging into an incident, triaging alerts, or reviewing access requests. Barney independently does the legwork and comes back with a recommendation (and, in some cases, automatically applies it). While our human engineers are still very much involved in day-to-day operations, Barney lets our team cover more ground by taking on the manual work, allowing us to think more long-term.
Getting Barney to this point wasn’t straightforward. We ran into some classic challenges when using AI for real-world workflows (like navigating context window limitations) and some security-specific ones (like teaching it how to properly balance curiosity and paranoia when approaching security problems). Solving these challenges required iterating on Barney’s architecture and shaping Barney to reflect our team’s culture: curious, problem-oriented, and appropriately skeptical.
How is Barney used today?
We originally built Barney as a simple query tool, one that could answer real security questions by querying our logs, looking up identities, and fetching from endpoint telemetry. As we saw success, we expanded its use from answering questions to serving as a holistic threat hunter. Today, Barney primarily lives in our terminals (though it’s also accessible via Slack) and handles the following tasks:
- Simple queries and tasks: Runs queries on how systems work (e.g., Which of our prod nodes are not monitored by GuardDuty?) and completes tasks with clear outcomes (e.g., Audit our cloud-warehouse admins and flag anyone whose footprint looks anomalous against their stated role). For diagnostic queries (e.g., Is this finding a duplicate?), Barney automatically applies the recommendation.
- Investigate alerts and detections: Investigates new detection end to end. It also runs a pass through our automated triage layer to further triage alerts and flag situations that require the on-call team’s attention.


- Incident investigation partner: When a human engineer is leading a live incident, Barney joins the channel as a working partner and assigns itself most of the investigative legwork (e.g., checking endpoints, gathering scope, drafting comms, tracking remediation status). The engineer can then react to Barney’s output and steer Barney if it goes off track.
- Adjacent skills: Over time, we’ve added a long tail of smaller capabilities, focusing on recurring, time-consuming tasks our Engineering team handles (e.g., reviewing access requests, writing post-incident reports, summarizing detection-tuning options, sweeping a specific class of risk).

Building Barney
Even though Barney does a lot of different tasks, it uses the same loop to accomplish all of them: understanding what needs to be done to achieve a specific task, accessing the right tools/skills to do the legwork, reporting back on the results, and taking feedback from human engineers to adjust direction as needed.
Getting Barney to this point required us to solve a few different challenges and answer some key questions, including:
- Giving Barney capabilities: To be an effective threat hunter, Barney would need access to and training in using different tools. What’s the best way to provide that knowledge?
- Making Barney’s results trustworthy: LLMs are notorious for hallucinating. How do we ensure that we trust its results before we hand off more responsibility to it?
- Teaching Barney judgment: To make the best recommendations, Barney needs to learn our human team’s judgment (which we affectionately refer to as “street smarts”). How do we efficiently give this knowledge to the LLM, especially given constraints on the memory window?
Let’s take a closer look at each challenge.
Giving Barney capabilities
Most of Barney’s capabilities come from a set of approximately 30 command-line tools we wrote, one per source system, and using APIs and prebuilt wrappers to give Barney access to third-party tools (like Slack and Databricks). By using composable command-line tools (rather than bespoke workflows), we can easily chain together commands on the fly to create new multistep investigation workflows, without having to anticipate them from the get-go. Plus, as we add new tools to our security stack, we can easily integrate them into Barney’s workflows without building any bespoke integrations.
But giving Barney capabilities is one thing—teaching it to fully leverage them is another story. For instance, like many security teams, we have logs from various systems (including our mobile device management and endpoint detection and response systems, as well as networking tools). They all contain different information about the host, but no single key links them together, making it difficult to piece together signals across them. To effectively use log data, Barney needs context to interpret these logs.
However, context windows are not infinite. There’s also a productivity tradeoff: Every kilobyte of standing context is a kilobyte of attention pulled off the work in front of the agent. That’s why we split what context lives where so that the most important context (what we call the “always-on” rules) is stored centrally and applied to each session. Meanwhile, we leverage skills for task-specific instructions, with references to detailed how-tos that are loaded whenever a particular skill needs them.
To further improve efficiency, we use background systems to precompute key context—including users, addresses, domains, and devices across all sources—into a context graph, rather than having Barney figure it out on the fly from the raw logs time and time again (and waste precious tokens!). Now, when Barney opens a case, it receives a ready-made packet of everything we know, turning what would be seven tool calls into a single lookup. This has turned out to be one of our highest-leverage investments, as equipping Barney with the right context efficiently not only makes the investigations more efficient but also makes us more confident in Barney’s output.
Making Barney trustworthy
Trust is a major concern when implementing AI agents; with Barney, it was even more so, given our vision for it to be a key player in our security operations. Our approach is simple: We assume that Barney will hallucinate, and we build systems to catch the hallucinations before they can be used to reach a verdict.
We explicitly instruct Barney to treat any AI-generated claims as a junior teammate’s first take, even if it’s from a critical system like our automated triage layer. To actually use an AI-generated claim in its decision-making, Barney needs to re-derive it using primary sources (e.g., a Slack conversation or meeting transcript).
To hold Barney accountable, we have a separate deterministic checker that reviews Barney’s report and verifies that no AI-generated claims are present. We intentionally built this as a separate system because we didn’t trust the system doing the work to verify its output honestly. Our deterministic checker is a simple Python script, and we explicitly chose not to use AI for this, given the risk of reintroducing hallucinations. The Python script works because every tool Barney runs leaves a structured, saved artifact, which the script can easily verify without needing to understand what Barney is trying to do. And it’s working—in a single session, this script caught three fabricated quotes, which Barney was then able to fix.
Improving Barney’s judgment
Even with the right context, tools, and trustworthiness, we have to make sure Barney applies the same judgment as our team in completing its work. However, AI systems are optimized to take the path of least resistance when completing a task, which sometimes leads them to take shortcuts that undermine the effectiveness of the final result. To improve Barney’s judgment, we took a few different approaches:
- Trained on structured prior knowledge: We documented our policies and prior knowledge, synthesizing what we’ve learned and how we work. For instance, we put together a document that summarizes what we’ve actually seen across each detection type: what we’ve learned from past cases, common shapes, and what would actually constitute new evidence. We then feed this data to Barney so it’s working off the same priors as our team.
- Added an investigator checker: A separate investigator checker agent (with no prior context on what Barney is up to) reviews Barney’s report, looking for any gaps in the investigation work. This provides “fresh eyes” on Barney’s work and helps improve the output. The investigator checker is also a gatekeeper—Barney can only consider a task done after this agent confirms there are no further threads for Barney to investigate.
- Codified Barney’s identity: Surprisingly, one of the most helpful changes we made was tying the behavior we wanted to see (e.g., not taking shortcuts in its work) to Barney’s identity. We codified Barney’s identity as a curious threat hunter (with a healthy dose of skepticism) that seeks to fully explore a given problem space. This change has helped Barney understand many intangible skills, such as how much to pressure-test a given claim.

The next chapter for Barney: Self-improvement
With these changes, Barney has become an effective threat-hunting agent that can help us with day-to-day security operations. But what we’re interested in now is getting Barney’s help with bigger-picture problems and solutions. We want to get Barney’s help with questions like, What does our exposure look like for this new attacker technique? or, Is there a hidden pattern in the way this detection family is failing? Doing so requires Barney to learn new skills—but what if Barney could teach itself?
While this is an exciting direction, there are some clear risks. For instance, on a bad day, Barney can talk itself into the wrong lesson, silently weaken a guard it doesn’t like, and drift toward whatever skills it finds easy (rather than what’s useful to our work). But with the right structure and guardrails, we think we can help Barney become more self-sufficient.
The main change has been adding a self-reflection prompt. After each batch of work, Barney runs a reflection exercise, thinking through what failed, what almost failed, what pattern recurred, and what change would have prevented the worst of it. It uses this information to propose changes to its codebase—whether it’s a new skill or a change to how it works—that it thinks can help it work more effectively. Note that there are limits to this, as Barney is not allowed to learn any skills that involve taking actions that impact real people (e.g., shutting down employee laptops or cutting access for a customer).
Before Barney’s changes are merged, a few different checks happen. First, a separate code reviewer agent reviews this output closely to ensure that the changes actually help improve how Barney works. If the code reviewer makes substantial changes, a stop hook is triggered, which blocks shipping any new behavioral logic until those changes are tested and reviewed, preventing us from accidentally shipping regressions. Finally, a human reviews the code before merging it in.
We still have a long way to go before Barney can autonomously improve itself, but early results are encouraging. For example, Barney realized that it could investigate new alerts in parallel while retaining the same depth, resulting in an 80% reduction in time to incident.
Barney’s impact
Barney proved something bigger than any single investigation: that with the right context, guardrails, and judgment, we can hand an entire class of work to an AI system and trust the result.
This proof point has made us excited to apply the same playbook to other security workflows, like detection engineering. A single pipeline now handles detection creation and tuning end to end, saving our team about 40 hours a week. It’s also transforming the role of the detection engineers on the team, who no longer write or tune detections at all; instead, they review the merge requests and the evidence behind them, do the data engineering to bring new surfaces into reach, and threat model the company to feed the system fresh ideas. We’ve also built a data harness to make it easier to launch new agents from scratch.
That free time has allowed us to take on work that moves the needle for our security posture. For example, we’ve now used AI to create a threat intelligence function that we didn’t have the resources to staff before. The AI identifies new entities and themes to monitor, validates intelligence against our environment and logs, and routes completed incidents to our incident response team—while managing and updating its own rules. As a result, we’re able to take on more ground, whether it’s fixing weak configurations in our product stack or building new systems to detect emerging threats. Examples of these systems include a deception system, machine learning models trained on our log sources for GuardDuty-style coverage across proprietary and SaaS data, and a live package-and-extension inventory of our fleet to stay ahead of a surge of supply chain attacks.
___
If you’re excited about applying security to AI problems, come work with us. Apply to open roles on our Security and Engineering teams.