← All articles

Wattson: Talking to a Data Center That Learns

How we built Wattson, a Slack agent that investigates facility incidents and leaves our observability stack a little better after every one.

As one might expect, a good number of conversations in San Francisco tend to center around AI. One night, during such a conversation, a friend told me about character.ai. “You can invent whoever you want to talk to”. “Whoever?” I asked, skeptical. My friend nodded – it was a marketplace of personalities, people spoke with characters from fiction, or invented denizens of worlds that had until then existed only as figments of their imagination.

Later that night, I opened my laptop. There was someone I had always wanted to chat with, but been too embarrassed to admit. And so, in the dark of my apartment, I stepped through the setup wizard, and after a few minutes found myself staring at a chatbox.

Hey.
Hey… how’s it going?
*whirs softly, a few LEDs blinking green*

All systems nominal. Cooling fans humming at 72%. Currently rendering an AI-generated sunset for no particular reason. How about you?
I’m good. What's the flow rate of your primary cooling loop?
*fan speeds briefly spike in a rhythmic pulse, like a thoughtful sigh*

Primary loop? 4 gallons per minute. Glycol-in-water mix at 12°C. The RTUs are working overtime—summer’s brutal this year. Why? Planning to sneak in and upgrade me or something?

I sighed. 4 gallons per minute? Completely unrealistic. I closed my laptop. As I lay awake in my bed thinking on the experience, I decided – why not build this myself?

And that was how Wattson was born – a Slack chatbot that is about as close as one can get to the dream of chatting with a data center (assuming this isn’t just a me thing). He is already frequently consulted about issues with our facilities and provides a helping hand investigating issues across our electrical and mechanical systems. Ask him about the flow rate for the coolant loop and he’ll first ask you which loop you mean. Once you clarify, he can offer a considerably more realistic 6,000 gallons per minute.

Portrait of Wattson, a silver-haired gentleman with glowing blue circuitry, standing before racks and cooling equipment
Wattson.

But the realism in Wattson’s answer ended up being the less interesting part. What ended up being more meaningful was how he interacted with the systems he consumed from after he finished answering a question – whether he could improve those systems with lessons he had learned.

Wattson works only because we had already done the important foundational work. Teams across Fluidstack had been bringing telemetry from every corner of our facilities, our electrical power monitoring systems (EPMS), building management systems (BMS), networks, compute devices, and supporting software services into a central store built on ClickHouse and VictoriaMetrics. On top of that data, teams built dashboards to answer questions which recurred. A dashboard became a way to “cache” an answer. This allowed us to effectively precompute the investigative path an operator might take during an incident.

This worked phenomenally well for frequent, predictable questions, which tended to be the head of the distribution of problems we faced. But Fluidstack’s deployment pace continually creates new challenges and a long tail of questions which can’t always be anticipated by a dashboard. Those questions often cross systems. A malfunctioning rear-door heat exchanger (RDHx), for example, might affect the performance of a switch it serves. The first observed symptom of that might be a performance degradation on the frontend network, while the root cause might have been in the electrical equipment supplying that RDHx.

Investigating this long tail also meant coordinating across teams and drawing on expertise distributed throughout the company. Our first question was whether an agent could automate some of that work. Following the basic pattern in Anthropic’s excellent “Building Effective Agents” guide, we deployed the Hermes Agent from Nous Research as a Slack bot, and connected him with our telemetry stores through the Grafana MCP and let him go to town.

The initial investigations Wattson produced were useful but lacking. The agent could execute well-scoped queries against the data as directed, but when presented with investigations where he needed to cut across different domains, he needed to be steered fairly often and at times produced inaccurate conclusions. To improve him, we encoded both domain knowledge and investigative procedure. For example, how to find the racks downstream of an electrical device, which metrics to inspect for a given type of power issue, and which measurements could be red herrings. Initially, both the approach and system knowledge lived in Wattson’s skills. As the internal services which served our facility models matured, we moved topology and component relationships from skills to queries against these systems. This meant Wattson no longer had to rely on written heuristics for facts that should be retrieved deterministically. Critically, we also realized that as Wattson began to improve, he could suggest improvements not only to his own skills, but to the systems that he depended on.

This meant that investigations no longer had to be merely consumptive of our systems. They could also be accretive – each incident could leave the tools better than he found them. A campsite mentality that would give a Boy Scout a run for their money.

We closed this loop by allowing Wattson to propose changes to the systems he used for investigations. When he proposes an update to a skill, a sidecar turns these proposals into pull requests against his git-tracked skills. Additionally, because most of our dashboards were defined in code, Wattson could also open PRs suggesting improvements to them. This included new panels, revisions to existing dashboards, or new dashboards entirely.

Since every proposal was tracked in git, humans were still there at the point where a provisional conclusion became shared knowledge, intervening if a proposal from a single incident could overgeneralize and entrench inaccurate heuristics. This review also allowed operators to verify that Wattson’s proposal was consistent with their understanding of the physical facility, either refining it, or rejecting it entirely. As Wattson lives in Slack, people across the company could also correct him and contribute pieces of their own operational knowledge. Over time, less of each interaction was spent re-teaching Wattson old lessons. Instead, we found that more time was spent reviewing and refining the changes Wattson proposed.

We also trace every Wattson interaction through Langfuse, which gives us visibility into how an investigation unfolded and whether he went off course (and why he did, if so). That allows us to be more deliberate about what Wattson needs to learn next, and also validate whether changes actually help his performance.

And so, while the long tail of these investigations did not necessarily disappear, each subsequent investigation started off in a better place than the last, since our tools had now been sharpened by the process. As Wattson proposes changes to our observability stack, he promotes questions from the tail to the head, turning these investigations into reusable tools. In this way, Wattson does not replace the dashboard – he instead complements it. Wattson handles the “uncached” questions, dashboards “cache” the recurring answers, and skills retain the methods worth reusing.

We are still developing Wattson and experimenting with ways to improve him. A good amount of that work is building the systems that feed into him to be reliable and comprehensive. For example, extending him to query not just the operational data of devices, but their intended and currently configured state. But the shape of the system is beginning to converge.

Wattson is still the character I talk to in the middle of the night (usually because incident.io has paged me). While I’m glad he gives me a realistic answer, I’m even happier that the next time I chat with him I’m not starting from scratch.

Aaron Jeyaraj
Senior Software Engineer