From maintenance plans to AI agents
From 2023 to 2025 I worked as a reliability and maintenance engineer in the energy sector. The work was RCM, FMEA, and risk assessment: working out how equipment can fail, what that failure would mean, and what maintenance is worth doing about it. Since 2025 I've been building AI agents and data systems for the same industry. On paper that looks like a career change. To me it feels more like the same job with different tools.
This note is about the habits I brought with me, and why I think they matter for anyone building agents for plants and equipment.
Every answer needs a source
An FMEA is only useful if each line can be traced back to something real: failure history, a manufacturer's recommendation, an engineer's documented judgment. A maintenance strategy doesn't get approved because a spreadsheet said so.
I hold AI to the same standard. If an agent tells a planner how long a job will take, the planner should be able to see why. That's the idea behind the Maintenance Duration Estimator: it runs retrieval over past work orders to estimate how long a maintenance job will take, and it cites the similar past jobs it based the estimate on. The estimate is useful. The citations are what make it something you can actually check.
Know when to ask a person
A lot of what keeps a plant running isn't written down anywhere tidy. It lives with the operators, technicians, and engineers who work there. An agent that guesses when it's missing that context is worse than no agent at all, because it's confidently wrong in a place where wrong is expensive.
ARIA is my attempt at the opposite. It uses MCP tools to do its work, but when it doesn't have the context it needs, it reaches out to real people to get it before it acts. For an agent in an industrial setting, asking isn't a failure mode. It's the feature.
Some decisions aren't the model's to make
Reliability work teaches you that decisions are not all equal. Deciding which report to read first is low stakes. Deciding whether safety-critical equipment can keep running is not. Risk assessment exists to separate those two kinds of decision and treat them differently.
Agents should respect the same line. They can gather the data, run the analysis, draft the recommendation, and show their working. The call on anything safety-critical stays with a qualified person. I'd rather build a tool that makes that person faster and better informed than one that tries to replace their judgment.
Make the maths usable
Weibull analysis is one of the most practical tools in reliability engineering. It takes failure data and turns it into an answer to a plain question: when should this equipment be replaced? But it only helps if the people making that decision can run it and read the result.
So I built Weibull Analysis, a browser-based tool for life-data analysis on energy assets. It runs in the browser, so there's nothing to install. Reliability maths, made usable.
The same goes for AI. A clever model behind a confusing interface doesn't help anyone on a plant. That's why I care about the unglamorous end-to-end parts: the data coming in, the pipeline in the middle, and the screen someone actually uses at the end.
What it adds up to
Show your sources. Ask when you don't know. Leave safety-critical calls to people. Build tools that people can actually use. None of that is new to reliability engineers. It's just new to a lot of AI software.
That's the kind of agent I want to build, here in Abu Dhabi and for the industry I came from. If you're working on something similar, I'd like to hear about it.