← The Maintenance Guru
THE MAINTENANCE GURU

AI in maintenance analytics: where LLMs help, and where they get it wrong

Every maintenance software vendor has an AI pitch this year. Fewer are willing to talk about what happens when you point a large language model at the kind of data most maintenance teams actually have — incomplete, inconsistent, and mostly typed by hand under time pressure.

The AI pitch every maintenance team is hearing

Predictive maintenance vendors have been applying machine learning to sensor data and failure histories for years — that's genuine, mature ground. What's new over the last couple of years is a second wave: generative AI and large language models (LLMs) being layered onto maintenance software as chat assistants, promising to let an engineer "ask their maintenance data anything" in plain English.

Those are two different technologies solving two different problems, and treating them as interchangeable is where teams start making poor tool choices. A model trained to classify a failure mode from sensor readings is answering a narrow, well-defined question. A large language model asked to reason freely over a plant's maintenance history is doing something much less constrained — and much less predictable.

Where large language models genuinely help

Used narrowly, LLMs are a real improvement on tasks that were previously manual and tedious: turning inconsistent free-text technician notes into structured categories, summarising a long work-order history for a shift handover, drafting a first pass at job instructions, or translating site-specific shorthand into plain language. These are bounded tasks. The output is easy for a human to sanity-check in seconds, and the cost of an occasional imperfect answer is low — a re-worded sentence, not a wrong engineering decision.

Where it gets dangerous: fluent answers built on shaky data

The core failure mode of a large language model is that it's optimised to produce fluent, plausible-sounding language — not to know when its answer isn't actually grounded in reliable evidence. That matters enormously in maintenance, where more than 95% of transactional data is entered by a human, in the field, often under time and safety pressure. That's exactly the kind of data that's full of gaps, inconsistent shorthand, and judgment calls — the same issues we covered in an earlier post on data quality.

Ask an LLM why a particular pump keeps failing, over a sparse and inconsistent maintenance history, and it won't necessarily tell you the data can't support a confident answer. It's liable to construct a plausible-sounding root cause anyway, because sounding plausible is what it was trained to do. That's a materially different failure mode from a wrong statistical estimate. A wrong estimate usually comes with a visible confidence interval attached. A hallucinated narrative usually doesn't come with any visible warning at all — it reads exactly as confidently as a correct one.

A model that sounds confident is not the same as a model that is right — and inconsistent maintenance data rarely gives it the grounds to be either.

Why deterministic methods still earn their place

Established reliability engineering methods — Weibull analysis, FMEA, RCM, structured defect-classification taxonomies — remain valuable precisely because they're transparent about their own uncertainty. A Weibull fit on eight run-to-failure events will visibly show wide confidence bounds. A root-cause analysis built on an incomplete work-order trail will visibly show gaps in the timeline. None of that transparency is guaranteed with a generative model unless it's deliberately engineered in, and most off-the-shelf LLM tooling doesn't do that engineering.

This isn't an argument against AI in maintenance. It's an argument for choosing the right kind of model for the job, and for treating any generative output as a first draft a qualified engineer reviews — not a finished answer.

The actual fix is upstream: data quality before model choice

The highest-leverage fix here isn't a bigger or newer model — it's the same recurring gaps we've written about before: incomplete notifications, uncaptured emerging work, and weak data integrity control at the point of entry. Any model, generative or otherwise, inherits whatever the underlying records actually contain. Investing in cleaner inputs pays off no matter which analytical method sits on top of them. Investing in a flashier model on top of the same messy inputs mostly just produces more confident-sounding versions of the same uncertainty.

How IronMan® applies AI differently

IronMan® uses machine learning models purpose-built and validated for specific, bounded maintenance tasks — classification, anomaly detection, defect elimination — rather than an open-ended chat interface asked to reason freely over incomplete records. A qualified engineer stays in the loop for judgment calls, and the platform is explicitly built to work with the messy, human-entered data that already exists, rather than assuming clean inputs it's unlikely to get.

AI has a real role in maintenance analytics. The practical question for any team evaluating a tool isn't "does it use AI" — it's what kind of model is actually doing the work, and what happens when the data it's fed doesn't fully support the answer.

Ready to see IronMan® in action?

We run live demonstrations and zero-risk pilot deployments — so you can see the value before you commit to anything.