← The Maintenance Guru
THE MAINTENANCE GURU

Do You Know Your Machine Availability?

Availability isn't just reliability by another name. Getting it right means accounting for maintainability, utilisation, shared infrastructure and how you classify every hour an asset isn't producing — not just how often it breaks.

Reliability, availability and maintainability aren't the same thing

The acronym RAM (or AR&M in the UK) covers three related but distinct properties, defined formally under the European standard EN 13306:2017. Reliability is an asset's ability to deliver its required functions, in a defined context, for a given period. Maintainability is how easily it can be retained in, or restored to, a state where it can deliver those functions, given the right resources and conditions. Availability is the ability to operate when required, in context, assuming the necessary resources are on hand. Acceptable availability generally implies acceptable reliability combined with fast, low-friction maintenance turnaround — you can under-deliver on either half and still end up with an availability problem.

Availability only means something in context

Availability has to be read against how much of that available time an operation can actually use profitably. In a 24/7 operation with strong demand and pricing, near-continuous utilisation may genuinely be worth pursuing — but no complex machine can run forever without maintenance, so the real planning problem is finding windows to release the asset for work with minimum disruption to operations. How much time that maintenance needs depends on the volume of work due, how it's packaged against the schedule, how the individual tasks can be sequenced given physical and spatial constraints, and the staffing and facilities actually available to do it.

Reliability and maintainability trade off against each other in ways that aren't always intuitive. Picture two otherwise-identical, fully sold-out systems: the first has a dominant failure mode that stops production often, but a median recovery time of only 10 minutes. The second fails far less often, but its dominant failure mode takes a median of 2 hours to recover from. Despite being "more reliable," the second system loses more availability — and more revenue — because maintainability is so much worse. Through-life cost has to account for all three RAM components together, not reliability in isolation.

A rarer failure that takes hours to fix can cost more availability than a frequent one you can clear in minutes.

Shared infrastructure and dependency make it worse

Assets that share infrastructure compound the problem: one breakdown can block every other asset relying on the same infrastructure, not just the one that failed — a train stalled on a single track halts every other train on that line until it's cleared or moved to a siding. Dependencies between assets create the same cascading effect more generally: a single failure can make otherwise-healthy downstream assets unavailable too. Redundancy is one mitigation, but it needs investment weighed carefully against the probability that the standby capacity is ever actually needed. Stockpiling finished material downstream of critical assets is another option, though it ties up working capital and carries its own financing cost.

Deciding what actually counts as "available"

Before you can measure availability meaningfully, you need to settle two questions: does planned outage time count against the overall figure, and how is delay time attributed across categories? Most systems exclude planned outage from the headline availability number, which makes the classification scheme behind that number the real design decision. A practical starting point is a defined timeline classification model for how an asset's time is spent — and that classification needs to be genuinely tailored to context. A steam plant, for instance, needs a deliberately slow, controlled start-up and shutdown to avoid stressing the plant with a temperature transient, whereas many other assets have no meaningful ramp time at all. Targets can then be set against each category in the classification and compared to actual performance — but the classification scheme itself has to stay practical enough that the underlying data can realistically be gathered.

How that time gets measured depends heavily on whether collection is manual or automatic. Automatic collection, via SCADA/DCS networks or data historians, can script simple timestamped state-change logging, from which elapsed and cumulative times roll up automatically against the targets. Manual collection, more common inside a maintenance work window itself, is trickier: several tasks might run concurrently, and a delay on a task that isn't on the plan's critical path may not actually affect the overall maintenance duration at all, even though it looks like lost time in isolation.

Measuring the maintenance window itself

Getting a real read on maintenance productivity means measuring individual tasks, then checking those against a Gantt-style view of the planned sequence to understand how delays actually propagate through the whole window — actual execution on the day will often diverge from the plan regardless. A simple front-line system letting staff mark tasks started and finished, with a note field for delays, can capture this without becoming a burden, and works best as a shared improvement exercise (in the spirit of Total Quality Maintenance's Kaizen approach) rather than a surveillance tool imposed on the team.

Task-level measurement also feeds back into planning accuracy: if actual times or resource needs consistently diverge from your estimates, that's the signal to correct the underlying task-list assumptions, which in turn tightens the target times built into your asset availability timeline. None of this is a one-off exercise — it's the ongoing discipline that turns availability from a lagging number you report on into a lever you can actually pull, and it's exactly the kind of fleet-level pattern that shows up clearly once the underlying data is tracked consistently across a fleet rather than reconstructed after the fact.

Ready to see IronMan® in action?

We run live demonstrations and zero-risk pilot deployments — so you can see the value before you commit to anything.