Ask five people on a plant floor what “maintenance is working” means and you’ll get five different answers. Two numbers cut through that: MTBF, mean time between failures, and MTTR, mean time to repair. Track both and a plant manager can tell, in minutes, whether a maintenance program is actually reducing breakdowns or just getting faster at reacting to them.
What MTBF actually measures
MTBF is the average running time between one failure and the next, measured on a specific asset, a filler, a case packer, a specific CNC spindle, not the plant as a whole. A rising MTBF means the asset is failing less often. A flat or falling MTBF, even with a fast repair team, means the plant is still absorbing the same number of unplanned stops, just recovering from them quicker.
What MTTR actually measures
MTTR is the average time from “this machine went down” to “this machine is back in production,” including diagnosis, parts, and the repair itself. A short MTTR is good, but it can hide a bad MTBF. A team that gets very efficient at 20-minute repairs might never ask why the same conveyor motor keeps failing every 10 days.
Why the two numbers only mean something together
A plant chasing MTTR alone optimizes for reaction speed. A plant chasing MTBF alone can under-invest in the repair process itself. The plants with the best real uptime track both, and use the gap between them to decide where the next maintenance dollar goes: toward preventing the failure, or toward repairing it faster when it happens.
Where the data for both numbers actually comes from
Neither metric means much if it’s built from a maintenance technician’s memory of “about how long that usually takes.” Nulogy Maintenance pulls MTBF and MTTR directly from the same event stream that tracks downtime and OEE, so a failure gets timestamped the moment the line actually stops, not when someone logs a work order later in the shift. That single source of truth is what makes the two numbers trustworthy enough to act on, and it’s the same real-time visibility into plant floor operations that powers OEE and downtime reporting. See how shortstops feed into the same downtime record that maintenance work orders draw from.
From reactive to preventive to predictive
Most plants start with reactive maintenance: fix it when it breaks. The next step is preventive, maintenance on a fixed calendar, which helps but treats a heavily used asset the same as a lightly used one. The most mature step is predictive, or condition-based, maintenance, where work orders trigger from the asset’s actual signal data rather than a date on a calendar. Read the full comparison of reactive, preventive, and predictive maintenance for how to know which mix fits your plant today.
What good looks like on a QBR
A plant manager reviewing MTBF and MTTR quarter over quarter should be able to answer three questions without pulling a spreadsheet together first:
- Which three assets have the worst MTBF this quarter?
- Has MTTR improved or gotten worse on those same assets?
- Are we opening more preventive work orders on those assets than we were last quarter?
If those answers require a week of manual reporting, the maintenance program is being run on instinct, not data. A supervisor can shortcut most of that reporting by simply asking Nora which asset has the worst MTBF this week, rather than building the report by hand.
What this changes for a maintenance lead’s day
The practical value of tracking MTBF and MTTR by asset isn’t the quarterly review, it’s the daily decision. A maintenance lead who can see MTBF trending down on one specific conveyor motor, in real time, can open a work order that week instead of waiting for that motor to cause a full line-down event that finally gets everyone’s attention.
A plant that can see MTBF and MTTR by asset, in real time, is a plant that can finally have the “fix it or replace it” conversation with numbers instead of gut feel.
To learn more, see how real-time OEE tracking works inside Smart Factory.