General Patton is written in history as one of the most brilliant generals and also as the one who said: “A pint of sweat will save a gallon of blood”, referring to the fact that perfect maintenance and prior preparation of troops and strategy before combat were the key to success. That maxim can also be applied to IT when we talk about computer maintenance, saving us from fires at 3 a.m. and the subsequent traumatic Vietnam flashbacks.
Because we have all been through them: the main server goes down in the early hours, the phone will not stop screaming, the coffee turns even more bitter and the team in charge makes that face in front of the terminal when it understands that it could have been avoided. It almost always can.
Because the difference between nights of chaos and peaceful mornings is computer maintenance, knowing which type should be applied, when and why.

What is computer maintenance?

The theory, as is usually the case, is not difficult. Computer maintenance is the set of actions, both planned and reactive to an event, that ensure the proper functioning of an organization’s IT systems.
We are talking about hardware, software, networks and infrastructure in general, but this is where the nuances come in, and that is where the important part is usually hidden.
Computer maintenance is not simply “fixing what breaks”, but actively managing the lifecycle of all our resources so that they are available, secure and performing optimally.
The key is to understand the difference between the reactive and proactive approach when we talk about that maintenance.
Reactive acts when the damage has already been done, but proactive works so that the damage does not happen. It is that pint of sweat that will save us the gallon of blood.
That said, we must treat it as an ideal to aspire to because, in day-to-day operations, no team will manage to live exclusively in one of the two worlds, especially the proactive one.
However, the maturity of an IT department is measured, to a large extent, by how far it has managed to shift maintenance into that more preventive territory, which would make General Patton proud.

The four types of computer maintenance

Having seen the general picture, let us increase the microscope zoom to dissect each form of computer maintenance.

1. Preventive maintenance

This type of computer maintenance is the one that happens before the problem appears.
Its goal is to reduce the likelihood of technical failures (and human heart attacks) through scheduled actions, such as:

  • Software updates.
  • Hardware checks.
  • Physical cleaning of equipment.
  • Cooling system checks.
  • Backup management.
  • Effective patch management, etc.

The result when it is done properly?
Fewer incidents, lower repair cost and, above all, fewer sleepless nights wondering why we did not become plumbers.
Well-executed preventive maintenance is the best thing you can be in IT: boring. Boring because nothing explodes, nothing interrupts operations and the infrastructure simply works.
But of course, preventive maintenance requires planning and discipline, two words that are also boring, but in IT they are the difference between thriving and barely surviving.
This includes tasks such as:

  • Implementing regular system reviews.
  • Establishing update schedules.
  • Documenting asset status…

These are tasks we explored in depth when discussing preventive maintenance in IT and they will not impress anyone in our Tinder bio, but who wants love when you can avoid trouble?
The way to measure this maintenance is, above all, with the time elapsed without incidents and their trend, which should be decreasing or stabilized at a very low level.

2. Corrective maintenance

Corrective maintenance is responding to the failure that has already occurred.
Depending on the nature of that failure, maintenance may consist of repairing the equipment, restoring the service, replacing the component, remediating the incorrect configuration…
No matter how many times we read Patton’s biography or The Art of War, it is impossible to avoid this type of computer maintenance because nothing is infallible and because, however preventive our approach may be, the song is always right and “life gives you surprises”.
Likewise, the key in this maintenance is not so much running around like headless chickens as the quality of the diagnosis.
An IT team that quickly reaches the problem but does so without context or tools may take hours to do what should take minutes.
That is where the following become important:

  • Remote access tools and what they allow us to do without physically going to the heart of the beast.
  • RMM (Remote Monitoring and Management) systems, which can provide that critical context.
  • Incident management solutions, which facilitate coordination and mitigation actions.

MTTR (Mean Time To Repair/Recovery) is the indicator that will best reveal how efficient we are in corrective maintenance. The lower it is, the better we are performing when unexpected events occur.
However, the key for elite IT teams is to go one step further and turn corrective maintenance into future learning that protects us from similar failures.
To do this, good root cause analysis after each incident is necessary, and it must be documented, shared and used for continuous improvement. Another of those boring terms that you only appreciate when you find yourself like Germany facing Patton, with a thousand open fronts and all of them in retreat.

3. Predictive maintenance

If preventive maintenance is prudent and corrective maintenance is inevitable, predictive maintenance is the most sophisticated of all.
Here we continue with that notion of going further and becoming elite operators through predictive maintenance that is based on system monitoring and analysis of the generated data to detect anomalies before they become failures.
The ideal thing for this would be to have the Enterprise computer from Star Trek, a system that monitors every parameter of every component in real time and alerts the crew before something fails.
With something like that, the goal is to star in the most boring episode in the world and for that computer not to say: “The warp engine has exploded”, but rather: “The warp engine will explode in four hours if no action is taken”.
That is predictive maintenance.
And even if we do not have Starfleet technology, that is fine, because professional tools such as Pandora FMS come into play here and make that predictive maintenance possible.
The key is constant visibility, performance metrics, well-configured thresholds, automatic alerts and the ability to correlate events. Connecting those pieces in advance is what will grant us prescience over what will happen, like Paul Atreides.
Thus, a disk starts generating read errors before failing completely, a server CPU spends weeks at anomalous values before collapsing, network latency rises steadily with no apparent cause… all of these are signals.
Capturing them and acting on them before the service goes down is the heart of predictive maintenance.
And also what most reduces long-term operating cost. Something to remind management of, in case they do not want to invest in tools that help us.

4. Evolutionary maintenance

Evolutionary maintenance is the least urgent of the four, truth be told, but that does not mean it is not important.
This type of maintenance consists of updating, improving and adapting IT systems to the organization’s new needs, technological evolution and changes in the environment.
It is difficult to cover everything it involves, but it includes, for example:

  • Migrations to new infrastructure.
  • Adoption of new platforms.
  • The scalability of systems.
  • IT change management in a structured way.

The key is not that everything works today (that too), but that it can work in the future with room for growth. Because if the infrastructure cannot scale with business growth or adapt to constant technological change, then the system has a bomb ticking under the chair, because time forgives nothing and neither do unexpected events.
Good evolutionary maintenance is a commitment to the continuous relevance of IT infrastructure and to giving it the ability to adapt without (too many) traumas.

Comparison between the types of maintenance

Let us remove that microscope zoom and now open the perspective to a bird’s-eye view, to recap the four types of maintenance we have seen.

Type

Goal

Time of action

Operational impact

Relative cost

Required IT maturity

Preventive

Avoid failures

Before failure

High (reduces incidents)

Medium

Medium

Corrective

Restore service

After failure

Variable (depends on MTTR)

High in emergencies

Low

Predictive

Anticipate failures with data

Before failure

Very high

Medium

High

Evolutionary

Improve and scale

Planned

Strategic

Variable

High

How to define an effective IT maintenance strategy

The question is this: the theory sounds great when preached from the pulpit, but the enemy is tenacious and is always preparing an offensive in the Ardennes.
Well, let us put the stripes back on and talk strategy to make that theory real.
No organization can live exclusively from a single type of maintenance, so the question is not which one to choose, but:
“In what proportion do we combine them and how do we prioritize them?”.

The first step in a computer maintenance strategy

The first thing is perhaps the most difficult for a human being (and also for an IT technician): knowing how to accept. In this case, the fact that corrective computer maintenance will always exist.
Failures are an inevitable part of life among cables and silicon, so we must be prepared to respond.
However, it is true that, the more investment there is in preventive and predictive maintenance, the less corrective maintenance will be needed, saving costs, time and sanity.
So let us put unhealthy perfectionism aside, because the goal of the strategy is not to eliminate reactive maintenance in one fell swoop, but to reduce it progressively and sustainably by adjusting the rest of the maintenance types.

How to decide which type of maintenance to invest in
Prioritization must be guided by criticality, since not all systems have the same impact if they fail.
That is why the first step before allocating or refocusing resources on computer maintenance is to identify those systems that are critical to business operations.
Once the critical elements have been identified, if we want maintenance not to devour margins and budgets, we need automation and standardization as a fundamental part of the strategy.
One of the factors for this is that “boredom” I mentioned. Because boring is good in IT, but the tedium inherent in maintenance tasks makes people become overconfident and careless.
Standardization and automation are the antidote to that “boring” quality, because they always act, without sighing or rationalizing why “nothing will happen if I do not make a backup today”.
Taking the above into account, from the identification of the critical elements we will go, in order of importance, designing the maintenance plan for each part of the infrastructure, determining its needs and the actions required to meet them. But once again the reality is that, if that plan depends on someone remembering to execute it manually, then it is a plan with an expiration date.
I will not be the one to say anything good about Skynet and the rise of the machines, but IT operational efficiency largely involves reducing dependence on the human factor in repetitive and predictable tasks, as many maintenance tasks usually are.

The relationship between computer maintenance and monitoring

Since maintenance goes beyond reactive work, all the way to predictive and evolutionary approaches, monitoring is the backbone of modern and complete maintenance.
Without real-time visibility into the status of systems, predictive maintenance is impossible and preventive maintenance becomes arbitrary.
Fortunately, IT monitoring makes it possible to detect trends, correlate seemingly unrelated events and anticipate failures before they occur.
Monitoring transforms data into actionable knowledge and that knowledge is the raw material of any maintenance strategy that aims to be proactive.

Tools to manage IT maintenance

Good luck building a cathedral without tools and good luck also managing computer maintenance in professional environments while trying to do without them too.
For modern maintenance that covers all the facets we have seen, we need to:

  • Centralize information.
  • Automate repetitive tasks.
  • Ensure the team operates with context about what is happening and traceability of what it does.

For this, in incident and ticket management (the first line of inevitable corrective maintenance), solutions such as GLPI or Pandora FMS’s ITSM module (whose installation can be checked here) allow incidents to be recorded, classified, escalated and resolved in a structured way.
In addition, they are not only useful for the corrective side: the history they generate is pure gold for subsequent analysis and continuous improvement, especially to reduce MTTR, prevent the same issues from recurring and thus improve preventive, predictive and evolutionary maintenance.
For their part, RMM tools are our best card until teleportation is invented, because they provide remote visibility and the ability to act on systems without physical travel.
In distributed environments, or when we are talking about teams covering multiple sites with limited resources, this is crucial.
However, we can do more in maintenance, because there is only one good incident: the one that does not happen. That is where monitoring tools such as Pandora FMS come in, essential if we really want to prevent, predict and evolve.

How Pandora FMS helps with computer maintenance

Pandora FMS acts as the central base, repository and brain of proactive maintenance. It is not (yet) that Enterprise computer, but it is the closest thing to it.
Thanks to its continuous monitoring of systems, networks, applications and services, it provides the visibility that preventive and predictive maintenance need to detect the anomalies that precede failures.
For their part, configurable alerts allow the team to receive notifications exactly when and where they need them, with smart thresholds that reduce noise and ensure that each alert is actionable.
Pandora’s ITSM module, which I already mentioned a little above, facilitates corrective maintenance by managing the full incident lifecycle, from automatic detection to resolution and documentation.
And for evolutionary maintenance, the historical visibility provided by the platform is key, with capacity trends, performance evolution, identification of bottlenecks that point to the need for an update before they become an urgent problem…
In short, Pandora FMS is like Odin’s ravens, providing information and clarity about everything that happens and needs to be done, together with the ability to do it faster, more effectively and more easily.
In the end, the main issue is that an IT department that operates mostly in reactive mode (putting out fires and chasing problems with a butterfly net…) is paying too high a price in time, resources and team burnout.
That is why the evolution toward a preventive and predictive approach must be the compass, because otherwise it will be the epitaph.
Our IT world has become too complicated and we must strategically combine the four types of computer maintenance, supported by tools that provide real visibility and automation.
Chaos at two in the morning has a solution and it almost always consists of having done the previous work at less untimely hours. And yes, I know we do not have time for those things, but will we have much more afterwards to fix the damage?

Shares