Thursday, July 30, 2026

< + > If A Server Fails Today, How Long Until Your Patients Can Be Seen?

The following is a guest article by Michael Ohayon, GM Nexcess Managed Cloud at Nexcess

Go from “backups running” to a clinical recovery mandate

When a server fails at your practice, specialty clinic, or healthcare software company, the questions that follow are immediate and high-stakes. You need to know which systems are affected, where recoverable copies are stored, and most importantly when access will return. These questions require tested answers before the next hardware failure, cyberattack, or accidental deletion interrupts care.

In healthcare, a server outage transcends the IT department. Your front desk may lose appointment schedules, while clinicians find themselves unable to open patient records or review critical medication histories. Staff may be forced to fall back to paper workflows designed for short interruptions, not the uncertain windows associated with modern technical recovery. This is why the most vital question is,  “How long will it take to get the right systems back?”

A successful backup is not a successful recovery

Backups are often treated as evidence of preparedness. A dashboard indicates a job completed, and the organization moves on. However, a successful backup job doesn’t tell you how long a restore takes, nor does it confirm that the copy is usable for clinical applications.

This distinction matters because downtime spreads fast. A scheduling system or patient portal may not be the EHR, but its failure can still paralyze patient-facing work. A practical managed backup and storage strategy starts with recovery requirements, not storage capacity, which means defining two metrics up front:

  • Recovery Time Objective (RTO): The target duration for getting a system operational again.
  • Recovery Point Objective (RPO): The maximum amount of recent data the organization can afford to lose (e.g., losing four hours of clinical documentation vs. 24 hours).

Recovery priorities should follow the patient workflow

Not every system needs to return simultaneously. During an incident, the priority should be the apps that allow staff to identify patients, understand their clinical needs, and document treatment. Administrative reporting or development environments can wait.

That sequence has to be decided before a crisis, not negotiated in the middle of one. For example, you may decide that identity management and clinical access return first, with the patient portal following. Deciding in advance also exposes the dependencies that break recovery plans in practice: restoring an application server accomplishes nothing if the underlying database is still down.

Protection beyond the primary server

A backup stored only on the affected server is vulnerable to the same event that took the server down. Hardware failures, facility-level incidents, or ransomware can  compromise the original data and the local copy together. This reality was highlighted during the global cPanel security incident in April, forcing emergency patching and temporary service shutdowns across hosting environments to head off unauthorized root access.

Off-server and offsite protection should be the baseline, and the “3-2-1” principle remains the best guide: three copies of your data, on two forms of storage, with one copy offsite. For healthcare, that offsite copy also has to meet security and compliance obligations for encryption and access controls, not only exist. Small restores shouldn’t require big escalations

Not every recovery is a catastrophe. Often, the issue is a deleted document or an overwritten configuration file. If every minor restore requires a high-level support escalation, a small error turns into a four-hour interruption. Self-service recovery options for authorized IT staff lets routine restores  begin immediately and keeps specialized support free for genuine system-wide failures.

Nobody wants backups; what people want is the ability to restore. Getting there requires understanding the difference between standard backups and disaster recovery, which covers how infrastructure returns after a large-scale disruption.

Test the theoretical

Healthcare data grows relentlessly, and a backup strategy has to keep pace without collapsing under its own weight. Review storage against usage and retention requirements to avoid overprovisioning or hitting silent quotas that halt protection.

A recovery plan is theoretical until it is tested. Testing doesn’t require taking production systems offline; restoring selected files or validating a database copy in an isolated environment will answer the questions a dashboard can’t:

  1. Can we access the backup?
  2. Is the data complete and usable?
  3. How long does the process take?
  4. Who makes the call to fail over?

Ask the question in minutes, not hours

“Backups are running” is not an answer you can use. A resilient organization answers  in specifics: the scheduling system has a one-hour recovery target, the most recent copy is 30 minutes old, and the hosting team is executing a documented recovery runbook right now.

Server failures are inevitable, but extended uncertainty about patient care isn’t. Move from a backup mindset to a recovery mandate, and when the next outage begins, the path back to normal care is already mapped.

Michael Ohayon is with Nexcess, a managed hosting platform that provides HIPAA-aligned infrastructure and migration support for specialty practices, regional health centers, and healthcare SaaS platforms.



No comments:

Post a Comment

< + > Improving Operational Efficiency with Back Office Health IT Systems

As we try to combat healthcare burnout, improving operational efficiency is a big part of that. Like everything in healthcare, there are ple...