Skip to content
Zero-Downtime Modernization

Modernise A System That Cannot Stop

Some systems have no maintenance window. We plan modernisation so that the old system stays in service until the new one has carried real traffic and matched its results.

Zero downtime is a design goal and a method. We describe how we work towards it, not a guarantee.

Why Modernisation Efforts Fail

  • The single big cutover

    Everything moves on one night. If anything is wrong, the only options are to push on or to reverse the whole change under pressure.

  • Undocumented behaviour

    The old system does things nobody wrote down. The new one leaves them out, and the gap appears only in production.

  • Data treated as an afterthought

    The code is ready but the data move is planned late. Long locks, lost writes and mismatched formats follow.

  • No tested rollback

    A way back exists on paper but was never rehearsed. When it is needed, it does not work as expected.

Four Principles

  • Parallel Execution

    Old and new systems run at the same time on the same inputs. The old system remains the source of truth until results match.

  • Traffic Validation

    The new system is judged on real traffic. We start with copies of requests, then a small share of live requests, then more.

  • Reversibility

    Each step is designed so that it can be undone. Rollback is rehearsed before the step is taken.

  • Blast-Radius Control

    Each change touches one bounded part and a limited group of users. A fault stays small and is easier to trace.

The Method In Five Steps

  1. 1
    Step 1

    Map

    We record traffic, dependencies, scheduled jobs and integrations. The result is an inventory of everything that must keep working.

  2. 2
    Step 2

    Route

    We place a routing layer in front of the old system. At first it passes everything through unchanged.

  3. 3
    Step 3

    Shadow

    The new component receives copies of real requests. Its responses are compared with the old ones and are not shown to users.

  4. 4
    Step 4

    Shift

    A small share of live traffic moves to the new component. The share grows only while error rates and results stay within agreed limits.

  5. 5
    Step 5

    Retire

    When the new component carries all traffic and has been stable, the old one is switched off and then removed.

The limits that allow or stop each traffic increase are agreed with you before the first shift.

Where Downtime Hides

Outages during modernisation rarely come from the new code itself. They come from the parts around it.

Where downtime hidesFailure patternPrevention
Data migrationA schema change locks a busy table, or writes made during the copy are lostAdditive schema changes, backfill in small batches, and row counts and checksums compared before the switch
DNS and certificatesCached DNS records keep sending users to the old address, or the new endpoint has no valid certificateShorten record lifetimes ahead of the move, test certificates on the new endpoint first, and keep the old endpoint serving until traffic drains
Background jobsA scheduled job runs in both systems, or in neitherOne named owner per job at any time, jobs written to be safe to run twice, and an explicit handover for each
Third-party integrationsA partner allows only the old IP addresses, or webhooks still point at the old URLAn inventory of every integration, partner changes arranged ahead of the move, and old callback URLs kept forwarding
Session stateSessions held in memory on old servers are lost, and users are signed out mid-taskSessions moved to a shared store first, in a format both versions can read

How AI Supports Safety

AI tools assist engineers. They do not approve cutovers or change production on their own.

  • Finding hidden dependencies

    AI tools help engineers search old code for calls, jobs and integrations that are easy to miss. Each finding is confirmed by a person.

  • Drafting comparison tests

    AI tools propose test cases for checking old results against new ones. Engineers review and correct them before use.

  • Sorting differences

    During parallel runs, AI tools help group similar mismatches so engineers can review them faster. People decide what each one means.

Your code and data are used only with AI tools you have approved.

Where Downtime Is Not An Option

  • Payments

    A failed or duplicated transaction affects real money. Parallel runs compare results before the new path handles a payment.

  • Healthcare

    Clinical and scheduling systems are used around the clock. Changes must not interrupt access to patient records.

  • Logistics

    Warehouses and fleets keep moving. A stopped system means goods that cannot be scanned, routed or dispatched.

  • SaaS with SLAs

    Availability is written into customer contracts. Planned outages count against the same commitment.

See how we work across sectors on the industries page.

Why CTOs Trust This Approach

  • Decisions rest on evidence

    Traffic moves forward because measured results match, not because a date has arrived.

  • Rollback is rehearsed

    The way back is tested before each step, so using it is a routine action.

  • Progress is visible

    Work is tracked in your tools and committed to your repository. You can see the state of each component at any time.

  • Terms stay flexible

    Engagements run month to month with no exit fee. You own all code and IP.

Why Most Firms Struggle

Why most firms struggle

  • Cutover by calendarThe go-live date is fixed early and the plan is bent to meet it.
  • Testing on synthetic dataThe new system passes tests that do not look like production traffic.
  • Rollback as a documentThe way back is written but never run.
  • Data moved lastMigration is planned after the code, when options are limited.

What we do differently

  • Cutover by evidenceTraffic increases only when agreed measures are met.
  • Testing on real trafficShadow requests show how the new system behaves under actual use.
  • Rollback as a drillEach rollback is rehearsed before the step it protects.
  • Data planned firstThe data move shapes the plan from the start.

The Zero-Downtime Assessment

A review of one system, focused on what could interrupt service during a migration.

What we analyse

  • Traffic patterns and peak periods
  • Data stores and how they are written to
  • Scheduled jobs and integrations
  • Current deployment and rollback process

What you receive

  • A list of downtime risks in order of importance
  • A proposed migration sequence
  • The measures we would use to allow each traffic shift

What happens next

  • An NDA is signed before any code access
  • A call with a senior engineer to agree scope
  • We present the findings and you decide whether to proceed

Request An Assessment

Tell us what the system does and why it cannot go offline.

No code is shared before an NDA is signed.

Frequently Asked Questions

Have More Questions?

No responsible engineer can guarantee that. What we offer is a method designed to avoid downtime: parallel runs, staged traffic shifts and a rehearsed rollback at every step.

Yes, for the period of overlap. You pay for extra infrastructure while both systems run. We plan the overlap per component so that it is as short as the evidence allows.

The old system stays the source of truth during the parallel run. Data is copied to the new store in batches and kept in step, and counts and checksums are compared before any switch.

Traffic is routed back to the old component, which is still running. The cause is investigated before another shift is attempted.

No. Adding one is part of the method. Many systems already have a load balancer or gateway that can be used for this.

Plan A Migration Without An Outage Window

Request the assessment. You will get a written list of downtime risks and a proposed sequence.