Live Fabric core cabinet rebuild

At a glance

Client
A global cloud services provider
Sector
Cloud services
Facility
London colocation site, 200+ cabinet data hall
Service
Cabinet rebuild and recabling on live infrastructure
Scale
4 core router cabinets, each with multiple routers, ultra high density patch panels, 100+ uplinks and around 40 cross connections
Duration
4 days preparation and audit, then 4 weeks on site — one cabinet a week with 2 days of suspension
Outcome
All four cabinets rebuilt and recabled, every connection labelled to the documentation, minimal troubleshooting at handback

The situation

The client is a global cloud services provider. These four cabinets carry the core connections for a 200+ cabinet data hall at a London colocation site, running a Fabric topology. Each one holds multiple routers with a dense mesh of cross connections between them and more than a hundred uplinks coming in from top of rack switches.

The cabinets had simply been outgrown. They were built with low-density patch panels, and as the Fabric grew and more routers were needed, those panels were taking up rack space the routers required. There was nowhere left to put more routers, and nowhere left to add more patch panels to serve them either. The only way forward was to rebuild the cabinets around ultra high density LC fibre patch panels on MPO trunks, which meant moving every router into a new position.

All of it was live throughout. The Fabric is built to tolerate the loss of a cabinet, so taking one out does not stop the service. What it does is reduce the available bandwidth, and that affects overall performance for as long as the cabinet is out. The aim was therefore to keep each cabinet out of service for as little time as possible.

The constraint

Time under suspension was the whole problem.

A cabinet like this cannot be rebuilt gradually. The routers all have to come out and go back in new positions, which means the cross connections between them and every uplink from the switches come out too. Around 140 fibre connections per cabinet, each of which has to go back to a specific port.

Every hour a cabinet is suspended is an hour the Fabric runs with less bandwidth than it should have. That is not an outage, but it is a degraded service, and every connection that goes back wrong extends the window while somebody traces it. The work is also dense enough that only a limited number of engineers can physically get to the rack at once, so the obvious answer of adding more people does not work.

DACPROS planned and ran the rebuild

The client agreed the approach of taking one cabinet at a time and verified the Fabric at the end of each one. Everything else was ours.

We audited the cabinets, wrote the method, decided what could be done outside the suspension windows and what could not, scheduled the suspensions, resourced the shifts and carried out the rebuild. The plan came from us.

What we did

Audited every cabinet before anything was touched. Four days were set aside for preparation and audit before the first suspension. The first job was to find out what was actually there rather than what the records said was there. Missing and unclear labels were identified and corrected. A rebuild depends entirely on knowing which cable goes where, and an unlabelled fibre found during a suspension window is an expensive thing to trace.

Moved every possible task out of the suspension window. We went through the work item by item and separated what genuinely required the cabinet to be down from what did not. Anything in the second group was collected into a list and carried out in advance by our field engineers during normal operation. By the time each suspension started, the only work left was work that could not have been done any earlier. Cabling was the clearest example. Because the routers were going back in different positions, a number of cross connections needed new runs. Those were pulled in advance and left hanging in the cabinet, so on the suspension day they only had to be plugged in. The old runs stayed live until the window, then came out on the housekeeping day afterwards.

Wrote the rebuild as a documented step sequence. Every stage was written down in order before the first suspension. That meant every engineer on shift was working to the same document rather than to their own reading of the job, which is what allowed the team to work quickly without checking with each other at every step.

Scheduled the suspensions in advance, with slack built in. The project ran to a four-week schedule, one cabinet a week. Each week carried two days of suspension, a day of housekeeping after it, and two days left as buffer. Windows were booked ahead against the plan rather than arranged as the project went along, so the client knew when each part of the Fabric would be affected and had the buffer if anything needed it.

Ran a two-shift day. Only a limited number of engineers can work in one cabinet safely and usefully. Rather than accept an eight-hour day, we proposed early and late shift teams, giving a sixteen-hour working day at the same density of people in the rack. The suspension window was being consumed by elapsed time, not by headcount, so doubling the hours in the day halved the number of days the cabinet was down.

Rebuilt one cabinet at a time. Every router came out and went back in a new position. New ultra high density fibre patch panels and cable management accessories went in. More than a hundred uplinks and around forty inter-router cross connections were rebuilt and labelled exactly as the documentation specified.

Repeated the same process three more times. The method did not change for the remaining cabinets. That is the point of writing it down.

Every connection accounted for

Every connection was labelled to the documented scheme as it went in, not afterwards. Because the audit and the labelling had been done properly beforehand, the troubleshooting at the end of each suspension was minimal. There was very little hunting for a connection that had not come back up, which is normally where a rebuild of this kind loses its remaining time.

The result

Two days of suspension per cabinet.

Each cabinet came back with its routers repositioned, new ultra high density fibre patch panels in place, and around 140 connections rebuilt and labelled to the documentation. All four were delivered inside the four-week schedule, one a week.

The buffer days were never used. They were there in case something unexpected came up, and nothing did, because the audit and the preparation had already found the things that would otherwise have surfaced mid-window.

The client’s own team had one job at the end of each window, which was to verify the Fabric. They did not have to plan the work, sequence it, resource it or troubleshoot it.

Time spent with reduced bandwidth was held to the lowest level the physical work allowed, and it was held there four times in a row rather than once.

A long-running relationship

This client has worked with us since DACPROS was founded, and is today one of our largest customers.

It started small. Over the years the work grew more complex as we proved ourselves on each piece of it, and as our engineers brought plans to the client rather than waiting to be told. This project is where that got to.

Services used

Cabinet rebuild and recabling on live infrastructure, with audit, method planning, suspension scheduling and shift resourcing run by DACPROS.

If you need work carried out on infrastructure that cannot be taken down for long, we will plan the sequence and the windows before anything is touched.

Work on live infrastructure?

Request a callback