Ask each resilience-related function how prepared we are, and you may get several reasonable answers.
Put those answers together and try to explain how a customer order gets dispatched during a major disruption. That can take a little longer.
Business continuity, cyber, IT recovery, risk, supplier management and asset management each see part of the operation. Keeping a service running depends on the connections between them. Who takes responsibility for understanding those connections and making sure they work under disruption?
That’s the challenge behind the overlapping circles in the diagram. To deliver what sits in the center, we need a shared understanding of how the company works and capabilities that support each other under disruption. I’d give an operational resilience lead responsibility for bringing that work together, starting with one service and the people who deliver, protect and recover it. Here’s what I’d ask them to work through.

1. Follow the work from customer to delivery
Take a company that receives and fulfils customer orders. I’d sit down with the people involved and follow an order through stock allocation, picking and dispatch. Use existing process information, asset records and supplier knowledge, then ask the people doing the work to fill in what’s missing.
Suppose IT’s view covers the order application and its hosting platform. The warehouse adds that dispatch also requires stock information and a separate carrier connection to produce shipping labels. Restoring the application alone won’t get an order out of the building. Someone needs to connect those two views before recovery priorities are agreed.
Sketch those connections together. Record what each activity needs and who provides it, including people and external services. Keep the first pass focused on what could stop delivery. We can resolve unknowns as we go; we don’t need to catalogue the entire company before having a useful conversation.
2. Make something unavailable
Now suppose the hosting environment is affected by a suspected compromise. We can’t trust it, and we don’t know how long recovery will take. Follow the consequences through the operation we’ve just sketched.
The warehouse still has people and stock, but can’t access the order application. Someone suggests manual processing. Before we count on that, I’d ask where the stock information comes from. In this example, it’s in the same unavailable environment.
That gives operations, IT and cyber a specific problem to work through. Can we make the information independently accessible, with suitable protection and a way to record stock movements during the outage? Does that access rely on the same identity service that’s affected?
Bring the supplier owner into this discussion too. If the provider restores its platform, what work remains on our side before the warehouse can use it? Agree who will coordinate the handover from the provider’s recovery to our own work, and follow it all the way back to dispatch.
3. Work backwards from the minimum service
For this example, suppose the company receives 120 orders a day. Thirty require same-day dispatch. After considering customer consequences and commitments, the business agrees that other orders can wait for up to three days.
Work backwards from those 30 orders. We need to receive and identify them, allocate stock, pick them and arrange delivery. A smaller volume may still require most of the supporting services. The carrier connection doesn’t become optional because we’re sending fewer parcels.
Now the business and technical teams can agree which capabilities must remain available or come back first. If the minimum operation can’t be supported, put that shortfall in front of someone who can decide what to change or fund. The requirement and the recovery sequence need to describe the same service.
4. Check whether we can sustain it
Suppose the proposed fallback can process 40 orders a day. That could cover the 30 priority orders, provided the team has the information and access it needs. We still have to try it with the people expected to do the work.
It also leaves 80 orders accumulating each day. After three days, 240 are waiting, assuming steady arrivals, no starting backlog and no cancellations. Restoring normal capacity of 120 a day won’t clear them. IT may have completed its recovery while operations still needs extra capacity and a way to reconcile the manual records.
Ask what the workaround requires across shifts and how recovery competes for those resources. Check whether several teams are counting on the same people. Being listed in three recovery procedures doesn’t give anyone three times the capacity.
Recovery estimates need this joint scrutiny too. If a four-hour restore assumed working access and a prepared environment, establish how we’ll get those conditions in place. That’s part of the time the business will be waiting.
5. Give the work between teams an owner
We’ve now found a gap that several teams need to fix. Operations must define the stock information and manual controls it needs. IT must make the information accessible, with cyber involved in protecting it. The resilience lead needs to bring those contributions together, arrange the trial and follow up on the result.
I’d give a resilience lead responsibility for keeping that work connected, with access to someone who can resolve competing priorities and resources. The service owner stands behind the business requirement; the relevant teams own the improvements. Agree names and dates while everyone is involved.
Keep the service shortfall visible until it’s resolved or management has made an explicit decision about the remaining exposure. Otherwise, each team can finish its own task while the original problem remains between them.
This is the planning work. Capture the decisions and working arrangements in plans people can use, including who can decide what during an incident.
6. Try the service through the disruption
I’d test the connections we’re least sure about. Try receiving and dispatching priority orders without the normal application. Can the team get the stock information, use the carrier connection and maintain reliable records? Can the next shift carry on?
Bring the technical recovery results into the same discussion. If the workaround needs data that only becomes available after restoration, testing both activities separately won’t resolve the dependency.
Smaller exercises and tests can build evidence when a full service test isn’t practical. Be clear about their limits: one successful shift won’t demonstrate three days of sustainable operation. Use what we learn to improve the capabilities and update the shared picture as systems, suppliers and working practices change.
So who owns that space?
I’d make the answer explicit: a named resilience lead coordinates the work across functions, backed by a service owner and a management route for decisions the teams cannot settle. Functional owners remain responsible for delivering the improvements.
That gives the center of the diagram practical meaning. For our order service, someone follows the connections from stock information and carrier access through to dispatch, checking whether the teams can collectively deliver what the business needs.
Owning that space means building the shared operational understanding, getting people to plan and exercise together, and following gaps through to a decision or improvement. The plans capture that work. The service depends on the people and capabilities behind them.
Six Take-aways
1. Follow the operation. Connect customer-serving activities to the people, information, technology and providers it needs.
2. Trace the disruption. Find where one team’s answer depends on something another team cannot yet provide.
3. Define the minimum service. Work backwards to the capabilities and recovery sequence required to deliver it.
4. Check capacity and duration. Include shared resources, backlog and the work left after systems return.
5. Name the owner of the space between teams. Give that person a mandate to coordinate improvements and escalate unresolved trade-offs.
6. Exercise together. Test whether the capabilities support the service, then improve and maintain them.
Keeping a service running depends on the connections between teams. Name an owner for that space and give them a mandate to coordinate improvements and escalate the trade-offs the teams can’t settle. Plans will keep improving, but resilience improves when someone owns the space between the teams.
Leave A Comment
You must be logged in to post a comment.