Resilience

Critical infrastructure does not fail alone

Every organisation in a dependency chain can be individually compliant while the chain as a whole is assured by nobody. That gap is where national scale outages actually come from.

There is a particular kind of incident report that has become very common, and it always reads the same way. A service that millions of people depend on stops working. The organisation that runs it did nothing wrong. Its own systems were healthy, its own controls were in place, and its own last assessment came back clean.

The failure came from four steps upstream, in a company most of the affected users have never heard of, which in turn depended on something further upstream again.

One failure, four consequences
One failure, four consequences

Everybody in that chain was individually compliant. The chain itself had no owner.

Why assurance stops at the organisational boundary

Compliance frameworks are written for organisations. That is entirely reasonable, since an organisation is the thing that can be held accountable, audited and fined. But it produces a specific blind spot, because the scope of every assessment ends exactly where the legal entity ends.

What sits beyond that boundary gets handled by a different mechanism, usually a supplier questionnaire or a contractual clause. Both are weak substitutes for assurance. A questionnaire tells you what a supplier was willing to assert about itself at a point in time, filled in by somebody in their sales organisation who was asked to complete it quickly. A contract tells you who is liable afterwards, which is a legal comfort rather than an operational one.

Neither one tells you what happens on Friday if that supplier is unavailable, and neither one knows what your supplier's supplier does.

The concentration nobody can see from inside

The more interesting problem is not any single dependency. It is the way dependencies converge without anyone noticing.

Consider a sector regulator looking at twenty organisations. Each one submits its own posture. Each submission looks reasonable. What the regulator cannot see from those twenty documents is that fourteen of them run their most critical workload in the same hosting region, and eleven of them use the same authentication provider.

No individual organisation has made a bad decision. Each one chose a reputable provider for sound reasons. But the sector has quietly arranged itself into a shape where a single provider failure takes down most of it simultaneously, and that shape is invisible to every participant, including the regulator, because nobody is looking at the connections.

Concentration risk is not created by any one decision. It is created by many reasonable decisions that nobody compared.

What would need to change

Making this visible does not require new obligations on anyone. It requires that the information already being collected is collected in a shape that can be compared.

Three things have to be true.

  • Submissions have to share a structure. Twenty organisations describing their posture in twenty formats produces a filing cabinet rather than a picture. When the underlying model is consistent, comparison stops being a research project.
  • Dependencies have to be first class. A dependency recorded as a sentence in a supplier annex cannot be traversed. A dependency recorded as a relationship can be, which means the question of what a provider failure reaches becomes a query rather than an investigation.
  • The view has to reach past the first hop. Most third party risk programmes stop at the direct supplier. The failures that cause national outages almost always originate one or two steps further out than that.

This is not a technology problem, but technology decides whether it is tractable

Every part of the above could, in principle, be done by hand. Analysts could read twenty submissions, build a dependency map, and identify the concentration. Some sector bodies genuinely do this, and it takes them months, which means the map is already out of date when it is finished.

The reason to build this into a system is not novelty. It is that a map maintained by hand decays faster than it can be produced, and a map that reflects last year's architecture is worse than useless because it inspires confidence it has not earned.

That is the reasoning behind ControlGraph. The same structure that lets a control answer several frameworks at once is the structure that lets you ask what a single system failure would reach. Those look like two different products until you notice they are both traversals of the same graph.

The wider point

We describe what we are building as work toward infrastructure that is safer, more compliant and more resilient. The third word in that list is the one that depends on the connections.

An organisation can be made safer on its own. It can be made compliant on its own. It cannot be made resilient on its own, because resilience is a property of the chain rather than of any link in it. Any serious attempt at national resilience has to be able to see across organisational boundaries, and right now very little of our assurance machinery is built to do that.

If you sit inside a sector where everyone quietly depends on the same few providers, that is exactly the problem we would like to hear about.

Common questions

What is concentration risk in cyber security

It is what happens when many organisations independently choose the same provider for good reasons, and the sector quietly arranges itself so that one provider failure takes down most of it at once. No single participant made a bad decision, which is exactly why nobody sees it.

Why are supplier questionnaires weak assurance

They tell you what a supplier was willing to assert on a particular day, usually completed at speed by somebody under deadline pressure. They do not tell you what happens on the Tuesday their platform is unavailable, which is the question that actually matters.

How far down the supply chain should we look

Further than most programmes do. Third party risk work commonly stops at the direct supplier, and the failures that cause national scale outages usually originate one or two steps beyond that.

Who is responsible for a dependency chain

Legally, each organisation for its own part. Practically, nobody for the chain as a whole, which is the gap. Making it visible needs submissions that share a structure and dependencies recorded as relationships that can be traversed rather than sentences in an annex.

ResilienceNational InfrastructureThird Party RiskOversight

Working on something this touches?

If this raises a question about your own compliance position or your security operations, the quickest route is a direct conversation.