Latest blog and updates | Apex Configuration Group

Recurring Incidents: Spotting Patterns Across Shared CIs

Written by Krzysztof Chrzanowski | Oct 1, 2026, 8:00:02 AM

Recurring incidents often don't look recurring. They arrive on different services, through different support groups, with different symptoms, and each one is resolved on its own. What they share sits a layer below: a load balancer, a certificate, a storage array or a database cluster that several services depend on. The configuration management database (CMDB) is the only place that connection is recorded. If incidents are tied to the right configuration items (CIs), and those CIs are related to the infrastructure beneath them, the pattern shows up in a report. If not, it shows up in the next outage.

This article covers why shared components hide patterns, what happened when one was missed, and the reports and triggers we set up so problem management sees the pattern first. If you want the model behind it, our guide to the configuration management database and how it works explains CIs and relationships.

Why incidents on different services look unrelated

Incident reporting usually groups by service, support group or category. Those are the dimensions people choose from when they log an incident.

A shared component cuts across all three. The payroll application and the customer portal belong to different services, different teams and different categories. If both sit behind the same load balancer, their incidents share a cause but not a single field on the incident record.

So each team sees a few incidents a month on its own service, none of which justifies a problem record. The pattern only exists when you add the incidents together at the component they have in common. That needs relationships.

What a missed shared component looks like

At a broadcaster with an estate of around 15,000 CIs, services had been mapped by tagging. Each application manager tagged infrastructure with the application they considered most important. A shared component therefore appeared under one application only.

When that component was taken offline for a change, several business applications and services failed at once. A later audit found that the applications that failed had no service maps at all. Problem management was able to link the incidents back to the change, but only after the fact, because nothing in the CMDB had tied those services to the same component beforehand.

The same gap that hid the risk from change management hid the pattern from problem management. We cover the change side in why change impact analysis is only as good as your CMDB. The fix for both was the same. We replaced tagging with top-down, pattern-based service mapping, so each shared component is related to every service that uses it.

Building the report that finds shared causes

You need three things in place before a pattern report is worth running.

Incidents resolved against the failed CI. Not the service, and not a catch-all record. Our article on why problem management can't find root cause without a CMDB explains how to make this stick at the service desk.

Relationships down to the shared layer. Application servers linked to the load balancers, storage, database clusters and network devices beneath them.

A roll-up that follows the relationships. The report takes each incident's CI, walks down to the shared components beneath it, and counts incidents per shared component across all services.

With those in place, the output is short and useful:

Shared componentServices affected in periodIncidents in periodSupport groups involved
A load balancer pairSeveralMany, each smallSeveral, none owning the component
A certificateEvery service using itA cluster on one dayWhoever's service broke first
A storage arraySeveralIntermittent slownessApplication teams, not storage

The row to look for is the one where many services and many support groups meet at a component none of them owns. That's a problem record waiting to be opened.

Turning the pattern into a trigger

A report that someone has to remember to run finds patterns late. These controls make the pattern raise itself.

  1. Set a cross-service trigger. Agree a rule, for example a set number of incidents across two or more services on the same shared component within a month, that raises a candidate problem automatically. Tune the threshold to your volumes.
  2. Review by component class every week. Put certificates, load balancers, storage and database clusters on the problem review agenda as classes, not services. These are where shared causes live.
  3. Line incidents up against changes. For each cluster, check for changes recorded against the same shared component in the days before. Correlation isn't proof, but it tells the investigator where to look first.
  4. Name an owner for every shared component. A component that serves five services and is owned by none of them will never have a problem raised against it. Record the owner and support group, and hold them to it in the RACI.
  5. Fix the map after every miss. Each time a pattern is found late, record the missing relationship and add it, so the next roll-up includes it.

For the reporting side of this, including dashboards built on CMDB data, see our page on CMDB reporting and dashboards.

Measuring whether patterns are being caught

Track these as key performance indicators (KPIs):

  • Problems raised from cross-service triggers, against problems raised after a major incident
  • Shared components with no owner or no support group
  • Incidents resolved against a specific CI rather than a service
  • Critical services with relationships mapped down to the shared layer

The first measure tells you whether problem management is working ahead of the outage or behind it. It should shift towards triggers as the relationships improve.

How Apex helps

Our data quality assessment measures your CMDB against weighted critical success factors, KPIs and metrics, including relationship coverage down to shared infrastructure and the accuracy of the CIs incidents are resolved against. It shows where patterns are hidden and why. It's scoped at an initial consultation against your estate.

For a quick, scored starting point, our two-week CMDB health baseline tells you which CIs are trustworthy and which are guesses, in a report you can take to your change advisory board.

Book a CMDB diagnostic call or arrange a meeting with one of our consultants.

Frequently asked questions

How do you find recurring incidents across different services?

Roll incidents up from the CI each one was resolved against to the shared components beneath it, then count incidents per shared component across all services. That needs relationships down to the shared layer.

Why don't normal incident reports show these patterns?

They group by service, support group or category. A shared component cuts across all three, so its incidents are split between teams that each see only a few.

What threshold should raise a candidate problem?

There's no universal figure. Start with a small number of incidents across two or more services on one shared component within a month, then tune it against your incident volumes.

Which components most often hide shared causes?

Load balancers, certificates, storage arrays, database clusters and core network devices. They serve many services and are often owned by none of them.