Recurring incidents often don't look recurring. They arrive on different services, through different support groups, with different symptoms, and each one is resolved on its own. What they share sits a layer below: a load balancer, a certificate, a storage array or a database cluster that several services depend on. The configuration management database (CMDB) is the only place that connection is recorded. If incidents are tied to the right configuration items (CIs), and those CIs are related to the infrastructure beneath them, the pattern shows up in a report. If not, it shows up in the next outage.
This article covers why shared components hide patterns, what happened when one was missed, and the reports and triggers we set up so problem management sees the pattern first. If you want the model behind it, our guide to the configuration management database and how it works explains CIs and relationships.
Incident reporting usually groups by service, support group or category. Those are the dimensions people choose from when they log an incident.
A shared component cuts across all three. The payroll application and the customer portal belong to different services, different teams and different categories. If both sit behind the same load balancer, their incidents share a cause but not a single field on the incident record.
So each team sees a few incidents a month on its own service, none of which justifies a problem record. The pattern only exists when you add the incidents together at the component they have in common. That needs relationships.
At a broadcaster with an estate of around 15,000 CIs, services had been mapped by tagging. Each application manager tagged infrastructure with the application they considered most important. A shared component therefore appeared under one application only.
When that component was taken offline for a change, several business applications and services failed at once. A later audit found that the applications that failed had no service maps at all. Problem management was able to link the incidents back to the change, but only after the fact, because nothing in the CMDB had tied those services to the same component beforehand.
The same gap that hid the risk from change management hid the pattern from problem management. We cover the change side in why change impact analysis is only as good as your CMDB. The fix for both was the same. We replaced tagging with top-down, pattern-based service mapping, so each shared component is related to every service that uses it.
You need three things in place before a pattern report is worth running.
Incidents resolved against the failed CI. Not the service, and not a catch-all record. Our article on why problem management can't find root cause without a CMDB explains how to make this stick at the service desk.
Relationships down to the shared layer. Application servers linked to the load balancers, storage, database clusters and network devices beneath them.
A roll-up that follows the relationships. The report takes each incident's CI, walks down to the shared components beneath it, and counts incidents per shared component across all services.
With those in place, the output is short and useful:
| Shared component | Services affected in period | Incidents in period | Support groups involved |
|---|---|---|---|
| A load balancer pair | Several | Many, each small | Several, none owning the component |
| A certificate | Every service using it | A cluster on one day | Whoever's service broke first |
| A storage array | Several | Intermittent slowness | Application teams, not storage |
The row to look for is the one where many services and many support groups meet at a component none of them owns. That's a problem record waiting to be opened.
A report that someone has to remember to run finds patterns late. These controls make the pattern raise itself.
For the reporting side of this, including dashboards built on CMDB data, see our page on CMDB reporting and dashboards.
Track these as key performance indicators (KPIs):
The first measure tells you whether problem management is working ahead of the outage or behind it. It should shift towards triggers as the relationships improve.
Our data quality assessment measures your CMDB against weighted critical success factors, KPIs and metrics, including relationship coverage down to shared infrastructure and the accuracy of the CIs incidents are resolved against. It shows where patterns are hidden and why. It's scoped at an initial consultation against your estate.
For a quick, scored starting point, our two-week CMDB health baseline tells you which CIs are trustworthy and which are guesses, in a report you can take to your change advisory board.
Book a CMDB diagnostic call or arrange a meeting with one of our consultants.
Roll incidents up from the CI each one was resolved against to the shared components beneath it, then count incidents per shared component across all services. That needs relationships down to the shared layer.
They group by service, support group or category. A shared component cuts across all three, so its incidents are split between teams that each see only a few.
There's no universal figure. Start with a small number of incidents across two or more services on one shared component within a month, then tune it against your incident volumes.
Load balancers, certificates, storage arrays, database clusters and core network devices. They serve many services and are often owned by none of them.