Why Problem Management Can't Find Root Cause Without A CMDB
Problem management can't find root cause reliably without a configuration management database (CMDB), because root cause analysis is a question about configuration items (CIs). Which component failed? What else depends on it? What changed on it recently? Which other devices share its model or version? If incidents aren't logged against the right CI, and that CI isn't connected to its services and changes, the investigation stops at the symptom. The problem record closes with "unable to determine cause", and the same fault comes back.
This article covers what a root cause investigation needs from the CMDB, where investigations usually stall, and how to design configuration data so problem management can use it. If you're newer to the model underneath, our guide to the configuration management database explains how CIs and relationships fit together.
What root cause analysis asks of the CMDB
Strip a problem investigation back and it asks four questions. Each one is answered by configuration data, or not answered at all.
| The investigator's question | What the CMDB has to hold |
|---|---|
| Which component actually failed in each incident? | The specific CI on every resolved incident, not the business service or a placeholder |
| Do these incidents share something? | Relationships from each CI to the infrastructure beneath it and the services above it |
| What changed before it broke? | Changes recorded against the same CIs, with the affected CIs filled in properly |
| Where else could this happen? | Attributes such as model, firmware, operating system version and software version, populated consistently |
A problem analyst with all four can move from a cluster of incidents to a cause in hours. A problem analyst with none of them is interviewing engineers about what they remember.
Where problem investigations usually stall
The failure points are consistent across estates, and most are set long before anyone opens a problem record.
Incidents logged against the service. The service desk records the business service the user reported, and the incident is resolved without anyone updating it to the CI that failed. Twenty incidents later, all you know is that the service is unreliable.
The wrong CI. An engineer picks the nearest match from a long list, or a generic record created years ago. The incident now points at a device that had nothing to do with it.
Relationships that stop halfway. The application is linked to its servers, but the servers aren't linked to the storage array or the network devices they share. The common factor sits one level below where the map ends.
Changes with no affected CIs. A change was made, but its affected CI list is empty or names only the service. The investigation can't see that the failing component was touched the night before.
Blank grouping attributes. Without model and version data, you can't ask which other devices run the same firmware. The problem is fixed on one device and waits on the others.
A CMDB built without the processes that use it
We saw what happens when the CMDB stops being designed for its consumers at a broadcaster with an estate running to tens of millions of CIs.
The organisation believed configuration management was a standalone process, so the person leading it didn't need a deep understanding of incident, problem or change management. On that basis, the experienced configuration management lead was made redundant and replaced by a lower-cost configuration manager.
The new owner looked after the CMDB as a data store. Over time it stopped meeting the needs of the processes that relied on it. Incidents took longer to resolve and more changes failed. Complaints arrived from across the service management processes that the CMDB was unusable.
Putting it right cost the organisation more than $100,000 in consulting fees to undo the design decisions. We carried out that work, and then took over running the configuration management process as a managed service, so the CMDB stayed aligned with the processes that consume it.
Problem management feels this first. It's the process that asks the hardest questions of configuration data, and it's the first to find out when the answers have gone.
Designing configuration data for problem management
These are the controls we put in place so problem investigations have something to work with.
- Make the failed CI mandatory at resolution. The service desk can log an incident against a service. The resolver must name the specific CI that failed before the incident can close. Report on incidents resolved against services or generic records, and take them to the next problem review.
- Retire the catch-all records. Find CIs with names like "unknown", "various" or "generic" that attract incidents. Each one hides a pattern. Replace them with the real CIs and stop them being selected.
- Map relationships to the shared layer. For critical services, make sure relationships go down to storage, network and platform components, since that's where the shared causes of problems live. Our page on dependency views and relationships covers how those links are structured.
- Require affected CIs on every change. A change with no affected CIs is invisible to the next problem investigation. Make the field mandatory, and check at post-implementation review that it names components rather than services.
- Populate the grouping attributes. Choose the attributes problem management needs to find similar devices, such as model, firmware and software version, and measure their completeness by class.
- Put the configuration manager in the problem review. Every problem that stalls on configuration data is a data defect with a name. Record it and fix it, and write that responsibility into the RACI for configuration management.
How to tell whether it's working
Track these as key performance indicators (KPIs) and review them alongside problem backlog figures:
- Resolved incidents with a specific CI rather than a service or a generic record
- Problems closed with a confirmed root cause, against problems closed as undetermined
- Changes closed with affected CIs recorded at component level
- Completeness of the grouping attributes for your critical classes
The second measure is the one to show leadership. It turns configuration data quality into a number the problem manager already cares about. For how these measures sit alongside completeness and correctness more broadly, see The 3 Cs of a CMDB.
How Apex helps
Our data quality assessment measures your CMDB against weighted critical success factors, KPIs and metrics, including the fields and relationships problem management depends on. It shows where incidents are being logged against the wrong records, where relationships stop short, and why. It's scoped at an initial consultation against your estate.
For a quicker first view, our two-week CMDB health baseline scores which CIs are trustworthy and which are guesses. If you'd rather hand the running of the process to specialists, we also offer managed CMDB services.
Book a CMDB diagnostic call or arrange a meeting with one of our consultants.
Frequently asked questions
Can problem management work without a CMDB?
It can investigate individual problems through interviews and log files. It can't reliably find patterns across incidents, because that depends on incidents, changes and components being linked in one place.
Should incidents be logged against a service or a CI?
Both, at different stages. Log against the service the user reports, then record the specific CI that failed before the incident is resolved. The second is what problem management needs.
Who should fix configuration data found wrong during a problem investigation?
The CI owner fixes the record, and the configuration manager tracks it to closure. Write both responsibilities into the RACI so the defect doesn't sit with the problem team.
Which CMDB attributes matter most for problem management?
The specific CI on each incident, relationships to shared infrastructure, affected CIs on changes, and the attributes that group similar devices, such as model and version.
- ServiceNow (25)
- CMDB (19)
- CMDB Data Quality (8)
- Cyber Security (5)
- Vulnerability Management (5)
- Service Mapping (4)
- Compliance & Audit (3)
- Enterprise Risk Management (3)
- IT Cost Optimisation (3)
- Agentic AI (2)
- CI Ownership (2)
- Change Management (2)
- Data Model Design (2)
- IT Asset Management (2)
- Incident Management (2)
- CI Reconciliation (1)
- Duplicate CIs (1)
- IRE (1)
- ITOM (1)
- Operational Efficiency (1)
- Service Graph Connectors (1)
- ServiceNow Advisory (1)
- ServiceNow Discovery (1)
