Run ServiceNow Discovery before your configuration management database (CMDB) is ready, and duplicates are the predictable result. Discovery sends payloads. The Identification and Reconciliation Engine (IRE) decides whether each payload matches an existing configuration item (CI) or creates a new one, using identification and reconciliation rules someone was supposed to design and test. If those rules are wrong, every scan adds records instead of updating them. The duplicates then either pile up in a de-duplication queue nobody works, or never reach the queue at all.
This article covers what "ready" means before a Discovery rollout, what happened on an estate where the rules weren't, and the sequence we follow so the first full scan adds clarity rather than noise. For how Discovery fits into the wider model, see our guide to how an enterprise CMDB is structured.
A duplicate is created when the IRE can't match an incoming payload to the CI that already exists. There are three common reasons.
In each case the IRE is doing exactly what it was told. That's why the duplicates keep coming until the rules change. Our article on how identification rules prevent duplicates goes into the matching logic in detail.
We saw where this leads at a consulting organisation with an estate of around seven million CIs.
The organisation believed the IRE was preventing duplicates. It had Discovery running, identification and reconciliation rules in place, and a de-duplication process on paper.
In practice, the existing scans were creating duplicate CIs, because the identification and reconciliation rules had been misconfigured. When we were brought in, we found more than 50,000 de-duplication tasks that had never been addressed.
Clearing them took two Apex analysts eight weeks. In parallel, an Apex architect fixed the IRE configuration, so the tasks stopped being created faster than they could be closed. Resolving the backlog without fixing the rules would have been work with no end.
Readiness is a set of decisions, each with an owner. We use this list before any rollout, or any major expansion of one.
| Decision | Ready when |
|---|---|
| Class model | Every class Discovery will populate is agreed, including where it sits in the Common Service Data Model (CSDM) |
| Identification rules | Each class has a rule naming the identifiers to match on, tested against real payloads |
| Data sources | Every source writing to those classes is listed, with the identifier it supplies |
| Reconciliation rules | For attributes written by more than one source, the authoritative source is named |
| De-duplication ownership | A named team works the queue, with a target time to close each task |
| Rollout plan | Ranges are phased, with a check between each phase |
If a row can't be completed, that part of the rollout waits. It's faster to delay a range by a fortnight than to spend two months unpicking it.
These steps keep a rollout from producing a backlog.
This is also why Discovery can't be treated as the CMDB itself. We cover that argument in why scanning isn't enough.
Track these as key performance indicators (KPIs), by class:
Age matters more than volume. A queue of a few hundred tasks worked every week is healthy. A queue of any size that nobody touches is how 50,000 accumulate.
Our data quality assessment measures your CMDB against weighted critical success factors, KPIs and metrics, and identifies where de-duplication, normalisation and reconciliation are failing and why. It's the right starting point if Discovery is already running and the duplicates are already there. It's scoped at an initial consultation.
If you're planning or expanding a rollout, our discovery and integration coverage review establishes what's missing from your estate view and which sources will collide, before the scans start.
Book a CMDB diagnostic call or arrange a meeting with one of our consultants.
Discovery sends the data. Duplicates come from identification and reconciliation rules that can't match it to existing CIs, or from other sources that created CIs without the identifiers Discovery uses.
You can clear the backlog, but new duplicates will keep arriving. Fix the identification and reconciliation rules first, or at the same time, so the queue stops refilling.
The engine only raises tasks for records it recognises as possible matches. If a source supplies no identifier it can correlate, its records look like new CIs, not duplicates.
By network range. Pilot a range you know, check the CI count against reality, close its de-duplication tasks, then move to the next.