Latest blog and updates | Apex Configuration Group

Run Discovery Before Your CMDB Is Ready And You Get Duplicates

Written by Paulina Nadera | Oct 4, 2026, 8:00:00 AM

Run ServiceNow Discovery before your configuration management database (CMDB) is ready, and duplicates are the predictable result. Discovery sends payloads. The Identification and Reconciliation Engine (IRE) decides whether each payload matches an existing configuration item (CI) or creates a new one, using identification and reconciliation rules someone was supposed to design and test. If those rules are wrong, every scan adds records instead of updating them. The duplicates then either pile up in a de-duplication queue nobody works, or never reach the queue at all.

Key Takeaways

  • ServiceNow Discovery sends the data. Duplicates come from identification and reconciliation rules that can't match a payload to the CI that already exists.
  • Three causes recur: identification rules that use the wrong attributes, other sources creating CIs without the identifier Discovery matches on, and reconciliation rules that let sources overwrite each other.
  • Before the first full scan, each class needs identification rules tested against real payloads, a named authoritative source for every shared attribute, and a named team working the de-duplication queue.
  • Pilot one range, compare the CI count with the real device count, and expand only when the previous range's de-duplication tasks are closed.
  • Track the age of open de-duplication tasks as well as their number. A queue nobody touches is how backlogs of tens of thousands build up.

This article covers what "ready" means before a Discovery rollout, what happened on an estate where the rules weren't, and the sequence we follow so the first full scan adds clarity rather than noise. For how Discovery fits into the wider model, see our guide to how an enterprise CMDB is structured.

How duplicates get made

A duplicate is created when the IRE can't match an incoming payload to the CI that already exists. There are three common reasons.

  • The identification rule uses the wrong attributes. It matches on something the payload doesn't carry, or on something that changes, such as a host name after a rebuild.
  • Another source got there first with different identifiers. An import or integration created the CI without the attribute Discovery matches on, so Discovery can't find it.
  • Reconciliation rules let sources overwrite each other. A source changes the identifying attribute on an existing CI, and the next Discovery run no longer recognises it.

In each case the IRE is doing exactly what it was told. That's why the duplicates keep coming until the rules change. Our article on how identification rules prevent duplicates goes into the matching logic in detail.

The queue nobody was working

We saw where this leads at a consulting organisation with an estate of around seven million CIs.

The organisation believed the IRE was preventing duplicates. It had Discovery running, identification and reconciliation rules in place, and a de-duplication process on paper.

In practice, the existing scans were creating duplicate CIs, because the identification and reconciliation rules had been misconfigured. When we were brought in, we found more than 50,000 de-duplication tasks that had never been addressed.

Clearing them took two Apex analysts eight weeks. In parallel, an Apex architect fixed the IRE configuration, so the tasks stopped being created faster than they could be closed. Resolving the backlog without fixing the rules would have been work with no end.

What "ready" means before the first full scan

Readiness is a set of decisions, each with an owner. We use this list before any rollout, or any major expansion of one.

DecisionReady when
Class modelEvery class Discovery will populate is agreed, including where it sits in the Common Service Data Model (CSDM)
Identification rulesEach class has a rule naming the identifiers to match on, tested against real payloads
Data sourcesEvery source writing to those classes is listed, with the identifier it supplies
Reconciliation rulesFor attributes written by more than one source, the authoritative source is named
De-duplication ownershipA named team works the queue, with a target time to close each task
Rollout planRanges are phased, with a check between each phase

If a row can't be completed, that part of the rollout waits. It's faster to delay a range by a fortnight than to spend two months unpicking it.

Rolling out so problems surface early

These steps keep a rollout from producing a backlog.

  1. Run a duplicate report before the first scan. Group existing CIs in the target classes by serial number and correlation identifier. Duplicates already in the CMDB will attract Discovery payloads unpredictably, so resolve them first.
  2. Pilot one range and compare counts. Scan a small range you know well. Compare the CI count before and after with the number of devices really in that range. If the count rises by more than the new devices, the rules aren't matching.
  3. Phase by range, not by date. Expand to the next range only when the previous one passed its count check and its de-duplication tasks are closed. Our page on ServiceNow Discovery setup covers scheduling and phasing.
  4. Watch the de-duplication queue daily during rollout. A rising queue during a phase is the earliest warning you'll get. Stop the phase, find the class, fix the rule.
  5. Check the sources the queue can't see. Compare CI counts per source against each source's own console. Duplicates created by an integration with no shared identifier won't always appear as de-duplication tasks at all.

This is also why Discovery can't be treated as the CMDB itself. We cover that argument in why scanning isn't enough.

What to measure during and after rollout

Track these as key performance indicators (KPIs), by class:

  • Open de-duplication tasks, and their age
  • CI count per class against the expected device count
  • Payloads the IRE rejected or couldn't match, by rule
  • CIs per source against each source's own count

Age matters more than volume. A queue of a few hundred tasks worked every week is healthy. A queue of any size that nobody touches is how 50,000 accumulate.

How Apex helps

Our data quality assessment measures your CMDB against weighted critical success factors, KPIs and metrics, and identifies where de-duplication, normalisation and reconciliation are failing and why. It's the right starting point if Discovery is already running and the duplicates are already there. It's scoped at an initial consultation.

If you're planning or expanding a rollout, our discovery and integration coverage review establishes what's missing from your estate view and which sources will collide, before the scans start.

Book a CMDB diagnostic call or arrange a meeting with one of our consultants.

Frequently asked questions

Does ServiceNow Discovery create duplicate CIs?

Discovery sends the data. Duplicates come from identification and reconciliation rules that can't match it to existing CIs, or from other sources that created CIs without the identifiers Discovery uses.

Can we clear duplicates without changing the rules?

You can clear the backlog, but new duplicates will keep arriving. Fix the identification and reconciliation rules first, or at the same time, so the queue stops refilling.

Why don't some duplicates appear in the de-duplication queue?

The engine only raises tasks for records it recognises as possible matches. If a source supplies no identifier it can correlate, its records look like new CIs, not duplicates.

How should a Discovery rollout be phased?

By network range. Pilot a range you know, check the CI count against reality, close its de-duplication tasks, then move to the next.