Data basics

What Is Data Deduplication?

Data deduplication removes repeat records so each person or company appears once. Here is how it works on lead lists, and what to match on.

Get your free API key →Free to start. No credit card. 1,000 records to spend whenever you like.

Data deduplication is the process of finding records that describe the same person or company and reducing them to one. In sales data it keeps one row per human, so nobody gets the same email twice.

Key takeaways

  • Deduplication means one record per real-world entity. For lead lists, that entity is a person or a company.
  • Match on a stable key (LinkedIn URL, work email, company domain), not on a name.
  • Exact matching is cheap and catches most repeats. Fuzzy matching catches the rest and needs a human check.
  • LeadOcean exports are deduplicated by person_id, and person_group_id lets you spot the same person across records.

What it is

Data deduplication is the process of identifying records that refer to the same real-world person or company and merging them into a single record.

The term has a second meaning in storage engineering, where it removes repeated blocks of bytes from backups. This page covers the sales and marketing meaning: duplicate contacts and accounts in lists, CRMs and exports.

Duplicates enter a list in three ways. You import two sources that overlap. A person changes jobs and gets a second record. Or two reps type the same lead by hand with slightly different spelling.

How it works

Deduplication follows five steps. The order matters, because matching only works on clean input.

  1. Normalize. Lowercase emails, strip www. and paths from domains, trim spaces, and put phone numbers in one format.
  2. Pick a match key. Use the field least likely to change or vary: LinkedIn URL first, then work email, then company domain plus full name.
  3. Match. Exact matching compares keys for equality. Fuzzy matching scores near misses such as "Jon" and "Jonathan".
  4. Choose a survivor. Keep the record with the freshest data, or merge fields so the survivor holds the best value of each.
  5. Log the merge. Record which rows collapsed, so a wrong merge can be reversed.

A worked example, with placeholder data:

RowNameEmailCompany domain
1Jane Doejane.doe@acme.comacme.com
2Jane M. DoeJane.Doe@Acme.comwww.acme.com
3J. Doejdoe@acme.comacme.com

After normalizing, rows 1 and 2 share the same email and collapse into one record. Row 3 has a different address and a shortened name, so exact matching leaves it alone. A fuzzy pass would flag it as a possible match, and a person decides.

Two real people can share a name and a company, so a name match alone is a prompt to look, never a reason to delete.

Three mistakes show up again and again:

  • Matching on name. Common names collide, and one person can appear as Jon, Jonathan and J. Doe.
  • Dropping the loser. Delete a duplicate without merging and you lose the phone number or title that only that row held.
  • Deduplicating once. A list that was clean in January has new repeats by March, so treat it as a recurring job.

The rule to keep: exact matches can merge automatically. Fuzzy matches need a review queue, because a false merge destroys a real contact.

Deduplication vs data cleansing

Deduplication is one task inside data cleansing, and people often use the two terms as if they meant the same thing. They do not.

DeduplicationData cleansing
Question it answersIs this the same entity as another row?Is this value correct and well formed?
What it changesRow countField values
Typical actionMerge or drop repeat rowsFix formats, fill gaps, verify emails
Match key neededYesNo
ExampleTwo rows for Jane Doe become oneACME.COM becomes acme.com
OrderAfter normalizingBefore matching

Run cleansing first, then deduplicate. If you match before normalizing, Jane.Doe@Acme.com and jane.doe@acme.com look like two people.

When it matters

Duplicates cost money and reputation in four situations.

Cold email campaigns

A duplicate means one person gets the same sequence twice, from you, in the same week. That produces spam complaints and unsubscribes, and both damage the sending domain. Deduplicate the list before every send, not once a quarter. Check people across campaigns too, since a contact in two active sequences is a duplicate even when each list is clean.

Buying data from more than one source

Two providers usually overlap on the same senior people. If you pay per record, you pay twice for the same human. Compare the overlap on a small sample before you buy a second list. A sample of a few hundred rows tells you how much of the new list you already own.

CRM hygiene

Duplicate accounts split a customer's history across two records. Reps then work the same account in parallel, and reporting double counts pipeline. Use the company domain as the account key and merge on it. Subsidiaries that share a parent domain are the exception, so review those by hand.

Enrichment runs

Repeated keys in an input file cost repeated lookups. Remove duplicate emails, LinkedIn URLs and domains before you enrich, and the bill falls with the row count.

How LeadOcean handles it

LeadOcean deduplicates its exports, and gives you one field to check the rest yourself. These are described in the OpenAPI spec and the exports docs, checked September 2026.

  • Exports. A POST /v1/exports file is deduplicated by person_id, up to 50,000 rows per export.
  • Records vs people. person_id names a record. person_group_id names the person. Two records that share a person_group_id are the same person as far as LeadOcean's de-duplication knows.
  • Where it appears. /v2/people/reverse returns person_group_id in data. Compare it, never look it up.

LeadOcean holds more than one record for some people. The same human can come back under different person_id values depending on the contact point you searched with. So when you merge LeadOcean output with another list, key on person_group_id where you have it, then on LinkedIn URL.

The example below looks up the same placeholder person (the values are made up) by email and by phone. Each call costs 1 record. If both answers carry the same person_group_id, they are one person.

bash
curl -s -G "https://api.leadocean.io/v2/people/reverse" \
  --data-urlencode "email=jane.doe@acme.com" \
  -H "x-api-key: $LEADOCEAN_API_KEY" | jq '.data.person_group_id'

curl -s -G "https://api.leadocean.io/v2/people/reverse" \
  --data-urlencode "phone=+1 415 555 0133" \
  -H "x-api-key: $LEADOCEAN_API_KEY" | jq '.data.person_group_id'

LeadOcean does not run a rep workspace or a CRM, so it does not merge duplicates inside your CRM. Do that step in the CRM, or in a script on the exported CSV. LeadOcean prices at Free (1,000 records, one-off) and Pro at $499 a month, so a clean export costs the same as a messy one. See pricing. To compare sources before you buy, start with best B2B databases, or the UK-focused list if your market is the United Kingdom.

FAQ

What is the difference between deduplication and merging?

Deduplication finds the repeats. Merging decides what the surviving record contains. Most tools do both, but they are separate decisions, and a bad merge rule loses good data.

Is deduplication the same as email verification?

No. Verification asks whether an address can receive mail. Deduplication asks whether two rows are one person. A list can be fully verified and still hold the same person three times. LeadOcean publishes an email_status on each row, so you can check both.

Should I deduplicate by name?

No. Names repeat across thousands of people and vary in spelling for one person. Use LinkedIn URL, work email or company domain as the key, and use name only as a tie-breaker.

How often should I deduplicate a lead list?

Before every send and after every import. Lists drift as people change jobs and new sources arrive. If you buy from several sources, the data providers comparison helps you pick fewer.

Does LeadOcean remove duplicates across my own CRM?

No. LeadOcean deduplicates its own exports by person_id and exposes person_group_id so you can match. It does not read or clean your CRM.

Start with one clean record per person

Free to start. No credit card. 1,000 records to spend whenever you like.

Get your free API key →