Data deduplication is the process of finding records that describe the same person or company and reducing them to one. In sales data it keeps one row per human, so nobody gets the same email twice.
Key takeaways
- Deduplication means one record per real-world entity. For lead lists, that entity is a person or a company.
- Match on a stable key (LinkedIn URL, work email, company domain), not on a name.
- Exact matching is cheap and catches most repeats. Fuzzy matching catches the rest and needs a human check.
- LeadOcean exports are deduplicated by
person_id, andperson_group_idlets you spot the same person across records.
What it is
Data deduplication is the process of identifying records that refer to the same real-world person or company and merging them into a single record.
The term has a second meaning in storage engineering, where it removes repeated blocks of bytes from backups. This page covers the sales and marketing meaning: duplicate contacts and accounts in lists, CRMs and exports.
Duplicates enter a list in three ways. You import two sources that overlap. A person changes jobs and gets a second record. Or two reps type the same lead by hand with slightly different spelling.
How it works
Deduplication follows five steps. The order matters, because matching only works on clean input.
- Normalize. Lowercase emails, strip
www.and paths from domains, trim spaces, and put phone numbers in one format. - Pick a match key. Use the field least likely to change or vary: LinkedIn URL first, then work email, then company domain plus full name.
- Match. Exact matching compares keys for equality. Fuzzy matching scores near misses such as "Jon" and "Jonathan".
- Choose a survivor. Keep the record with the freshest data, or merge fields so the survivor holds the best value of each.
- Log the merge. Record which rows collapsed, so a wrong merge can be reversed.
A worked example, with placeholder data:
| Row | Name | Company domain | |
|---|---|---|---|
| 1 | Jane Doe | jane.doe@acme.com | acme.com |
| 2 | Jane M. Doe | Jane.Doe@Acme.com | www.acme.com |
| 3 | J. Doe | jdoe@acme.com | acme.com |
After normalizing, rows 1 and 2 share the same email and collapse into one record. Row 3 has a different address and a shortened name, so exact matching leaves it alone. A fuzzy pass would flag it as a possible match, and a person decides.
Two real people can share a name and a company, so a name match alone is a prompt to look, never a reason to delete.
Three mistakes show up again and again:
- Matching on name. Common names collide, and one person can appear as Jon, Jonathan and J. Doe.
- Dropping the loser. Delete a duplicate without merging and you lose the phone number or title that only that row held.
- Deduplicating once. A list that was clean in January has new repeats by March, so treat it as a recurring job.
The rule to keep: exact matches can merge automatically. Fuzzy matches need a review queue, because a false merge destroys a real contact.
Deduplication vs data cleansing
Deduplication is one task inside data cleansing, and people often use the two terms as if they meant the same thing. They do not.
| Deduplication | Data cleansing | |
|---|---|---|
| Question it answers | Is this the same entity as another row? | Is this value correct and well formed? |
| What it changes | Row count | Field values |
| Typical action | Merge or drop repeat rows | Fix formats, fill gaps, verify emails |
| Match key needed | Yes | No |
| Example | Two rows for Jane Doe become one | ACME.COM becomes acme.com |
| Order | After normalizing | Before matching |
Run cleansing first, then deduplicate. If you match before normalizing, Jane.Doe@Acme.com and jane.doe@acme.com look like two people.
When it matters
Duplicates cost money and reputation in four situations.
Cold email campaigns
A duplicate means one person gets the same sequence twice, from you, in the same week. That produces spam complaints and unsubscribes, and both damage the sending domain. Deduplicate the list before every send, not once a quarter. Check people across campaigns too, since a contact in two active sequences is a duplicate even when each list is clean.
Buying data from more than one source
Two providers usually overlap on the same senior people. If you pay per record, you pay twice for the same human. Compare the overlap on a small sample before you buy a second list. A sample of a few hundred rows tells you how much of the new list you already own.
CRM hygiene
Duplicate accounts split a customer's history across two records. Reps then work the same account in parallel, and reporting double counts pipeline. Use the company domain as the account key and merge on it. Subsidiaries that share a parent domain are the exception, so review those by hand.
Enrichment runs
Repeated keys in an input file cost repeated lookups. Remove duplicate emails, LinkedIn URLs and domains before you enrich, and the bill falls with the row count.
How LeadOcean handles it
LeadOcean deduplicates its exports, and gives you one field to check the rest yourself. These are described in the OpenAPI spec and the exports docs, checked September 2026.
- Exports. A
POST /v1/exportsfile is deduplicated byperson_id, up to 50,000 rows per export. - Records vs people.
person_idnames a record.person_group_idnames the person. Two records that share aperson_group_idare the same person as far as LeadOcean's de-duplication knows. - Where it appears.
/v2/people/reversereturnsperson_group_idindata. Compare it, never look it up.
LeadOcean holds more than one record for some people. The same human can come back under different person_id values depending on the contact point you searched with. So when you merge LeadOcean output with another list, key on person_group_id where you have it, then on LinkedIn URL.
The example below looks up the same placeholder person (the values are made up) by email and by phone. Each call costs 1 record. If both answers carry the same person_group_id, they are one person.
curl -s -G "https://api.leadocean.io/v2/people/reverse" \
--data-urlencode "email=jane.doe@acme.com" \
-H "x-api-key: $LEADOCEAN_API_KEY" | jq '.data.person_group_id'
curl -s -G "https://api.leadocean.io/v2/people/reverse" \
--data-urlencode "phone=+1 415 555 0133" \
-H "x-api-key: $LEADOCEAN_API_KEY" | jq '.data.person_group_id'LeadOcean does not run a rep workspace or a CRM, so it does not merge duplicates inside your CRM. Do that step in the CRM, or in a script on the exported CSV. LeadOcean prices at Free (1,000 records, one-off) and Pro at $499 a month, so a clean export costs the same as a messy one. See pricing. To compare sources before you buy, start with best B2B databases, or the UK-focused list if your market is the United Kingdom.
FAQ
What is the difference between deduplication and merging?
Deduplication finds the repeats. Merging decides what the surviving record contains. Most tools do both, but they are separate decisions, and a bad merge rule loses good data.
Is deduplication the same as email verification?
No. Verification asks whether an address can receive mail. Deduplication asks whether two rows are one person. A list can be fully verified and still hold the same person three times. LeadOcean publishes an email_status on each row, so you can check both.
Should I deduplicate by name?
No. Names repeat across thousands of people and vary in spelling for one person. Use LinkedIn URL, work email or company domain as the key, and use name only as a tie-breaker.
How often should I deduplicate a lead list?
Before every send and after every import. Lists drift as people change jobs and new sources arrive. If you buy from several sources, the data providers comparison helps you pick fewer.
Does LeadOcean remove duplicates across my own CRM?
No. LeadOcean deduplicates its own exports by person_id and exposes person_group_id so you can match. It does not read or clean your CRM.
Start with one clean record per person
Free to start. No credit card. 1,000 records to spend whenever you like.
Get your free API key →