Ultimate HubSpot Deduplication Guide
From prevention to automation: how duplicates enter HubSpot, what HubSpot handles natively, and how to clear the rest.
HubSpot ships its own duplicate management, and it still stops short of merging in bulk. That gap is why Dedupely <> HubSpotexists. This guide covers what HubSpot handles on its own, and what to do about everything it leaves behind.
A complete guide to duplicate management in HubSpot: where duplicates come from, what HubSpot handles, and how to clear the contacts and companies it does not.
Table of Contents
- HubSpot Duplicates 101
- Step One: Finding and merging existing duplicates in HubSpot
- Addressing merge risk
- Best practices to prevent merge disasters
- Finding duplicates in HubSpot
- Start merging duplicates one-by-one
- Using merge rules to save time
- Set the primary record
- Audit your merged HubSpot records
- “Help! I’m still seeing duplicates in HubSpot”
- How long should the initial HubSpot de-dupe take?
- Step Two: Preventing duplicate record entry in HubSpot
- Step Three: Automate
- Budgeting your HubSpot De-dupe
- In-house vs. HubSpot Marketplace Solutions
- API solution built by an in-house developer
- HubSpot Native Duplicate Management
- “Heck, it’s the sales team’s job!”
- Why HubSpot Marketplace solutions are a no-brainer
- And no! HubSpot doesn’t earn kick-back from their integration providers, as some might have you believe.
- What to look for in a de-dupe provider
HubSpot Duplicates 101
How duplicate contacts enter HubSpot
From our own data and what customers tell us, duplicates enter HubSpot these ways:
- Manually creating new records by hand without the proper checks beforehand.
- Improperly imported records or importing the same records more than once without including emails (the import de-dupe key) in the CSVs.
- Third-party integrations that don't run checks for existing contacts before creating new records.
- Web-form-2-lead setups that just insert and don't properly check for existing records.
- Incorrect in-house API implementations.
HubSpot is as vulnerable to duplicates as any spreadsheet or database.
Why HubSpot duplicates are costly and who should worry?
Duplicate records cost money and they reach customers. Both are worth your team's attention.
These roles feel it first:
- Sales reps
- Marketing or sales managers
- Sales and Marketing VPs, CMOs, Marketing Directors...
What that costs you:
- Makes sales reps have to wade through more data, contacting the same person twice by accident, wasting time researching and preparing for calls and follow-ups already made.
- Skews valuable numbers from attribution reports to financial accounts.
- Employees lose up to 50% of their time to routine data cleanup (MITSloan).
- Prevention costs $1, correction costs $10 and leaving duplicates cost up to $100 (SiriusDecisions).
- Hurts your campaign performance, brand and customer relationships.
Should I worry about deduplicating HubSpot today?
Correcting duplicates costs a fraction of leaving them. The sooner you start, the less goes out the window.
Some data should be duplicated. This can tell us how many return customers we have or repeat orders.
Data decay is like smoking, a dirty house or weight gain. The sooner you do it the more of a gain you get.
A study on lead data health shows that 15% of company sales leads are duplicated. Another study shows that companies can recover up to 70% of their revenue just on clean data alone.
Do not wait until three days before a campaign. Today is the right time.
What HubSpot currently does to prevent duplicates
Duplicates are common enough that HubSpot has safeguards against them.
Preventing duplicates with emails and company domain name
HubSpot doesn't allow you to add two contacts with the same email. Yes! This means you have fewer duplicates long term. This also means you're extra careful not to use company-generic (info@) email addresses that more than one employee uses.
For companies, Company Domain Name is the unique key that can't be used more than once. For example, you can't have two companies with example.com.
(Easily) Import de-dupe in HubSpot
Again, when importing in HubSpot, it's easy to import duplicates if you forget to add the de-dupe key (email for contacts or domain for companies).
We've covered this before but it's as easy as simply importing with the correct keys in your CSV or Excel.
Also, you can use existing HubSpot IDs to re-import back into HubSpot, updating existing records...
"You can use an object ID to specify any records that already exist in your CRM. All objects include contact, company, deal, ticket, and product. If you import an object that already exists in HubSpot, any matching properties will be updated with the latest data from your import." -- HubSpot
Using the HubSpot deduplication tool (for HubSpot Pro users)
HubSpot is one of the few CRMs that ships a duplicate finder at all. According to HubSpot, the duplicate finder uses AI to know which contacts are similar, improving the duplicate matches.
It finds duplicates periodically and lists them two at a time, paginated.
Manage duplicates is a Professional and Enterprise feature. On a Free account you will not see it.
It is a genuinely useful tool, and it stops short in a few places.
Two records per match. Duplicates arrive in any number.
No bulk merging. That is a defensible choice on HubSpot's part (the tool cannot tell a partial match from an exact one) but it leaves the work with you.
Step One: Finding and merging existing duplicates in HubSpot
The first job is the duplicates sitting in your account right now. Automation comes after that, and it is built on what you learn here.
Note: We're going to use Dedupely for the following examples. You can sign up for a free trial to see how this works on your account.
Addressing merge risk
Are two people with the same first and last name in real life duplicates of each other? No! However, unless your company has 100 million customers, you probably don't have to worry about people with the same name.
So we can start by preventing ourselves from making common sense mistakes with duplicate matches that are natural. Records sometimes have the same phone numbers, or other matching attributes.
In short, the more fields we use to match duplicates with the lower the chance of incorrect merging. However, the more fields we use, the fewer duplicates we're going to have. So there's a risk trade-off that, at some point, you'll have to make.
The amount of risk you take depends entirely on you and I urge you to proceed as risk-adverse as possible. Why? Because once done, it's nearly impossible to fully and quickly recover from large amounts of incorrectly merged contacts. With the amount of data that moves around in a merge, it's very hard from a technical standpoint to undo merges (we still haven't found a reliable way to do it, and therefor we don't).
Best practices to prevent merge disasters
Disasters happen, and they are all preventable. These are the habits that stop one:
- Always have backups of your data handy. Create a special backup before large bulk merges.
- Always test and audit your matching setup. It’s easy to make a mistake by rushing. Dedupely does everything possible to prevent users from making simple mistakes. However, always closely review before bulk merges and audit the changes after the merge.
- Be aware of how and when your data evolves. Adapt your match setup and merge rules accordingly to avoid collisions with changes in your HubSpot data.
- Be in-tune with errors and nuances in your data. Having defaults like “000-0000” in a phone number which would match all records with blank phone numbers. To a computer non-blank defaults don’t look blank. Dedupely does try to make sure each input passes validation of what it’s suppose to look like.
- Never automate merges that you haven’t tested and reviewed. Never run bulk merges until you’ve looked over a good portion of the duplicates and are confident of the results.
- Take attribution into account. Decide how lead and contact owners are preserved through merge rules. Sam’s lead might become Jack’s lead because Jack has a duplicate of Sam’s lead. Who wins the ownership in this case?
Finding duplicates in HubSpot
Start with the fields you want to match on. First and last name is enough to begin; add fields later to narrow what the search returns.
When the search finishes, read down the list of matches and look for anything that is not a duplicate.
A wrong match tells you how far to trust the Match Options behind it. Tighten them before you merge anything in bulk.
Start merging duplicates one-by-one
Merging one at a time at the start shows you which fields actually differ between duplicates, and which ones matter.
Manual Review shows every match before it merges, and lets you set which record is the Primary and which value wins each field.
Using merge rules to save time
Merge Rules make that choice for you. Set one per field: keep the newest value, take a value over a blank, or pin the field to the Primary.
That is what makes bulk merging and Auto Merge safe to run.
Set the primary record
The Primary is the record that survives a merge. The Secondaries are absorbed into it, and Merge Rules decide field by field which of their values come across.
Audit your merged HubSpot records
Merge History lists every merge and what it kept. Refresh the record in HubSpot to see the result there.
“Help! I’m still seeing duplicates in HubSpot”
Normal. No first pass catches every duplicate. Start by widening the Match Options.
Dedupely sets text fields to similar matching by default. Each matcher finds a different shape of duplicate.
Change it and new matches appear:
- Exact match is the strictest match type. It ignores uppercase and lowercase but must match the text exactly.
- Similar match is also fairly strict but ignores punctuation marks among other commonalities that deem removable.
- Match first similar word/match last similar word both work to match the first or last words. This can work well for company names or first names that have the middle name added.
- Fuzzy match is the most aggressive match type and works roughly on sound-alike matches. However, this is the most inaccurate match type and should never be relied on to produce correct matches.
Then look at the prefixes and suffixes in your data. Matching improves once you start ignoring common terms. By adding ignored terms you reduce the extra noise preventing proper matching.
How long should the initial HubSpot de-dupe take?
Anywhere from a few hours to a few weeks, depending on how much data you have and what shape it is in.
Step Two: Preventing duplicate record entry in HubSpot
Preventing duplicates costs far less than fixing them, and fixing them costs less than living with them.
HubSpot blocks some of them at the door. Nothing blocks all of them.
Review the common duplicate sources
Knowing where duplicates come from is what stops the next batch.
If your developers write against HubSpot's APIs, ask them whether their code creates records without checking first.
Check every third-party app that creates records: web-to-lead forms, and any integration that syncs across platforms without a duplicate check.
Educating your team on data etiquette
Anyone who enters or edits records in HubSpot should know:
- How a field should be formatted
- How to check for a record before creating one
- Which tool to use to find or merge a duplicate
- How merging affects sales attribution
A team that understands this stops most duplicates before they exist.
Step Three: Automate
Step one is done and step two is in place. Now automate.
Prevention is the best way to solving most your duplicate problems (and maybe all problems). However dupes are inevitable and automation will pay dividends over the days, weeks, months and years your team is interacting with HubSpot. They will thank you as well!
Use auto merge to pick up daily duplicates
Auto Merge clears the obvious duplicates so nobody has to look at them.
You still review week to week. Dedupely flags what the non-automated searches catch, and you decide on those.
Between your team and a few preventative measures, you stay ahead of it.
Budgeting your HubSpot De-dupe
How de-dupes are priced
Record count is the largest factor in what any vendor charges.
Ask each one how they count records, whether the price is annual or one-off, and what happens when your database grows.
Dedupely bills on records synced and charges no per-user fee.
Costs are also influenced by:
- the level of customization of your HubSpot account
- consulting bases instead of self-serve
- the amount of time required to complete the de-dupe
Vendors price differently enough that the only comparison worth making is a quote against your own record count.
Realistic time frames for initial HubSpot de-dupe
Time scales with record count the same way pricing does.
A small account takes an hour or so. A large one takes days, because there is more data, more care needed, and more shapes of duplicate to catch.
Start weeks before a deadline, not hours. What decides how long it takes:
- How fast whoever you contract works.
- The testing and care it takes to avoid losing data.
- Errors already in the data that have to be fixed first.
- Time to sync and move records around.
The expensive ones are the surprises that land on the day of a launch. Be safe, do it ahead of time.
Who should be involved in the de-dupe?
Whoever owns the data, the sales managers, and anyone whose reporting changes when records merge.
How do I calculate the ROI of de-duping HubSpot?
Plenty of studies show what duplicate data costs. None of them justify a line in your budget.
The simplest measure is the hours your team loses finding and merging duplicates by hand. You can measure the amount of time it takes to find one group of duplicates and merge them by hand. Then ask your reps how many times a day/week they have to fix duplicates. Multiply those together and you have an hour figure over a certain period of time. How much are those hours costing your company instead of benefiting it?
The calculation could look like this:
((Number of duplicates merged by reps daily) * (minutes to merge a duplicate / 60) = (lost hours)) * (hourly base pay)
Time one person merging one group of duplicates. Multiply by the groups you have, then by the number of people doing it.
(20 * 4 minutes = 80 minutes) * 30.00/hr = $39 per day
Put your own hourly cost against that total and you have the monthly price of leaving the duplicates where they are .
Add lost opportunities if you want to. If the number still looks small, you do not have a duplicate problem. If you do it should be pretty easy to convince your team and peers to get on board with you in your de-dupe efforts.
This is an estimate, not an audit. It is enough to decide whether the spend is worth making.
In-house vs. HubSpot Marketplace Solutions
Everyone has been down the DIY rabbit hole and ended up paying someone else anyway.
Then there are scenarios where in-house just makes so much sense.
What are the options we have for effective de-duping?
API solution built by an in-house developer
You have developers. Why not put them on a solution that fits every requirement exactly? Not so fast!
Some developers love writing their own apps. Believe me, at one point I would have designed our entire software stack–and nearly did–before I was reminded of the true costs of in-house.
Your developers are there to build what you sell. The first version takes 1) time of initial development 2) time of bug fixes and specs back and forth 3) maintenance and updates over the years 4) costs of hosting and running.
Then it needs maintaining every time the CRM's API changes, and it is months before it runs unattended.
HubSpot Native Duplicate Management
HubSpot has more native duplicate management than most CRMs.
Where it stops, as covered above:
- You cannot choose how fields are matched. HubSpot decides.
- Nothing merges automatically, and nothing merges in bulk.
- A match holds two records.
- Overall somewhat cumbersome
The HubSpot native duplicate management, while a cut above other CRMs, still misses the mark in terms of time saving duplicate management.
“Heck, it’s the sales team’s job!”
When was the last time “data cleansing and cleaning up duplicates” was in the job description of a sales rep? Do any of your sales reps think their job is cleaning up bad data?
Training personnel to clean up after themselves is pretty basic. Everyone should actively clean up the data as they enter it into HubSpot. It’s just proper etiquette not to mention respect for fellow team members.
However if you’re not aiding sales in gathering, handling and cleaning data you’re making their lives harder, which in turn make sales harder and mean less sales for your company. By “aiding” I mean providing them with the proper tools to do their job.
Why HubSpot Marketplace solutions are a no-brainer
And no! HubSpot doesn't earn kick-back from their integration providers, as some might have you believe.
There are a few de-dupe apps on the HubSpot Marketplace. Some of them are not so great and others are pretty awesome. These apps are built by companies that dedicate large amounts of resources on building solutions to save you insane amounts of time and money.
You skip the DIY detour, the internal project that pulls your team off customers, and native features that stop short of merging.
What to look for in a de-dupe provider
Most listings on the HubSpot Marketplace handle the basics. What separates them is what happens during the merge itself.
With a consultancy, make sure they understand your fields. Take nothing for granted: they will not know what X and Y are for, or that dropping one costs you a report.
Self-serve puts your team in control of the whole process. Expect a learning curve and some documentation.
Whatever you choose should do all of this:
- Match on any field, exact or similar, with common terms ignored.
- Allow you to merge in bulk, customize merges and automate merges once you’ve done the homework and are confident in your settings.
- Merge Rules that cover Primary selection, attribution and your custom fields.
- Tell you about new duplicates as they appear in HubSpot.
- Answer support quickly.
If self-serve does not fit, our team sets it up with you at no extra cost. Take a closer look at the HubSpot <> Dedupely connector here.
Contact us
We’d be happy to help you get this set up.
Write us a message
We probably know the answer to your question already
Book a Zoom
Whether you’re getting started or getting intense.
Get in touch!
Discover Related Blog Posts
Stay updated with our latest articles and insights.
Get started for free
20 free trial merges and as much free support as you need to get your duplicates under control.
H Chitty
HubSpot User
Alina T
Sales Operations Specialist, FieldBee
R McNaught
HubSpot User
John K
Head of Marketing, David J Anderson School of Management
Samo J
Founder & CSO, TapHome
M Weppner
Yuliya
Pipedrive User
D Wright
HubSpot User
Shawnee K
Salesforce User
DiBlasio S
HubSpot User
Laura R
Salesforce User
Emily K
Mercy Housing
Paweł S
Pipedrive User
Grattan H
Pipedrive User
Scott B
VP Platform Ecosystem, HubSpot
Burchard J
HubSpot User
Sean B
Managing Director, Legal CPD
Isaac J
Salesforce User
Larry D
Pipedrive User
Andy G
Pipedrive User
Simon W
Tillhub
Allan R
Co-founder & Managing Director, Target3D
Marco S
Information Systems Manager, Efecte
A Team
HubSpot User
J Eddie
HubSpot User
A Grogan-Crane
HubSpot User
Wasmer D
HubSpot User
Running CRM cleanup for clients? See the Partner Program








