The short version
- Clean data means a different problem in every system and every department, so define it by the fields your next piece of work actually depends on.
- The figures everyone quotes for what bad data costs do not survive being opened. The error rate is the part anyone has actually counted: 47% of newly created records carried at least one critical error.
- The real choice is who owns the cleaning, everybody or one person. Tooling reads like a third answer, but it is what the owner does the work with, not an owner itself.
- A quarterly cleanup works. Its weakness is the month before the next one.
To start with, what does clean data mean? Cleaning data can mean clearing a database or system of invalid fields, duplicate accounts, wrong or missing data, an overload of data, and spelling errors. It takes on many different meanings depending on the type of business, the department it's being mentioned in and where the data sits. Data can sit across many platforms like website analytics, inventory management, HR systems, SRMs and CRMs, and marketing engagement tools. Pretty much every business uses and relies on data in some form.
Businesses live and die by their data quality and quantity. Every business wants clean and reliable data in every system they use, as it allows informed decision making up and down the channels. Quantity alongside quality is also important: if a business does not have enough data, they cannot make accurate decisions. So clean data is about both the quantity and the quality. By cleaning your data, you then get an accurate measurement of the quantity you actually have. Both go hand in hand.
What's the impact of unclean data on a business?
Bad and unclean data can be dangerous for a business to have. It can send an entire marketing campaign in the wrong direction, make wrong product decisions, report inaccurately, and it becomes hard to trust your data once you spot a few bad examples. This costs money, time and trust in the short and long run.
If you ask anyone in an organization, they will say the main place bad and unclean data lives in any business is in the CRM. Some examples of it:
- Sales team metrics and KPIs take a hit on outbound. Let's say each rep has a list of accounts to go outbound to, and 20% of that data has the wrong number or the wrong website for the company. That's extra time wasted from the rep's point of view and a loss of activity.
- It is not just marketing's problem. Marketing person running a mailing campaign to a list, if 15% of those email addresses are unverified, you risk tanking your domain rating from a deliverability perspective, and everybody else in the business sends from that domain too. Google asks bulk senders to keep their spam complaint rate under 0.3%, which is a narrower margin than most lists are built to.
- Dashboards give wrong readings. Most leaders rely on dashboards to do the heavy lifting and to keep them informed. If the data funnelling into the dashboard is bad, so are the readings, and off that, bad decisions are made.
You would expect there to be a good figure for what all of this costs, and there is not one. The two most quoted are Gartner's, that poor data quality costs an organization at least $12.9 million a year, and IBM's, that it cost the US economy $3.1 trillion in 2016. Neither comes with a method you can check. Gartner attributes its number to its own 2020 research, which it has not published. The IBM one appeared in a Harvard Business Review column with no workings behind it, and it works out at roughly a sixth of US GDP that year.
What does hold up is smaller and more useful. In 2017 Tom Redman, Tadhg Nagle and David Sammon asked 75 managers to take 100 records their own department had recently created, and mark each one against the handful of fields that department actually depends on.
- Records checked per departmentrecently created, not a sample of the whole database
- 100
- Records carrying at least one critical erroraveraged across the 75 departments
- 47%
- Departments whose score was acceptableon the loosest standard the authors would accept
- 3%
That is not a database quietly rotting. It is a measure of data being created wrong, today, by the people whose job it is to create it.
So how do you clean data across an organization?
Cleaning data can be done in many ways. How you do it depends on timing and budget. Either you keep it running in the background or you do it in one go, and the ongoing version splits into three. They work together more than they compete.
1. Set the culture
Ingrain in every user the habit of checking their data quality and accuracy at all points. Salespeople must keep their lists accurate and up to date at all times, marketing must do the same. If every channel of the business is focused on not letting bad data creep in, it's easier to spot when bad data appears, and you have multiple people flagging it.
Upsides
- Data first culture means with everyone focused on data, you can be sure the standard is the same throughout the organization, accurate data is flowing and users are ready to throw out bad data.
- Continuous improvement.
Downsides
- It can take away from other responsibilities and be used as an excuse for not fulfilling other duties.
- Hard to enact, especially at larger organizations.
- Requires the correct tools for everyone with a data first responsibility, for example enrichment, data verification and cleanup.
- Takes time to show effectiveness.
2. Dedicated role
By having a team or individual dedicated to keeping the data clean, it then means the responsibility falls on them to check, clean and verify data across an organization.
Upsides
- Easier to manage data when there's an assigned responsibility.
- Likely an expert in the field of data, which means they or the team become the source of truth across the organization.
- Easier to make serious improvements to both the quality and quantity of an organization's data with a dedicated role or team.
Downsides
- Expensive to do, and their ROI isn't directly distinguishable, as other departments will pick up the benefit.
- Everybody else stops owning it. The moment one person is accountable for data quality, it is nobody else's problem, which is the opposite of what the first option is trying to build.
- It becomes a queue. One person cannot keep pace with every team at once, so the work gets prioritized and some of it never comes up.
3. Technology and AI
There are plenty of tools out there, AI ones included, that can help keep your data clean and verified within your CRM and other systems. This option is not really a third owner though. Whoever owns the cleaning, a whole culture or a single person, still needs something to do the work with. The tools divide by the job they do:
- Verification. Checking that an address, a phone number or a company still exists before anybody acts on the row.
- Deduplication and merging. Finding the same person or company written two or three ways, and deciding which version survives.
- Standardization. Making one field say the same thing in every row, so that reports group the way you expect.
- Enrichment. Buying in what you never had in the first place, usually priced per record.
- Monitoring. Watching for bad data as it arrives, rather than finding it a quarter later.
Upsides
- Eases the pressure of responsibility.
- Able to process high levels of data and connect different sources together.
- Speed. It's much quicker to have technology and AI processing data than a human.
Downsides
- Expensive to run and operate at scale.
- AI can manipulate data and invent data, so it still needs close supervision.
- Set-up time can take a while, and so can getting the business used to it.
- Requires maintenance, which means a person or a team to adjust it for any new conditions.
By doing any one of these three, you are making sure your data is being cleaned, and you can trust what you are working with, to some extent. By implementing all three, you have an ecosystem which means your data is constantly being prioritized.
One-off cleanup
Some organizations don't make cleaning an ongoing priority, or they do it in conjunction with the three above. What this involves is, say every quarter, doing a cleanup across their organization whether it be in sales lists, marketing lists or HR records. This then becomes the responsibility of each department to make sure their data is fully accurate and in check. The reason this can work is because it acts almost like an end of financial year for each team and department to get their data sorted.
Another example of a one-off cleanup is during a CRM migration, or any system migration. This is usually smaller scale, and it is the one moment where nobody argues about whether the cleaning is worth doing. The hidden gap in migrating your CRM data covers what gets missed even then.
Upsides
- Much more cost effective as each department has ownership over their own data.
- A step towards setting the culture, without committing to it.
- Easier to see gaps during the cleanup per department.
Downsides
- Data goes out of date quickly, and waiting too long between one-off cleanups can mean there's a period where the data is not optimal. You do a cleanup every quarter and by month three, the data is already diminishing or diminished.
- The tools to clean data and verify it are still required in some capacity, which has a cost and a training angle that comes with it.
- When cleaning isn't directly tied to employees' targets, it can mean a rushed job is done on any cleanup. For example, a salesperson who has targets to hit would rather be selling than cleaning data, even if it meant they would be working with better data.
Cleaning what is already in the CRM
Of the three ongoing options, tooling is where we sit, and the half of the job we take is the harder one: the data already in the system rather than the data about to go into it.
A CRM works a record at a time. That is the right shape for the job it does every day and the wrong shape for correcting tens of thousands of rows at once, which is why cleanup inside one either never happens or gets bought in as a fixed price project.
So the round trip goes out and back. You export what is in there and work on it as a file, which is the one thing a spreadsheet does better than a CRM: the same correction applied to every row at once. You find the duplicates across the whole export rather than pair by pair, check every address against the mailbox instead of assuming it, and standardize the fields you report on in the same pass. Then it goes back over the originals, matched on email so an existing record is updated and not duplicated.
What the round trip cannot do is merge or delete the duplicate records already sitting in your CRM, because a push updates a record and cannot collapse two of them into one. That part stays an in-CRM job. How to clean up your CRM data is that round trip end to end, and CRM data quality: what to measure and what to fix first is how to decide what is worth fixing before you start.
Health Check, deduplicating, merging and the CRM Formatter all run on the free plan. Email verification beyond a quick clean, and pushing the cleaned file back to the CRM, begin on Starter. The pricing page has the split.
None of that answers who owns the job, which is the question this post is really about. What a tool changes is how big that job is, not whether somebody has to do it.
So overall, whether you do an ongoing cleanup or a one-off policy or both, it is good practice to clean your data and put the steps in place to be able to achieve that clean data title.
Questions
What does clean data actually mean?
It means the data is accurate, current, consistent and complete enough for whatever you are about to do with it, which is why the definition changes with the department. To a sales team it usually means the phone number connects and the account is not in there twice. To a marketing team it means the address is real and the person still works there. To a finance or reporting team it means one field says the same thing in every row, so the totals group correctly. The useful version of the question is not whether your data is clean in the abstract, but whether it is clean on the handful of fields the next piece of work depends on.
How much does bad data cost a business?
Nobody has a defensible number, and the widely quoted ones do not hold up. The $12.9 million a year figure is Gartner's, attributed to unpublished research of its own, and the $3.1 trillion figure is an IBM estimate for the US in 2016 that appeared without workings and amounts to roughly a sixth of US GDP that year. What can be measured is the error rate. When 75 managers each scored 100 records their own departments had recently created, 47% of those records carried at least one critical error, and only 3% of the departments scored acceptably on the loosest standard. The cost is real; the headline figures are not evidence of its size.
Should one person own data quality, or everyone?
Both work and they fail in opposite directions. Give it to one person or a small team and it gets done properly, but it becomes a queue that cannot keep pace with every department, and everybody else stops treating it as their problem. Spread it across the business and bad data gets caught closer to where it is created, but it is slow to take hold, hard to enforce in a big company, and easy to drop when targets are tight. Most organizations that get this right do both: a data first expectation on everyone, and somebody whose actual job it is to hold the standard.
How often should you clean your CRM data?
Quarterly is the common rhythm and it works, as long as you know what it leaves exposed. A quarterly cleanup behaves like an end of financial year for each team, which is useful because it creates a deadline that would not otherwise exist. The weakness is the month before the next one, when the data has been decaying for the best part of a quarter and everybody is still working from it. If the data feeds something continuous, like outbound sending or lead routing, the checks that matter most are worth running as records arrive rather than in a batch every ninety days.
Can you clean data that is already inside your CRM?
Yes, but not one record at a time, which is the only way a CRM is built to work. The practical route is a round trip: export what is already in there, do the work on the file, and send it back over the original records. A file is the one place you can apply the same correction to every row at once, compare every record against every other record for duplicates, and check whether an address is still live, which is something a CRM cannot tell you at all. That is what Operelio does. The limit is that a push updates records rather than merging them, so collapsing two existing CRM records into one stays a job for the CRM itself. The cleaning tools run on the free plan, and pushing the cleaned file back to the CRM starts on a paid one.
Does AI clean data reliably enough to leave alone?
Not on its own. AI is good at the parts that are judgement calls at scale, like deciding whether two differently spelled company names are the same company, and it is fast in a way no person can match. What it also does is change values and invent plausible ones, and a wrong value that looks right is worse than a blank field, because nothing downstream flags it. Let it do the work and have a person approve the result, and keep the fields your process actually depends on under review.
Written by

Daniel Pank, Founder
He spent seven years leading commercial and operations teams at a B2B outbound agency, running prospecting programmes for enterprise sales teams and building the systems underneath them, including an in-house CRM and a sales data consultancy. Operelio comes from years of working with data providers, and watching good data leave one and land in a CRM in worse shape than it left.