Tools·2 min read

Deduplicate

Find and remove duplicate rows from your spreadsheet. Choose which columns define a duplicate, match exactly, smartly or by similarity, and keep either the first or last occurrence.

By Operelio team

On this page8
  1. 1.What it does
  2. 2.How to use it
  3. 3.What the results show
  4. 4.Exact, smart and similarity matching
  5. 5.Matching options
  6. 6.Duplicates export
  7. 7.Run it automatically with a trail
  8. 8.Frequently asked questions

What it does

Scans your file for rows that share the same values in one or more key columns, then removes the duplicates. You choose which columns to check and whether to keep the first or last occurrence of each duplicate group.

How to use it

1

Upload your file

Any supported format (.xlsx, .xls, or .csv).

2

Pick key columns

Select one or more columns that define a unique row. For example, "email" alone, or "first_name" and "last_name" together.

3

Choose which to keep

First occurrence (default) or last occurrence.

4

Set matching mode

Exact is on by default. Smart treats emails, phone numbers and company names that mean the same thing as the same. Similarity adds a threshold slider for catching typos. All three on every plan.

5

Run

Click Run. The output contains only unique rows.

What the results show

After running, the results page shows the total input rows, unique rows kept, duplicates removed, and the number of duplicate groups found (a group is two or more rows that match on your key columns).

You also see the settings that were used: which columns Operelio matched on, which keep strategy was applied (first or last), whether case was ignored, and whether whitespace was trimmed. If you ran in similarity mode, the threshold percentage shows too.

A separate file with just the removed rows is available for download alongside your cleaned output, so you can review what was dropped before importing.

Exact, smart and similarity matching

Exact matching (available on all plans) compares cell values character by character. Two rows are duplicates only if the key columns match perfectly.

Smart matching (on every plan) reads what kind of column the key is and compares the meaning. Gmail addresses match with dots and plus tags ignored, so john.doe+shop@gmail.com and johndoe@gmail.com are one mailbox. Phone numbers match whatever the punctuation, with a US or UK country code dropped, so +1 (415) 555-0123 and 415-555-0123 are one number. Company names match with suffixes like Ltd, Inc and GmbH stripped, so Westmarch Trading and Westmarch Trading Ltd are one company. Any other column is compared exactly.

Similarity matching (on every plan, including Free) uses a similarity score from 0 to 100. You set a threshold with a slider, and any two rows with a similarity score at or above that threshold are treated as duplicates. This catches typos and minor variations like "Jon Smith" vs. "John Smith". The score measures how many character edits separate the two values, so closer values score higher.

The default threshold is 80, which works well for matching contact names. Go higher (90+) if you want to be more conservative and only catch very close matches. Go lower (70) for more aggressive matching that catches bigger variations.

Matching options

OptionDefaultWhat it does
Case insensitiveOnTreats "WESTMARCH" and "westmarch" as the same value.
Trim whitespaceOnStrips leading and trailing spaces before comparing.
Full row matchOnDefault scope. Compares every column. Switch to specific columns to match on a subset (for example, email only).
Blank or unreadable email addresses never count as a matchOn, for an email keyTwo rows with no usable email address are different people unless every other column matches too. On a company key the switch reads "Blank values never count as a match". Not used under similarity or a full-row match.

Duplicates export

On every plan, Operelio writes a second file containing just the rows that were removed. This lets you review what was dropped without re-running the job. If no duplicates were found, no extra file is created.

Run it automatically with a trail

Once the settings are right for the files you get every week, a trail can run them for you. A trail watches a folder and removes the duplicates, with the same match settings you chose here, on every file that lands there, then pushes the result into HubSpot, Salesforce or Pipedrive, with the seven safety checks and approval built in. Set it up once and the cleanup, the verification, the lead routing and the CRM import are automated from then on.

Trails run on Starter and above.

Frequently asked questions

What if I want to match on every column?

Pick "Entire row". That's the default. It compares all columns and treats two rows as duplicates only if every cell matches. Switch to "Specific columns" if you want to match on a subset, like email only.

What if my file has more rows than my plan allows?

Operelio processes the first N rows up to your plan limit and shows a banner with the total, the limit on your plan, and how many rows were processed. Free is 2,000 rows, Starter is 50,000, Pro is 75,000, Agency is 150,000. To process more, upgrade your plan or use the Split File tool to break the file into smaller chunks.

What happens to my original file?

Nothing. Operelio reads your upload and writes a new file. Your original is untouched. The cleaned output is a separate download.

Can I use this in a workflow with other tools?

Yes, on Starter. Workflow Builder lets you chain Deduplicate with other tools in one job. A common combination is Clean Headers, then Deduplicate, then CRM Formatter, all in a single run.

What plan do I need?

Every matching mode, exact, smart and similarity, plus the separate file of removed rows, is on every plan including Free. What changes is size: Free handles 2,000 rows per job, Starter ($39/month) 50,000, Pro ($99/month) 75,000, Agency ($299/month) 150,000.

Ready to get started?

Upload a file and run your first transformation. Free, no credit card required.