Duplicate Row Finder

Find duplicate rows in a CSV or delimited table, group them with their line numbers, and see which columns drive the duplication. Runs entirely in your browser — nothing is uploaded.

Try:
Duplicate report

Find duplicate rows in a CSV or table

Paste a CSV (or any delimited table) and this tool finds the rows that repeat, groups them together with their line numbers, and tells you which columns are driving the duplication. It is read-only: it reports the duplicates, it does not delete or rewrite anything. Everything runs in your browser — nothing is uploaded.

By default two rows match only when the whole row is identical. Give one or more key columns to match on just those (e.g. treat rows with the same email as duplicates even when the name differs). Turn on Ignore letter case and Ignore extra whitespace to catch near-duplicates that differ only by casing or stray spaces.

Worked example

Input (with the header row on):

name,email
Alice,[email protected]
Bob,[email protected]
Alice,[email protected]
Carol,[email protected]
Bob,[email protected]

Report output:

Scanned 5 data rows (2 columns), keyed on the whole row.
Found 2 duplicate groups covering 4 rows; 1 row is unique.

Duplicate groups (most repeated first):
  x2  Alice | [email protected]   (lines 2, 4)
  x2  Bob | [email protected]   (lines 3, 6)

Columns driving duplication (repeated rows / distinct values):
  name: 4 / 3
  email: 4 / 3

Line numbers count from the top of the input (the header is line 1), so you can jump straight to each duplicate in your source file.

Options

Limits & edge cases

FAQ

Does it remove the duplicates?

No — this tool only finds and reports them. Choose the Duplicate rows only (CSV) output to export the offending rows, or use the csv-dedupe / fuzzy-dedupe tools when you want the cleaned data back.

What does "columns driving duplication" mean?

For every column it counts how many rows carry a value that also appears in another row (repeated rows / distinct values). A column near the top of that list is where the repetition is concentrated — for example a repeated email while every id stays unique. It helps you pick the right key column to match on.

Can I match on more than one column?

Yes. List the key columns comma-separated, as header names (email,company) or 1-based indices (1,3); mixing names and indices works too. A row counts as a duplicate only when all the key columns match an earlier row.

How are near-duplicates handled?

With Ignore letter case and Ignore extra whitespace on (both default on), Alice matches alice and New York matches New York. That covers casing and spacing differences; it does not merge misspellings — for that, try the fuzzy-dedupe tool.

Is my data uploaded?

No — it's processed locally with WebAssembly. Nothing leaves your browser.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool duplicate-row-finder "name,email
Alice,[email protected]
Bob,[email protected]
Alice,[email protected]"

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/duplicate-row-finder/?data=name%2Cemail%0AAlice%2Ca%40x.com%0ABob%2Cb%40y.com%0AAlice%2Ca%40x.com&columns=email&header=true&delimiter=%2C&ignore_case=true&ignore_whitespace=true&output=report

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.