Remove accidentally repeated words from text

Catch the doubled words a spell-checker walks straight past — "the the", "is is", a word repeated across a line wrap in scanned or hard-wrapped text. Legitimate English repeats like "had had" and "that that" are protected by an editable keep list, and you can preview every deletion struck through before you accept it.

Try:
Result

About this tool

repeated-word-remover deletes the words you accidentally typed twice. It is the class of typo a spell-checker never flags, because every word in I think the the cat sat down is spelled correctly — only the pair is wrong.

The scan is deliberately adjacent-only. Two identical words next to each other are a candidate; a word that simply reappears later in the sentence is left alone, because that is normal English, not a typo. Within a run of repeats the first occurrence always wins, so its capitalisation, indentation and the punctuation around it survive untouched.

Worked example

Input:

I think the the cat sat on on the mat.

Cleaned output:

I think the cat sat on the mat.

Switch Result view to Marked-up changes and the original text comes back with each deleted copy struck through, so you can check the edit before you accept it:

I think the ~~the~~ cat sat on ~~on~~ the mat.

Switch it to Audit report and you get counts plus a line and column for every spot found.

Repeats that are meant to be there

Plenty of correct English doubles a word: He had had enough, the fact that that happened, what it is is simple, a long long time ago. These are protected out of the box by the Never collapse these words list, which starts as:

had, that, is, do, no, very, long, many, far, ha, blah, bye, night, so, chop, tut, yum

Add your own words to it, or clear the list entirely if you want every repeat collapsed regardless.

Options and limits

FAQ

Why did it leave "had had" and "that that" in my text?

Because those are grammatical. He had had enough uses the past perfect, and the fact that that happened uses that as a conjunction and then as a determiner. Both words are in the Never collapse these words list by default. Remove them from the list — or clear it — if you want them collapsed anyway.

Will it remove a word that appears twice in the same sentence?

No. Only adjacent repeats are collapsed. In the cat sat on the mat, the second the is separated by other words, so it is normal English and stays. If you want whole-line or whole-list deduplication instead, that is a different job — use a duplicate-line or list-dedupe tool.

Can I see what would change before applying it?

Yes. Set Result view to Marked-up changes. You get your original text back with every copy that would be deleted wrapped in markdown strikethrough (~~the~~), so nothing is removed until you decide. Audit report goes further and lists the line and column of each spot along with before/after word counts.

Does it handle a word doubled across a line break?

Yes, and that is on by default. Text that was hard-wrapped or run through OCR often ends one line with a word and starts the next with the same word. Turn Catch repeats split by a line break off if you want repeats confined to a single line. A blank line between the two words is treated as a paragraph break and is never collapsed.

What happens to my capitalisation, spacing and punctuation?

The first occurrence in a run is kept exactly as you wrote it, and everything from the end of that word to the end of the run is deleted. So The the cat becomes The cat, indentation on a list item is preserved, and the trailing punctuation after the last copy stays attached.

Is my text uploaded anywhere?

No. The same Rust core runs locally in the WebAssembly page and in the CLI, so your text is processed in your browser or terminal and never sent to a server.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool repeated-word-remover "I think the the cat sat on on the mat."

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/repeated-word-remover/?input=I%20think%20the%20the%20cat%20sat%20on%20on%20the%20mat.&output=clean&keep_words=had%2C%20that%2C%20is&case_sensitive=true&across_line_breaks=true&ignore_punctuation=true&include_numbers=true&min_length=1

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.