RAKE Keyword Extractor

Paste any document and pull out its most relevant keywords and keyphrases using the RAKE (Rapid Automatic Keyword Extraction) algorithm — ranked by score, no training data needed. Runs in your browser; nothing is uploaded.

Keywords

About this tool

The RAKE Keyword Extractor finds the most relevant keywords and keyphrases in any document using RAKERapid Automatic Keyword Extraction. RAKE is an unsupervised, language-agnostic algorithm: it needs no training data and no model, so it runs instantly and entirely in your browser. Nothing you paste is uploaded.

How RAKE works

  1. Split into candidate phrases. The text is broken into runs of content words at every stopword (the, of, and, …) and punctuation mark. Each run becomes a candidate keyphrase.
  2. Score each word. A word co-occurrence graph is built over the candidate phrases. Every word gets a score of degree ÷ frequency — words that appear in longer phrases and co-occur with many others score higher.
  3. Score each phrase. A phrase's score is the sum of its member word scores, so meaningful multi-word phrases naturally rise to the top.
  4. Rank. Phrases are sorted by score, highest first.

Options

Good for

FAQ

Does it work on languages other than English?

Partially. The RAKE algorithm itself is language-agnostic, but the built-in stopword list is English (the NLTK-style list RAKE was published with). On other languages phrases still get split at punctuation and scored, but common function words won't be filtered out, so expect noisier, longer candidate phrases.

Why do long phrases dominate the top of the list?

That's inherent to RAKE: a phrase's score is the sum of its member word scores, and each word's score (degree ÷ frequency) grows when it appears in longer phrases. If you want short, tag-like keywords, set Max words per phrase to 2 or 3 — longer candidates are dropped before scoring.

What does the score actually mean? Can I compare it across documents?

The score is the sum of degree÷frequency values for the words in the phrase — a purely relative ranking signal within one document. A score of 9 in one article and 9 in another don't mean the same thing, so use the ordering (and gaps between scores), not the absolute numbers.

Why are the extracted phrases all lowercase?

The text is lower-cased and whitespace-normalized during tokenization so that "Machine Learning" and "machine learning" count as the same phrase. Words keep internal apostrophes and hyphens (don't, state-of-the-art), while any other punctuation acts as a hard phrase boundary.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool rake-keywords "Paste a document to extract keywords from…"

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/rake-keywords/?text=Paste%20a%20document%20to%20extract%20keywords%20from%E2%80%A6&top_n=10&max_words=0

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.