Keyword Clustering Tool

Paste a keyword list and see which of them are one page. Clusters on shared stemmed terms and then splits anything that mixes intent.

One keyword per line. Add a comma and a monthly volume if you have one — backlink checker, 40000 — and it will be used to pick the primary keyword. Volumes are whatever you paste; nothing is looked up.
How much two keywords must share before they get the same page. Drag left for fewer, broader pages; drag right for more, narrower ones. There is no correct value — it encodes how much content you can afford to write.
Options

Nothing you type here is uploaded, stored or sent anywhere. It all runs in your browser.

The clustering is the easy half

Grouping keywords by how much vocabulary they share is arithmetic, and any tool can do it. The part that decides whether the work pays is the mapping: one cluster, one page, one intent. Get that wrong and you publish two pages that want the same result, Google picks one of them, and the other sits there earning nothing while still costing you crawl budget and internal links.

That is the actual failure mode. It is not a penalty and nothing warns you about it. You just notice, eventually, that the page you wrote second never ranked and the page you wrote first dropped a few places when it went live.

How this groups them

  • Normalise and stem. Every keyword is lowercased, split into words, stripped of stop words and lightly stemmed, so backlinks, backlink and backlink's are one token. The stemmer is deliberately gentle — it will not turn business into busi, because you have to be able to read the output.
  • Score every pair. Similarity is the Jaccard index over those stemmed token sets — shared tokens divided by total distinct tokens — plus a bonus when the two keywords share a head term, the last content word. Link building services and link building agency share two of four tokens; the head terms differ, so they land together only at a loose threshold.
  • Merge, then stop. Agglomerative clustering with average linkage: the closest two groups merge, distances are recalculated, and it repeats until nothing is above your threshold. No fixed number of clusters, because you do not know in advance how many pages a list is worth.
  • Split on intent. Clusters are then checked against the intent rules and broken up where members disagree. This is the step that stops a guide and a service page being planned as one URL.

The primary keyword is a decision, not an output

If you pasted volumes, the highest-volume member becomes the primary keyword and everything else is a supporting term. That is a reasonable default and it is wrong often enough to check: the highest-volume phrase in a cluster is sometimes the vaguest one, and a page targeting it competes with everything. When volumes are missing, this picks the shortest keyword that still contains the cluster's shared term, which is usually the head phrase.

What to do with the clusters

One cluster becomes one URL. The primary keyword goes in the title, the H1 and the first paragraph. The supporting keywords are the H2s and the sentences under them — not a list to sprinkle, but a checklist of things the page has to actually cover. A cluster of one keyword is a signal too: either it is genuinely its own page, or it belongs to a cluster you have not pasted yet.

Once the map exists, run the list through the keyword intent classifier to sanity-check the page types, then turn each cluster into an outline with the content brief generator.

Questions people ask

Why did two obviously related keywords end up in different clusters?

Because they share fewer words than your threshold requires, or because they were split on intent. Drag the threshold left and watch them merge. If they only merge at a very loose setting, that is information: they are related topically but not the same search, and they probably are two pages.

Does this look up search volume?

No. Nothing on this page makes a network request. Volume is whatever you paste after the comma, and it is used only to pick the primary keyword and total up each cluster. Export from your own keyword tool and paste the two columns.

Is clustering by shared words the same as clustering by SERP overlap?

No, and SERP overlap is the stronger method — it groups keywords by which URLs actually rank for them, which is Google telling you directly that two queries are the same job. It also needs live ranking data for every keyword, which is why it costs money. Term-based clustering is the free approximation, and it is close enough to plan with.

How many keywords can I paste?

Two hundred are clustered. Beyond that the pairwise comparison starts to make the page feel slow in a browser, so the list is trimmed and the page tells you. For a bigger list, cluster it in batches by topic — you were going to do that anyway.

What threshold should I use?

Start at 0.45. If you are getting more pages than you can realistically write, loosen it. If clusters contain keywords you would not put on the same page, tighten it. The threshold encodes how much content you can afford to produce, which is not something a tool can know.

A cluster still needs links to rank

Clustering tells you what to write and how many pages to write. It does not make the page competitive. That part is links.

Book a Call More free tools