Data Binning Tool

Bin a numeric CSV column into equal-width, quantile (equal-frequency), or custom-edge buckets and label every row with the bucket it lands in — choose the method, bucket count, labels, and interval style. Runs entirely in your browser, nothing is uploaded.

Try:
Binned CSV

Bin a numeric column of a CSV into buckets

Paste a CSV, pick one numeric column, and this tool sorts each row into a bucket and labels it — locally in your browser, nothing is uploaded. Choose how the buckets are drawn:

Label each bucket yourself (Custom bucket labels, one per bucket) or let the tool auto-label them as an interval range (like (50, 75]) or a 1-based bucket index. Right-closed intervals decides whether a boundary value falls in the lower or upper bucket, and Output mode either appends a new <column>_bin column or replaces the source column with the label.

Worked example

With the default Equal-width method and 4 buckets on the score column, this input:

name,score
Ann,0
Bo,40
Cy,70
Di,100

becomes:

name,score,score_bin
Ann,0,"[0, 25]"
Bo,40,"(25, 50]"
Cy,70,"(50, 75]"
Di,100,"(75, 100]"

score runs 0100, so four equal-width buckets have edges 0, 25, 50, 75, 100. 40 lands in (25, 50] and 70 in (50, 75]. The first bucket is written [0, 25] because the lowest edge is always included; interval labels that contain a comma are quoted by the CSV writer.

Limits & edge cases

FAQ

What is the difference between equal-width and quantile binning?

Equal-width cuts the value range into buckets that each span the same distance — with bins = 4 over 0100 you get 0–25, 25–50, 50–75, 75–100 regardless of how many rows land in each. Quantile (equal-frequency) instead moves the edges so every bucket holds roughly the same number of rows; it's the better choice for skewed data because no bucket ends up nearly empty.

How do I set my own bucket boundaries?

Choose Custom edges as the method and type strictly-ascending, comma-separated numbers under Custom edges — for example 0,18,65,120 makes the buckets [0, 18], (18, 65], (65, 120]. Any value below the first edge or above the last one gets a blank label. Pair it with Custom bucket labels (like child,adult,senior) to name each band.

What do the labels like "(50, 75]" mean?

They are interval notation for the bucket's boundaries. A square bracket [ or ] means the endpoint is included; a round bracket ( or ) means it is excluded. So (50, 75] covers values greater than 50 up to and including 75. Toggle Right-closed intervals off to flip this to [50, 75), or switch Auto-label style to Bucket index for a plain 1, 2, 3.

Which value goes into which bucket at a boundary?

A value exactly on an edge follows the Right-closed intervals setting. When it is on (the default), intervals are (a, b], so a boundary value falls in the upper bucket — except the very lowest edge, which is always included. Turn it off for [a, b) intervals, where a boundary value falls in the lower bucket and the very highest edge is included.

Is my data uploaded anywhere?

No. The whole computation runs locally with WebAssembly; your CSV never leaves your browser.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool data-bin "name,score
Alice,12
Bob,55
Carol,88
Dan,73" 'column=score'

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/data-bin/?input=name%2Cscore%0AAlice%2C12%0ABob%2C55%0ACarol%2C88%0ADan%2C73&method=equal_width&column=score&bins=4&edges=0%2C18%2C65%2C120&labels=low%2Cmid%2Chigh&label_style=range&right=true&precision=3&output=append&header=true&delimiter=comma

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.