Target-encode a CSV categorical column

Paste CSV, pick a category and numeric target, and turn the category into its per-category target mean — with optional smoothing and leave-one-out.

Try:
Encoded CSV

About this tool

Target Mean Encoder replaces a categorical CSV column with the mean of a numeric target for each category — also called mean, impact, or likelihood encoding. It turns a high-cardinality text column (city, product, merchant, ZIP) into a single model-ready numeric feature without one-hot exploding your column count.

Pick the category column and the numeric target column, and each category becomes the average target value for its rows. Optional m-estimate smoothing blends rare categories toward the overall mean, and leave-one-out excludes each row's own target so an in-place encoding leaks less signal into the model.

Worked example

CSV:

city,churn
A,1
B,1
A,0
B,1
C,1
B,0

With category city, target churn, replace output, and 3 decimals, each city becomes its churn rate — A = 1/2 = 0.500, B = 2/3 = 0.667, C = 1/1 = 1.000:

city,churn
0.500,1
0.667,1
0.500,0
0.667,1
1.000,1
0.667,0

Switch output to append to keep the original column and add city_target_enc instead of overwriting city.

Limits & edge cases

FAQ

What is target (mean) encoding and when should I use it?

Target encoding replaces each category with the average of a numeric target for that category's rows. It is handy for high-cardinality categorical features where one-hot encoding would add too many columns, giving models a compact numeric signal instead.

How does smoothing help with rare categories?

A category seen only once or twice has a noisy mean. Smoothing (the m-estimate) pulls those means toward the global target mean: the higher the smoothing weight m, the more a small category is shrunk toward the overall average, which reduces overfitting to rare values.

What does leave-one-out do?

Leave-one-out excludes each row's own target value from its category statistics before encoding that row. This reduces target leakage when you encode a training set in place, because a row can no longer "see" its own label through the encoded value.

What happens to blank or unseen categories?

Rows whose category cell is blank — or whose category has no usable target statistics — take the value from the Unknown / blank category setting. You can use the global mean (a safe default), NaN (to flag them for later handling), or zero.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool target-mean-encoder "city,churn
A,1
B,1
A,0
B,1
C,1
B,0" 'category=city' 'target=churn'

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/target-mean-encoder/?data=city%2Cchurn%0AA%2C1%0AB%2C1%0AA%2C0%0AB%2C1%0AC%2C1%0AB%2C0&category=city&target=churn&smoothing=0&leave_one_out=true&output=replace&unknown=global-mean&decimals=4&has_header=true&delimiter=comma

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.