PCA and t-SNE visualizer

Paste a labelled CSV or whitespace table and project high-dimensional rows into a 2D scatter plot. Use PCA for variance-explained axes or deterministic t-SNE for local clusters; points are coloured by a label column and the output stays browser-local.

Try:
Projection

About this tool

High-dimensional tables are hard to inspect directly: six measurements per sample already means no ordinary x/y chart can show the whole row. This tool turns each row into one point in a two-dimensional scatter plot so clusters, outliers, and label separation become visible.

Use PCA when you want a linear projection with interpretable axes. The x and y captions show how much variance PC1 and PC2 explain, and the calculation is the same deterministic Jacobi PCA engine used by the sibling PCA calculator. Use t-SNE when the goal is a visual cluster map: it preserves local neighbours rather than global distances. This implementation is deterministic too — it starts from PCA scores instead of a random layout — so rerunning the same table produces byte-identical output.

One non-numeric column can carry a class, species, cohort, cluster, or other label. The tool drops that column from the math, uses it to colour the points, and draws a legend. Leave Label column empty to auto-detect the only text column, or name it explicitly if the table has several text columns or a numeric group code.

Worked example

Paste this small labelled table:

sepal_len,sepal_wid,petal_len,petal_wid,species
5.1,3.5,1.4,0.2,setosa
4.9,3.0,1.4,0.2,setosa
5.8,2.7,4.1,1.0,versicolor
6.4,3.2,4.5,1.5,versicolor
6.5,3.0,5.8,2.2,virginica
7.6,3.0,6.6,2.1,virginica

Keep Projection method as PCA, Label column as species, and the output as SVG. The result is a standalone scatter plot: every row is a coloured circle, the legend lists the three species, and the axes are labelled with the variance share explained by PC1 and PC2. Switching Output to CSV returns rows like index,label,pc1,pc2, ready to feed into another charting tool.

For a non-linear cluster map, switch Projection method to t-SNE. Start with perplexity around 3–10 for a tiny table, 30 for a few hundred rows, and increase Iterations if the layout still looks compressed.

Limits and edge cases

FAQ

When should I use PCA instead of t-SNE?

Use PCA first when you want a fast, stable overview and axes that mean something: PC1 and PC2 are linear combinations of your original variables, and the captions report explained variance. Use t-SNE when you mostly care about whether nearby points form clusters. t-SNE can make local groups clearer, but its axes are not interpretable measurements.

How does the tool choose the label column?

If Label column is empty, the tool looks for a single non-numeric column and uses it as the label. That works for tables like x,y,z,class. If there are multiple text columns, or if the group column is numeric, set the label column by header name (species) or by 1-based index (5). The selected column is excluded from the projection and used only for colours and legend text.

Why does t-SNE look different from PCA?

PCA preserves the broad linear variance structure, so points that are globally far apart in the original variables tend to stay far apart. t-SNE optimizes neighbourhoods: points with similar neighbours are pulled close, and unrelated clusters are pushed apart for readability. That makes it great for cluster maps, but the exact distance between two separate clusters is not a reliable quantity.

Should I standardize the columns?

Usually yes. Without standardization, a column measured in thousands can dominate a column measured between 0 and 1, even if both are equally important. Keep Standardize numeric columns enabled for mixed units such as height/weight/age, gene counts, survey scales, or financial ratios. Turn it off only when raw variance is meaningful.

Can I use the result outside this page?

Yes. The default SVG is standalone and can be saved directly as a vector chart. CSV output gives index,label,pc1,pc2 or index,label,tsne1,tsne2 coordinates for another plotting tool. JSON output includes the full projection, categories, variable names, PCA explained variance, and the t-SNE perplexity actually used after small-table clamping.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool pca-visualizer "sepal_len,sepal_wid,petal_len,petal_wid,species
5.1,3.5,1.4,0.2,setosa
4.9,3.0,1.4,0.2,setosa
6.4,3.2,4.5,1.5,versicolor
6.9,3.1,4.9,1.5,versicolor
6.5,3.0,5.8,2.2,virginica
7.6,3.0,6.6,2.1,virginica"

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/pca-visualizer/?data=sepal_len%2Csepal_wid%2Cpetal_len%2Cpetal_wid%2Cspecies%0A5.1%2C3.5%2C1.4%2C0.2%2Csetosa%0A4.9%2C3.0%2C1.4%2C0.2%2Csetosa%0A6.4%2C3.2%2C4.5%2C1.5%2Cversicolor%0A6.9%2C3.1%2C4.9%2C1.5%2Cversicolor%0A6.5%2C3.0%2C5.8%2C2.2%2Cvirginica%0A7.6%2C3.0%2C6.6%2C2.1%2Cvirginica&method=pca&label_column=species&scale=true&perplexity=30&iterations=500&learning_rate=200&show_labels=true&point_size=4&title=Iris%20measurements%20%E2%80%94%20PCA&width=720&height=520&format=svg

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.