Train a local Naive Bayes text classifier

Paste labeled examples, classify new text, and see the predicted label, class probabilities, model settings, training statistics, and the tokens that drove the decision.

Try:
Classification result

About this tool

Naive Bayes Text Classifier trains a small supervised text model from the examples you paste, then immediately classifies new text in the browser or CLI. Each training line is label<separator>text, such as spam,win a free prize now. The separator can be auto-detected from tab, comma, pipe, or colon, or forced when your examples need it.

The default multinomial model is the classic bag-of-words baseline for spam, topic, ticket, and short-message classification. Bernoulli mode uses token presence/absence, and complement mode is useful when one label has many more examples than another. Controls cover smoothing alpha, n-grams, case folding, English stop-word removal, rare-token pruning, class priors, class-list length, explanations, and report vs JSON output.

Worked example

Training data:

spam,win a free prize now
spam,free money click here
spam,claim your free gift
ham,meeting at ten tomorrow
ham,lunch with the team
ham,project update attached

Text to classify:

claim your free money now

With default settings, the model predicts spam, lists the class probabilities, and shows the tokens that pushed the decision toward spam over the runner-up. Switch input_mode=lines to classify each non-blank input line as a separate document.

Limits and edge cases

FAQ

How should I format the training data?

Use one labeled example per line: label,text. The first separator on the line splits the label from the example text, so the text may contain more commas after that. Tabs, pipes, and colons are also supported; choose the separator explicitly if auto-detection guesses wrong.

Which model variant should I choose?

Start with multinomial for ordinary word-count text classification. Use bernoulli for very short texts where word presence matters more than repeated words. Try complement when the labels are imbalanced and the largest class otherwise dominates too easily.

What does alpha smoothing do?

alpha adds a small count to every token/class pair so a word that was unseen for one label does not make that label impossible. The default 1 is Laplace smoothing. Lower values make the classifier more confident in the pasted examples; higher values flatten the probabilities.

Can I reuse or export the trained model?

No. The tool trains from scratch on every run and returns only the prediction/report. That keeps the tool stateless and local. If you need persisted models, cross-validation, or deployment, use a dedicated machine-learning workflow after validating the idea here.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool naive-bayes-text-classifier "spam,win a free prize now
spam,free money click here
ham,meeting at ten tomorrow
ham,lunch with the team" 'text=claim your free money now'

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/naive-bayes-text-classifier/?training_data=spam%2Cwin%20a%20free%20prize%20now%0Aspam%2Cfree%20money%20click%20here%0Aham%2Cmeeting%20at%20ten%20tomorrow%0Aham%2Clunch%20with%20the%20team&text=claim%20your%20free%20money%20now&separator=auto&input_mode=single&model=multinomial&alpha=1&ngram_max=1&lowercase=true&remove_stopwords=true&min_count=1&priors=empirical&top_k=3&explain=true&output=report

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.