Identify a data sample's format, delimiter and columns

Paste rows of data and get the detected format with a confidence score, the delimiter and quote character, line endings, encoding, header row, column count and per-column types.

Try:
Detection report

About this tool

Data Format Sniffer inspects a pasted sample and reports the most likely structure before you hand it to a parser or converter. It recognises CSV, TSV, semicolon/pipe/custom-delimited text, JSON, JSON Lines, XML, HTML, Markdown tables, fixed-width text, marker-led YAML, and common binary containers such as Parquet and Avro when their bytes are supplied as base64 or hex.

The report includes a confidence score, encoding note, line-ending style, delimiter and quote character, a per-candidate delimiter score table, header-row guess, column count, inferred column types, a row preview, and warnings for ragged rows or sampled-only analysis. Use output=json when another workflow needs the same facts as machine-readable fields.

Worked example

Input:

name,age,city,joined
Ada,36,London,1815-12-10
Alan,41,Wilmslow,1912-06-23
Grace,45,New York,1906-12-09

With the default settings the tool reports CSV with a comma delimiter, UTF-8 text input, LF line endings, four columns, a likely header row, and column types such as integer and date. If your data uses a less common delimiter, put it in extra_delimiters, for example ^ for caret-delimited records.

Limits and edge cases

FAQ

Can it detect Parquet or Avro from a real file?

Yes, if you provide the beginning of the file as bytes using input_form=base64 or input_form=hex. The tool checks magic bytes such as PAR1 for Parquet and Obj\x01 for Avro. The page itself does not read uploaded files, so paste a byte sample instead of a filename.

Why does pasted text always say UTF-8?

A browser text field contains Unicode text, not the original file bytes. By the time the tool receives input_form=text, the original encoding has already been decoded. Use base64 or hex input when you need BOM or statistical encoding detection over real bytes.

How reliable is delimiter detection?

The sniffer tries comma, tab, semicolon, pipe, colon, tilde, space, and any extra delimiters you provide. It prefers candidates that produce at least two columns with consistent row widths, then reports every candidate's column count and consistency so you can spot ambiguous samples.

Does it validate every row in the file?

No. It samples leading lines for speed and reports early ragged rows when the winning delimiter gives inconsistent column counts. For a complete validation pass with row-by-row errors, use a validation-specific tool after the format has been identified.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool data-format-sniffer "name,age,city,joined
Ada,36,London,1815-12-10
Alan,41,Wilmslow,1912-06-23"

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/data-format-sniffer/?data=name%2Cage%2Ccity%2Cjoined%0AAda%2C36%2CLondon%2C1815-12-10%0AAlan%2C41%2CWilmslow%2C1912-06-23&input_form=text&sample_lines=100&extra_delimiters=%5E%23&comment_prefix=%23&detect_types=true&preview_rows=5&output=report

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.