Rule-based text extractor

Write a list of named regex rules and pull structured fields out of any text — log lines, invoices, emails, scraped pages. Grok-style %{PATTERN:field} shortcuts included, output as JSON, CSV or a readable listing. Runs entirely in your browser; nothing is uploaded.

Try:
Extracted fields

Extract structured fields with rules

Rule-based extraction is for the messy middle ground between a one-off regex and a full parser. Paste text, write one rule per field, and get back structured JSON, CSV, a readable listing, or a rule report. It is useful for logs, invoice snippets, support emails, scraped pages, clipboard dumps, and any text where the same date, code, email, IP address, ticket ID, or amount appears in a predictable shape.

Each rule is one line. The most common form is field = regex, where the field name becomes the output key and the first capture group (or the whole match if there is no group) becomes the value. You can also write a bare regex with named groups, or use Grok-style shortcuts such as %{DATE_ISO:date} and %{IPV4:client} to keep patterns short. Define your own reusable macro with @NAME = regex, then use %{NAME} or %{NAME:field} later in the rule block.

Worked example

Input text:

Invoice INV-2026-014 dated 2026-08-15, billed to [email protected], total $1,240.00 (VAT 21%).

Rules:

invoice = INV-\d{4}-\d+
date = %{DATE_ISO}
email = %{EMAIL}
total = %{MONEY}
vat = %{PERCENT}

With Output set to JSON, the extractor returns:

{"invoice":"INV-2026-014","date":"2026-08-15","email":"[email protected]","total":"$1,240.00","vat":"21%"}

For logs, switch Split input into records by to Lines so each line becomes one output object. For every email address or every URL in a record, switch Matches per rule to Every match; the JSON value becomes an array.

Rule syntax

Built-ins include WORD, INT, NUMBER, EMAIL, URL, IPV4, IPV6, MAC, UUID, HASH, DATE_ISO, DATE_US, TIME, TIMESTAMP_ISO, SYSLOG_TIME, LOGLEVEL, HTTP_METHOD, HTTP_STATUS, HOSTNAME, PATH, MONEY, PERCENT, PHONE, SEMVER, and TICKET.

Options

Limits and edge cases

FAQ

How is this different from a normal regex tester?

A regex tester tells you whether a pattern matches. This tool turns a whole set of named patterns into structured output. You can split text into records, extract several fields per record, choose JSON or CSV, and use the report view to see which rules never matched.

Can I use one rule block for many log lines?

Yes. Set Split input into records by to Lines. Each line is processed as one record and the JSON output becomes an array of objects. Turn on Skip records with no matches to ignore headers, blank lines, or unrelated lines.

What does `%{DATE_ISO:date}` mean?

DATE_ISO is a built-in pattern for dates such as 2026-08-15, and :date names the capture group. %{DATE_ISO:date} is equivalent to writing a named regex group by hand, but it is shorter and less error-prone.

Why did my lookbehind or backreference fail?

The tool uses Rust's regex engine, which deliberately excludes look-around and backreferences so matching remains predictable and fast in WebAssembly. Rewrite the pattern with explicit captures, split the input into smaller records, or use a separate parser if you need context-sensitive matching.

How do I debug a rule that is not matching?

Set Output to Rule report. The report lists every field, the rule line that produced it, how many records it hit, and which fields never matched. That lets you adjust one pattern without guessing from an empty JSON result.

Developer & Automation Access

Run it from the terminal

Same engine as this page, headless — via the gizza CLI:

gizza tool rule-based-extractor "Paste the log lines, invoice text or records to extract from…" 'rules=date = %{DATE_ISO}
level = %{LOGLEVEL}
client = %{IPV4}'

New to the CLI? Get gizza →

Open it by URL

Pre-fill and auto-run this tool with query parameters — the names match the API/CLI:

https://gizza.ai/tools/rule-based-extractor/?text=Paste%20the%20log%20lines%2C%20invoice%20text%20or%20records%20to%20extract%20from%E2%80%A6&rules=date%20%3D%20%25%7BDATE_ISO%7D%0Alevel%20%3D%20%25%7BLOGLEVEL%7D%0Aclient%20%3D%20%25%7BIPV4%7D&split=whole&split_pattern=%5E-%7B3%2C%7D%24&matches=first&ignore_case=true&multiline=true&dotall=true&trim=true&unique=true&on_missing=skip&skip_empty_records=true&max_records=5000&max_matches=1000&output=json&pretty=true

Machine-readable descriptor: tool.json — title + parameters JSON Schema for agents.