Tabular Cleaning Workbench
Table cleaning is a sequence of decisions. This workbench keeps the original records separate from the result and replays your ordered steps whenever you run them. You can remove any step or undo the last addition, then compare the new output with the original and inspect every changed cell or excluded row.
Key features
- Preserved original CSV/JSON records with stable source references
- Ordered trim, number, boolean, date, text and missing-value steps
- Explicit keep, null or row-exclusion policy for conversion failures
- Composite-key duplicates with keep-first or keep-last selection
- Remove any step or undo the last addition, then replay the original
- Paginated original/result/history/excluded tables and three CSV exports
How to use
- Paste a table or choose a UTF-8 CSV/TSV/JSON file and read its columns.
- Choose an operation, target column and relevant failure or duplicate policy.
- Add steps in the order they should run; fill uses text so add numeric conversion after it when needed.
- Run from the original and review the per-step counts and cell history.
- Remove or undo steps and rerun if necessary, then download the cleaned CSV, changes and excluded originals.
Use cases
- Normalize padded region names before deduplicating customer records.
- Convert grouped invoice amounts while retaining unrecognized text for review.
- Exclude missing measurement rows while preserving an export of their original values.
- Reproduce and review a small table cleanup without overwriting its input.
Frequently asked questions
Does undo restore the original values?
Undo removes the last pipeline step. Removing any step invalidates the old result. Run again to replay the remaining steps from the preserved original, including its original strings and JSON types. No prior output is used as a new source unless you explicitly import it.
How are failed type conversions handled?
Choose keep original, replace with null or exclude the whole row. Each stage counts failed cells separately from changed cells and removed rows. Numbers use finite decimal/exponent syntax; an explicit option accepts English grouping commas. Leading-zero strings become numbers only when you add numeric conversion.
What counts as missing, and how does fill work?
Null, absent JSON keys and whitespace-only strings count as missing; literal NA and null strings do not. Fill inserts the exact text you enter. To obtain a numeric filled value, put a number-conversion step after the fill step. Nullable values stay null during type conversion.
How are duplicate keys compared?
The selected columns are compared as an exact tuple after earlier steps. Case, whitespace and scalar types matter. Null, empty string and a space are distinct until another step normalizes them. Keep-first or keep-last chooses a source occurrence and preserves the surviving rows’ order.
What is stored in the change and exclusion files?
The change CSV contains every applied cell change with step, source record, column, original/updated type and value. The exclusion CSV contains step, reason and the unmodified original row, even if earlier steps changed it. Numbered original-column prefixes avoid collisions with metadata names.
What are the limits and CSV caveats?
The input is limited to 1 MiB, 5,000 records, 32 columns and 100,000 cells. A plan has at most 12 steps and 10,000 changed-cell records; exceeding a limit fails without keeping a partial result. CSV cannot preserve JSON scalar types or distinguish null from an empty cell. Formula-like text receives a leading apostrophe; numeric values remain numeric.
Privacy
Tables, settings and intermediate values stay in the current page’s browser memory. This tool does not send table contents to a server, URL or analytics event, and does not autosave them to browser storage. Only explicit downloads create files. Editing the input discards its parsed table, pipeline and results.
Comments & questions