Scientific Data Profiler
A table can be syntactically valid while hiding missing measurements, inconsistent types or a column that never changes. This profiler scans each column and connects a visual distribution to record-level evidence. It does not modify the source or automatically classify a valid observation as an error.
Key features
- CSV/TSV and flat JSON scalar records with explicit parsing errors
- Column type counts, missingness, typed distinct values and constants
- Up to ten numeric histogram bins or eight leading category frequencies
- Linear quartiles, mean, median and transparent 1.5×IQR fences
- Suspect source records and complete JSON report download
- Cancellation and bounded local processing with safe CSV export
How to use
- Paste CSV/JSON records or select a UTF-8 file.
- Choose the format and CSV delimiter; enable grouping commas only if they mean thousands.
- Run the profile and inspect the column overview.
- Select a column to inspect its distribution, numeric statistics and suspect records.
- Export the JSON report or suspect-record CSV and review flagged values against the original source.
Use cases
- Check a sensor export for missing measurements and isolated extreme values.
- Find a constant field or mixed date and text column before analysis.
- Compare the actual value frequencies in a survey export.
- Retain source record references while reviewing a spreadsheet import.
Frequently asked questions
What counts as missing?
JSON null, a missing JSON key, an empty CSV cell and whitespace-only strings count as missing. The literal strings null and NA remain text. A fully missing column is empty, not constant; one distinct nonmissing typed value makes a column constant.
How are numbers, dates and identifiers inferred?
Finite decimal and exponent forms within ±9,007,199,254,740,991 are numeric. Leading-zero strings such as 001 stay text. Group commas require the explicit option and strict 1,234.5 syntax. True/false are boolean; valid YYYY-MM-DD strings are dates. Other values remain text, and columns with multiple kinds are mixed.
How are suspected extreme values selected?
For at least four numeric values in a column, sorted positions (n−1)×q define linearly interpolated quartiles. Values below Q1−1.5×IQR or above Q3+1.5×IQR are listed. Mixed columns use their numeric subset. This is a review rule, not proof of a data error; no row is removed.
What does the chart show?
A numeric subset is divided into up to ten equal-width bins, with the final upper boundary included. A constant has one bin. Without numeric values, the eight most frequent typed values and a combined other count are shown. Missing values are counted separately.
Can JSON structures or very large integers be loaded?
JSON input must be an array of flat objects containing scalar values. Duplicate keys and nested objects/arrays are rejected. Numeric tokens that overflow, underflow to zero or exceed the safe magnitude are rejected before conversion; quote exact identifiers as strings. The supported numeric calculations use JavaScript floating-point arithmetic.
What is included in a download?
The JSON report includes every column, rules, statistics, histogram bins and suspect values. The CSV contains suspect evidence with source-record numbers and fences. CSV record numbers count logical records including the header, while JSON numbers start at the first object. Formula-like text is prefixed with an apostrophe in CSV exports.
Privacy
Input and results stay in this page’s browser memory. This tool does not upload table contents, put them in URLs or analytics events, or automatically persist them in browser storage. Downloads are created only when you request them. Editing input clears the previous result.
Comments & questions