Stratified Data Splitter
A model evaluation can look better than it is when rows from the same person or household appear in both training and validation. Choose an optional label and group column, enter a seed and proportions, and assign every input row exactly once. Group members stay together; the result shows where indivisible groups prevent ideal balance.
Key features
- The same seed, original row order, columns and percentages reproduce the assignment.
- With a label column and no group column, each label is split separately; rare labels receive explicit warnings.
- With a group column, every row of a group remains in one partition and actual-versus-target deviations are shown.
- Export train, validation and test CSV files, a full input-record assignment CSV and a reproducibility JSON report.
- Strict UTF-8 CSV/TSV parsing preserves quoted line breaks; CSV exports protect spreadsheet-like formulas.
How to use
- Paste a headed CSV or choose a UTF-8 file and inspect the columns.
- Optionally choose a target label column and a subject/group column.
- Enter train, validation and test percentages totaling 100, plus a seed, then run the split.
- Inspect actual partition sizes, label distributions and warnings.
- Download each partition CSV and the assignment and JSON reports to compare with the source.
Use cases
- Keep a customer's transaction rows together by grouping on customer ID.
- Preserve class proportions as far as the selected grouping permits in a small classification experiment.
- See an explicit warning when one large group makes an intended percentage impossible.
- Share a seed and original row-order contract so a teammate can reproduce the same split.
Frequently asked questions
Does it perfectly satisfy both stratification and group isolation?
Group isolation is enforced, while percentages and label balance are heuristic targets. An indivisible group can make exact counts or class coverage impossible. Inspect the shown deviations and missing-label warnings.
Is the same seed alone enough to reproduce a split?
Use the same CSV text and row order, selected columns, percentages and seed. Changing row order or group spelling can change the result. The pseudorandom generator is not cryptographic.
How is this different from the random generator?
The random generator produces values. This tool partitions every supplied CSV row into disjoint datasets, keeps groups together and reports class distributions.
Does the CSV leave my browser?
Input and splitting run in page memory. Downloaded CSV and JSON files can contain the original rows, so review them before sharing elsewhere.
Privacy
The selected CSV and seed are processed in this page's memory and are not automatically uploaded or stored. Downloaded partitions and JSON can contain source rows and selected labels or groups.
Comments & questions