Decision Tree Lab
Learn how a classification tree divides a small labeled table into branches. This lab fits its own greedy binary Gini model using only the training rows and reports the remaining rows separately. You can inspect every node, compare training and held-out results and follow the route for a new row. The displayed scores describe this one sample and split; they are not a promise about future accuracy.
Key features
- Explicit numeric and exact-string categorical feature choices; the target column is excluded from features.
- Reproducible seeded row split before fitting, with training-only candidate thresholds, categories and class counts.
- A binary tree with a depth limit, minimum child size, full branch table and scrollable SVG.
- Separate training and held-out accuracy and confusion matrices, plus unseen-label and duplicate-feature warnings.
- New-row prediction with the actual yes/no path; complete CSV, JSON and SVG downloads.
- Strict CSV errors, cancellation and a fixed operation budget, with no partial successful report.
How to use
- Paste CSV or choose a UTF-8 file, select the delimiter and inspect the columns.
- Choose the target column, select up to eight features and explicitly set each feature to numeric or categorical.
- Set the seed, held-out ratio, maximum depth and minimum child size.
- Train the tree and compare the two scores, confusion matrices, full node table and sample-limit notices.
- Enter a new row to follow its prediction path, or export the complete results.
Use cases
- Teach how a numeric threshold or a categorical equality separates a small labeled sample.
- Compare a shallow and a deeper tree while watching the training versus held-out accuracy gap.
- Inspect wrong held-out predictions and trace the actual branch decisions behind them.
- Reproduce a classroom example using the same CSV row order, options and seed.
Frequently asked questions
Which CSV data can I use?
Use 5–2,000 data rows and 2–32 columns, at most 1 MiB of UTF-8. Choose one categorical target and one to eight other features. Numeric and categorical feature types are explicit. Blank selected cells, invalid quotes, irregular row widths and duplicate headers are errors; unused empty cells are allowed. Blank records are skipped. The target may contain up to 12 distinct labels across all input rows.
Is the held-out data used to build the tree?
The seeded shuffle splits row indices first. Only training rows supply class counts, numeric candidate boundaries and category vocabulary. Reading all rows to validate limits or compute the final confusion matrices is distinct from fitting. A new held-out class cannot become a predicted leaf class. A categorical feature may have at most 32 distinct training values; unseen held-out values are reported separately.
What exact algorithm does the lab implement?
It uses its own greedy binary Gini profile, not an imported training library. Numeric splits are x ≤ an observed training value; categorical splits test equality to one category versus all others. It considers all valid candidates in that scope and picks the lowest weighted child impurity with a strictly positive gain. Integer count ratios are compared exactly. Equal gains keep selected feature order and then ascending numeric or code-unit category order; tied leaf classes use code-unit order. This is not a global-optimal-tree guarantee or an exact clone of another library.
How should I read the two scores and the gap?
Accuracy is correct predictions divided by rows in that partition. A confusion-matrix row is the actual label and a column is the predicted label. The displayed gap is training accuracy minus held-out accuracy in percentage points. A positive gap can prompt an overfitting review, but one small random split is uncertain. Repeatedly tuning settings on the same held-out rows weakens their value as an independent check.
Does the seed prevent every kind of data leakage?
No. It reproduces a simple row shuffle for identical row order and settings. It does not stratify classes, isolate people or groups, preserve time order, or detect all leaked columns. Held-out rows with feature vectors also found in training are counted as a warning, not removed. Correlated rows and target-derived input columns can still inflate scores. A dedicated grouped or time-aware split needs separate preparation.
What do the downloads and tree picture include?
CSV contains every input row’s selected feature values, partition, actual/predicted label and node path, independent of screen filtering or pagination. JSON includes the same full model, options, warnings and both scores. Numeric cells use parsed binary64 values rather than original numeric spelling. SVG contains all nodes, at most 127, with compact labels and full title text; the on-page branch table keeps complete conditions. Files are created only when you click download.
Privacy
The tool reads only pasted text or a file you explicitly select and works in page memory. It does not upload input, write it to analytics or URLs, or save it automatically. Downloads contain selected feature values and target labels; review them before sharing. Reloading or leaving the page discards unsaved data.
Comments & questions