Basket Pattern Analyzer
Count which items appear together in the transaction baskets you supply. Repeated occurrences of an item within the same transaction count once. The analyzer finds complete frequent itemsets within your chosen size limit and generates directed rules with their actual counts and descriptive measures. These observations are not causal conclusions or predictions of purchases, recommendations or revenue.
Key features
- Explicit transaction/item CSV schema and duplicate membership counts
- Exact support counts for itemsets of one to four items
- Directed rules with confidence, lift, leverage and conviction
- Separate finite, infinite and undefined conviction results
- Bounded, cancellable work with no silent partial success
- Top co-occurrence graph and complete independent CSV/JSON exports
How to use
- Paste a CSV or select a UTF-8 file with transaction and item columns.
- Choose the delimiter, whitespace policy, minimum support/count and confidence.
- Set maximum itemset length and the hard rule limit, then run analysis.
- Inspect the denominator, duplicate count, itemsets, directed rules and bounded graph.
- Download all itemsets, all rules or the complete JSON; preview filters do not restrict exports.
Use cases
- Explore anonymous items commonly present in the same sample basket.
- Compare conditional frequencies in A→B and B→A directions.
- Check a small labeled dataset against hand-counted association measures.
Frequently asked questions
What CSV shape is required?
Use exactly two columns named transaction and item; either order is accepted. Each row is one item membership in an anonymous transaction ID. Quoted commas and escaped quotes are supported, and you choose comma, semicolon or tab. Invalid quotes, ragged rows, empty labels and extra columns are errors. Physically blank records are ignored. Empty baskets cannot be represented; the denominator contains only transactions with at least one valid item.
How are duplicates and names treated?
The same item appearing repeatedly in one transaction is counted once and reported as a duplicate membership. Optional trimming removes surrounding whitespace before grouping. Case, Unicode spelling and leading zeroes remain distinct; no number conversion or Unicode normalization is applied. IDs and items are at most 120 UTF-16 units and cannot contain embedded controls, bidirectional control characters or unpaired surrogates.
Which itemsets and rules are included?
An itemset must meet both the percentage support threshold and minimum transaction count. Percentages allow two decimal places and boundaries use integer arithmetic. Within lengths 1–4, every observed qualifying itemset is retained. Each nonempty proper partition generates a directed rule if its confidence meets the threshold. Maximum rules is a hard limit, not a top-N truncation. Excessive input or work rejects the whole result.
What do the measures mean, and why can conviction be undefined?
Support is the fraction of transactions containing both sides. Confidence conditions that count on the antecedent. Lift divides confidence by consequent support; leverage subtracts the independent-support product. Conviction divides one minus consequent support by one minus confidence. Positive/zero is marked infinite, while zero/zero is undefined. JSON uses null plus an explicit state for both; these are not missing observations.
What does the graph show?
It is an undirected co-occurrence view, not a recommendation or causality graph. Up to 12 frequent single items with the largest counts become numbered nodes. Up to 20 frequent pairs between those nodes are shown, ranked by count; width reflects pair support. Table filters and rule-confidence thresholds do not change it. With maximum length 1, no pairs are calculated. The tables and exports cover the complete configured scope.
What limits and export guarantees apply?
Input is at most 1 MiB and 5,000 data rows; there may be 2,000 nonempty transactions, 200 unique items and 20 items per basket. Generation stops at 20,000 distinct combinations, 10,000 frequent sets, the chosen rule cap (at most 5,000) or one million work units. Cancel and limit errors discard the whole report. All output rows are exported regardless of filtering or pagination. CSV uses UTF-8 BOM, quoted fields and CRLF; sets are JSON arrays. JSON retains full numeric precision and omits transaction IDs.
Privacy
Data stays in page memory and is not uploaded or saved automatically by this tool. Use anonymous transaction IDs. Selected file bytes are read only after selection; analysis and download occur on request. Reports contain item labels and aggregate counts but omit transaction IDs and raw CSV. Leaving the page removes unsaved work.
Comments & questions