Text Tokenization Lab
Apply browser Intl.Segmenter and two explicit rules to the same text, then see exactly where their source boundaries differ.
Key features
- Compare Intl.Segmenter word, sentence and grapheme output with nonspace and Unicode letter/mark/number runs
- Show each exact source substring with half-open UTF-16 start and end offsets
- Distinguish code point positions from grapheme positions and mark boundaries inside a grapheme
- Count boundaries unique to each of two selected methods and show their surrounding source text
- Try abbreviations, Hangul, combining accents and ZWJ emoji in the sample
- Process up to 12,000 UTF-16 code units entirely in the browser
How to use
- Load the sample or enter multilingual text.
- Choose the locale requested from Intl.Segmenter: Korean, English, Japanese or Thai.
- Run the analysis and compare segment and token counts for five methods.
- Select methods A and B to inspect internal boundary differences and UTF-16/code point positions.
- Inspect the original source spans and download JSON if needed.
Use cases
- Find why string offsets grow after emoji even when they look like one character
- Observe how the current browser splits sentences containing abbreviations
- Compare punctuation behavior for nonspace runs and Unicode letter/mark/number runs
- Inspect character, code point and code unit counts for Hangul and combining accents
Frequently asked questions
Does this reproduce search engine tokenization?
No. It only compares the current browser's Intl.Segmenter with two stated, simple rules. It does not produce search index terms, morphological analysis, part-of-speech labels or semantic analysis of Korean eojeol.
What is a UTF-16 offset?
It is a zero-based JavaScript string code unit position. A supplementary-plane emoji commonly occupies two UTF-16 units even though it is one code point, and a joined emoji cluster can be longer. End offsets are exclusive.
Why do grapheme and code point counts differ?
Combining accents, skin-tone modifiers and joined emoji can contain several code points but appear as one grapheme. A grapheme position is shown only where a boundary exactly matches an Intl.Segmenter grapheme boundary; a boundary inside one is shown as a dash.
What do the explicit rules include?
The nonspace rule captures uninterrupted non-whitespace text, including adjoining punctuation. The Unicode-run rule captures consecutive Letter, Mark or Number characters and skips spaces, punctuation and emoji. Neither infers sentence structure or grammar.
Will every browser give the same result?
Intl.Segmenter locale data and implementations may change by browser or operating system version. The page displays the locale actually resolved at runtime. The two regular-expression rules are fixed in this code.
Is my text sent to a server?
No. Input and segmentation stay in browser memory. A downloaded JSON file includes the exact source text and every segment, so review it before sharing.
Privacy
Source text and segments are processed only in browser memory and are not sent to the server. Downloaded JSON includes the original text and all segments.
Comments & questions