Text Tokenization Lab

Text Tokenization Lab

Compare actual segment boundaries and source string positions across methods.

This is not search engine tokenization or morphological analysis. Intl output depends on the running browser's locale data. The explicit rules only capture nonspace runs or consecutive Unicode Letter, Mark and Number characters.

Text and locale

52 / 12000 UTF-16 code units

Enter text, then run the analysis.

Comments & questions

Text Tokenization Lab

Apply browser Intl.Segmenter and two explicit rules to the same text, then see exactly where their source boundaries differ.

Key features

  • Compare Intl.Segmenter word, sentence and grapheme output with nonspace and Unicode letter/mark/number runs
  • Show each exact source substring with half-open UTF-16 start and end offsets
  • Distinguish code point positions from grapheme positions and mark boundaries inside a grapheme
  • Count boundaries unique to each of two selected methods and show their surrounding source text
  • Try abbreviations, Hangul, combining accents and ZWJ emoji in the sample
  • Process up to 12,000 UTF-16 code units entirely in the browser

How to use

  1. Load the sample or enter multilingual text.
  2. Choose the locale requested from Intl.Segmenter: Korean, English, Japanese or Thai.
  3. Run the analysis and compare segment and token counts for five methods.
  4. Select methods A and B to inspect internal boundary differences and UTF-16/code point positions.
  5. Inspect the original source spans and download JSON if needed.

Use cases

  • Find why string offsets grow after emoji even when they look like one character
  • Observe how the current browser splits sentences containing abbreviations
  • Compare punctuation behavior for nonspace runs and Unicode letter/mark/number runs
  • Inspect character, code point and code unit counts for Hangul and combining accents

Frequently asked questions

Does this reproduce search engine tokenization?

No. It only compares the current browser's Intl.Segmenter with two stated, simple rules. It does not produce search index terms, morphological analysis, part-of-speech labels or semantic analysis of Korean eojeol.

What is a UTF-16 offset?

It is a zero-based JavaScript string code unit position. A supplementary-plane emoji commonly occupies two UTF-16 units even though it is one code point, and a joined emoji cluster can be longer. End offsets are exclusive.

Why do grapheme and code point counts differ?

Combining accents, skin-tone modifiers and joined emoji can contain several code points but appear as one grapheme. A grapheme position is shown only where a boundary exactly matches an Intl.Segmenter grapheme boundary; a boundary inside one is shown as a dash.

What do the explicit rules include?

The nonspace rule captures uninterrupted non-whitespace text, including adjoining punctuation. The Unicode-run rule captures consecutive Letter, Mark or Number characters and skips spaces, punctuation and emoji. Neither infers sentence structure or grammar.

Will every browser give the same result?

Intl.Segmenter locale data and implementations may change by browser or operating system version. The page displays the locale actually resolved at runtime. The two regular-expression rules are fixed in this code.

Is my text sent to a server?

No. Input and segmentation stay in browser memory. A downloaded JSON file includes the exact source text and every segment, so review it before sharing.

Privacy

Source text and segments are processed only in browser memory and are not sent to the server. Downloaded JSON includes the original text and all segments.

Related Tools

Sentence Structure LabCorpus ConcordancerUnicode InspectorDOCX Style AuditorCitation ManagerDocument Search IndexSigned PDF InspectorPDF/A Archival PreflightPrint Preflight AuditorSearchable Scan PDF MakerPublication Accessibility AuditorResearch Evidence MatrixPDF Form DesignerPDF Annotation StudioPDF Object InspectorPDF Outline EditorEPUB Authoring WorkbenchDocument Version ComparatorBooklet Imposition DesignerSensitive Text RedactorHanja Reading WorkbenchGrammar Production LabDialogue Script EditorBook Index BuilderBilingual QA CheckerText Diagram EditorANSI Art StudioEmail Thread ExplorerPDF Redaction StudioRich Text SanitizerFixed Width Record DesignerPresentation Rehearsal StudioThree Way Text MergeKorean & English Braille ConverterText to SpeechTyping speed testKorean Text Pattern ReviewKorean Spelling QuizFont Preview and ComparisonHTML to MarkdownMojibake RepairSort LinesFind & ReplaceReverse TextRoman Numeral ConverterHangul Jamo ConverterKorean Initial ConsonantsAdd or Remove Line NumbersNumber to Korean WordsMorse Code ConverterROT13 & Caesar ConverterURL Slug GeneratorHTML Tag RemoverEmoji CollectionCharacter CounterUnit ConverterFile Size ConverterColor Code ConverterText DiffPDF Merge/SplitBusiness Day CalculatoriCalendar Rule LabvCard Address Book EditorJPG to PDFKorean Name RomanizerOrganize PDF PagesFancy Font GeneratorContrast CheckerColor Palette GeneratorPDF to JPGLorem Ipsum GeneratorText Template WorkbenchKorean-English Typo ConverterText CleanerOnline NotepadWord Cloud GeneratorPDF UnlockEnglish Address ConverterPDF Compressor
Explore all Text/Convert tools →Image/Media →Life/Fun →Dev Tools →