Translation Memory Workbench
Keep sentence pairs you translated or reviewed yourself in one browser tab. Align equal-count lines manually, edit and reorder pairs, split or merge segments, and search a new source against saved sources. Similarity scores are deterministic editing aids; choosing a candidate only fills an unsaved draft. The page never translates or confirms a translation automatically.
Key features
- Manual one-to-one alignment of equal-length source and target line lists
- Stable pair IDs, editing, ordering, splitting and adjacent merging
- Matching numbered placeholders with source and target code payloads preserved
- Deterministic Unicode NFC, lowercase, whitespace-normalized character-bigram fuzzy retrieval
- Strict versioned JSON and a documented plain-text-plus-ph TMX 1.4 profile
- Formula-safe CSV report and explicit unsupported-TMX errors
How to use
- Load the example or start a project with distinct source and target language tags.
- Paste one sentence per line on each side, in the order you want paired. Both lists must have the same number of nonempty lines.
- Select a pair to edit, reorder, split at both character offsets, or merge with the next pair. Add matching numbered placeholders in both texts when a code must be retained.
- Enter a new source in Search. Inspect ranked saved pairs and copy a target into the draft only if it is useful; save the draft explicitly after your review.
- Download JSON to continue editing, TMX 1.4 for the supported exchange profile, or CSV for review. Import supported JSON or TMX only after checking its format.
Use cases
- Reuse a reviewed button label when translating a related interface sentence
- Align a short bilingual instruction set before handing it to an editor
- Preserve inline placeholder codes while splitting a long pair
- Inspect likely matches without sending text to a translation service
Frequently asked questions
Does a similarity match produce a translation?
No. The tool compares your new source with saved sources and displays the saved target. Copying it fills a draft only. You must inspect context, placeholders and meaning, then save the pair yourself.
How is the fuzzy score calculated?
Both sources are normalized to Unicode NFC, lowercased and whitespace-collapsed. Exact normalized matches score 100. Otherwise character bigrams use multiset Dice overlap, rounded to two decimals. Placeholders are excluded. Punctuation remains, and scores do not measure translation quality.
What TMX files can I import?
This page uses a deliberately narrow TMX 1.4 profile: one source and one target TUV per TU, sentence segmentation, text plus numbered ph inline codes, and no extra metadata or other inline code elements. Unsupported structures produce an error rather than silently losing content. Exported TMX can be imported back.
How do numbered placeholders work?
Write a marker such as ⟦1⟧ once in each side of a pair and add a placeholder row with ID 1 and the source/target code payloads. Both sides must contain the same IDs exactly once. TMX export writes matching ph elements; changing a marker without changing its row is rejected.
What happens when pairs are split or merged?
Split takes separate UTF-16 offsets on the source and target, preserving exact text on either side. Each placeholder must remain on the corresponding side of the new pair. Merge inserts one space between adjacent segments and renumbers collisions in the second pair's placeholder IDs while preserving their payloads.
Will this preserve every translation-memory file?
No. Only the supported JSON v1 and restricted TMX 1.4 profile round-trip. In particular bpt/ept/it/hi markup, properties, notes, extra variants and custom metadata are rejected. Keep an original external TMX separately.
Privacy
All sentence pairs and searches stay in this browser tab. There is no account, remote translation, collection or automatic save. Download JSON before leaving if you need to resume.
Comments & questions