Mojibake Repair
Repair cases such as café becoming café, where UTF-8 bytes were read using a Western single-byte encoding. You choose the mistaken encoding and number of repair passes. This is a limited, reversible conversion rather than a universal encoding detector.
Key features
- Recover bytes from Latin-1 or Windows-1252 characters, then decode strict UTF-8.
- Verify that re-encoding the result reproduces exactly the same UTF-8 bytes.
- Choose one to three explicit passes for repeated mis-encoding.
How to use
- Paste the affected text. If correct non-Latin text is mixed in, isolate the broken section first.
- Select the mistakenly applied encoding. Windows-1252 with one pass is a useful first case to check.
- Run and review the result. If rejected, check the encoding, pass count, or the original file’s actual encoding.
Use cases
- Repair Western text such as café
- Recover UTF-8 Korean text displayed as 안녕
- Attempt reversible repair of repeatedly mis-encoded logs or exports
Frequently asked questions
Can this repair broken Korean text?
Yes, when Korean UTF-8 bytes were decoded as Latin-1 or Windows-1252. For example, Windows-1252 text 안녕 becomes 안녕. Mis-decoding involving EUC-KR, CP949, or other encodings is not supported.
Can replacement characters or question marks be restored?
Input containing � is rejected because the original bytes may have been lost. Information replaced by ? cannot be restored either, although ? can also be a legitimate character, so the tool cannot detect every case of data loss.
How many passes should I choose?
One mistaken decoding usually needs one pass. If the broken text was re-encoded as UTF-8 and misread again, it may need two. For example, café needs two passes to become café. Too many passes can produce an error.
What happens to already correct text?
ASCII-only input may remain unchanged. Correct Korean and other characters outside the chosen single-byte encoding are rejected. A correct standalone é can also fail because its recovered single byte is not valid UTF-8.
Does this detect or change a file’s encoding?
No. It does not upload raw file bytes or automatically detect an encoding. It reverses information still present in pasted text using two specific single-byte mappings. If bytes were lost, reopen the original source with the correct encoding.
Privacy
Input and results are processed in browser memory. This tool does not upload or save their contents. Shared site advertising, visitor analytics, and aggregate tool-use events may operate separately; those events do not include your input.
Comments & questions