Text File Encoding Workbench
Changing a filename extension does not change the bytes in a CP949 file from an older Windows program. This workbench reads a selected file using the character set you specify and writes a newly encoded byte file. For CP949 it uses the WHATWG EUC-KR index, including extended Korean syllables. If a character has no target mapping, its position and code point are shown and download is blocked. It neither guesses the source encoding nor repairs text already saved in a corrupted form.
Key features
- Read actual file bytes strictly as an explicitly selected UTF-8 or CP949 input
- Re-encode extended Korean syllables using the WHATWG EUC-KR index
- List unmappable character positions and code points and block lossy output
- Detect and preserve, add or remove a UTF-8 BOM; never add a BOM to CP949
- Decode output again to verify identical text and disclose byte identity
- Inspect input/output sizes, a bounded preview and control-character warnings
How to use
- Choose a local text file and specify the encoding used when that file was saved.
- Choose the target encoding and, for UTF-8 output, select whether to preserve, add or remove a BOM.
- Run the local conversion and inspect the preview, byte sizes and BOM status.
- If decoding fails or characters cannot be represented, correct the source setting or target encoding and analyze again.
- Download the new file only after the Unicode round-trip passes, then check it in the receiving application.
Use cases
- Convert a Korean CP949 CSV from older Windows software to UTF-8
- Check whether all Korean text in a TXT file can be sent to a legacy CP949 system
- Add a UTF-8 BOM for an application that needs one, or remove it
- Preserve Korean text and line endings while changing subtitle or log file bytes
Frequently asked questions
Does it automatically detect the source encoding?
No. Specify the actual UTF-8 or CP949 source yourself. Some byte sequences are valid under both choices, so inspect the Korean text in the preview before downloading.
What happens to emoji that CP949 cannot represent?
Their character positions and U+ code points are shown, and file download is blocked. The tool never silently substitutes a question mark or HTML code. Choose UTF-8 output to keep emoji.
What does CP949 mean here?
CP949 here follows the web standard's EUC-KR index, including extended Hangul. Nonstandard code-page variants in other software and UTF-16 files are outside this tool's scope.
Can it recover text already saved as � or garbled characters?
No. Once characters have been lost, their original values cannot be inferred. An existing U+FFFD character is flagged and preserved. For pasted mojibake text, use the separate Encoding Fixer tool.
Does it upload or overwrite my file?
No. One file up to 4 MiB is read locally and a newly named copy is downloaded. The original is untouched. CP949 output has no BOM; UTF-8 output follows your BOM choice.
Privacy
The file is read and converted in this browser tab's memory; this feature does not upload it to the site. The new file is saved only when you click Download, and the original file is untouched. Review downloaded files before sharing sensitive text.
Comments & questions