Tools › Text
Encoding Fixer
Repair text that was read in the wrong encoding (“café” for “café”, “’” for “’”), convert a Windows-1252, Latin-1 or UTF-16 file to UTF-8, with or without a byte order mark, and make its line endings all the same.
About this tool What it's for, how to use it and an example
What it's for
Repair text that went through the wrong encoding, and save it as clean UTF-8. Garbled characters such as é for
é or ’ for ’ (known as mojibake) appear when UTF-8 text is read as Windows-1252 or Latin-1. Files saved by
older Windows programs, or exported as “ANSI” or UTF-16, are often turned down by imports that expect UTF-8.
For example, when a CSV exported from Excel shows Renée in another system, or a connector rejects a file
because of its encoding, open the file here and save it as UTF-8, with a byte order mark (BOM) if the import wants
one.
How to use it
Choose a file or drop it on the page, or paste text. For a file, File’s encoding is detected from its bytes: a BOM if it has one, UTF-8 if the bytes make valid UTF-8, UTF-16 if many alternate bytes are zero, and otherwise Windows-1252 (which also covers Latin-1). Choose it yourself if the guess is wrong. Pasted text is already decoded, so only the repairs and line endings apply to it.
- Repair garbled characters finds each run of characters that could be garbled UTF-8 and mends it in place, twice over if it was garbled twice. Runs that don’t turn back into valid UTF-8, such as real accented letters, are left alone.
- Line endings: keep them, or make them all LF or all CRLF. The message line says which kinds the text had.
- Start saved files with a UTF-8 BOM adds the three bytes some programs (Excel among them) look for.
The message line says what was found and done. Copy the result, or Save as a file, which adds -utf8 to
the file’s name. Save keeps the line endings exactly; the Result box, and so Copy, shows them as plain line breaks. The settings, apart from the encoding, are remembered in this browser.
Example
Paste this:
Café, café and it’s, twice: café
The result is Café, café and it’s, twice: café, and the message says 3 garbled runs were repaired. The first
“Café” was already right, so it’s left as it was.
Good to know
Detection is a guess from the bytes: a short Windows-1252 file with no accented characters is the same as UTF-8, which does no harm. Other legacy encodings, such as Shift JIS or KOI8-R, aren’t detected or offered. A garbled run directly next to a correct accented letter may not be repaired. The Character Inspector shows the code points and bytes, if you need to see what’s in the text.
Private: this tool runs in your browser. Nothing you type, paste or choose leaves this page.
Saved you a few minutes? Say thanks with a coffee.
Something wrong with this tool, or missing from it? Report a bug or suggest a feature.