Your files stay in this browser

Advertising and analytics are currently off. We remember only a dismissed-notice preference for this tab. Cloudflare may use technical security cookies. Cookies & storage details · File privacy

Practical guide

Fix garbled characters in a CSV

Understand UTF-8, the byte order mark, and what an encoding change can recover.

Separate bytes from characters

A text file stores bytes. An encoding tells a reader how those bytes represent characters. Choosing the wrong encoding can make an accented name or non-Latin text appear garbled even when the original bytes are still intact.

id,name
00127,Café mug

This teaching example contains é. Compare that character with your source, rather than deciding an encoding is correct merely because the first row of ordinary ASCII letters looks right.

What auto mode actually does

TableMender’s auto mode recognizes a UTF-16 byte order mark when present. Otherwise it attempts a strict UTF-8 decode. If the bytes are not valid UTF-8, it shows an error instead of silently replacing them. This is a check of a proposed encoding, not universal detection of every legacy format.

Try the encoding used by the exporting system

  1. Keep an unchanged copy of the original export.
  2. Choose the file in CSV check & repair.
  3. If auto mode fails, select the original encoding if you know it. This version offers UTF-8, UTF-16 little/big-endian, Windows-1252, Shift JIS and GB18030.
  4. Preview and compare meaningful characters in several records with the exporting system.
  5. Download UTF-8 output only when the decoded values are correct.

Inspect the file encoding →

When a UTF-8 byte order mark helps

A byte order mark at the beginning of a UTF-8 file can help an application recognize its encoding. Microsoft specifically describes opening UTF-8 CSV normally in Excel when the file includes a BOM, and provides import alternatives. Some other systems prefer a file without it. TableMender lets you choose rather than making one option mandatory.

Know what cannot be recovered

If a previous program already saved replacement characters such as �, changing the encoding will not reconstruct the missing text. Likewise, a file can be valid UTF-8 while containing text that was damaged earlier. The tool flags replacement characters as a reason to check the original, not as proof that it can repair them.

Work from the original bytes. Encoding conversion is not a substitute for checking whether the exported values match the source system.

After exporting

Import the new file with UTF-8 selected when your receiving application offers that choice. Check record counts, names and identifiers. Encoding fixes do not set spreadsheet column types; if you also need to preserve leading zeros, use the text-preserving XLSX converter.

Sources:

Examples are original teaching data. Tool-specific behavior describes this version of TableMender; software interfaces can vary by version.

← All guides