← Guides

TOOLDAM GUIDE

Why Korean text can break in CSV files

If the program that creates a CSV and the program that opens it assume different character encodings, Korean text can appear corrupted. UTF-8 and BOM handling are common checkpoints.

Key point

CSV files do not always communicate character encoding clearly, so different assumptions between programs can produce broken Korean text.

CSV structure and character encoding are separate

Correct comma-separated columns do not determine how bytes become characters. Reading UTF-8 bytes with another character set can turn Korean into question marks or garbled text.

Why UTF-8 can still behave differently

Some programs use a BOM or the operating system default to guess encoding. The same CSV can therefore open differently in a browser, text editor and spreadsheet app.

Do not overwrite the original when text is broken

Saving a garbled file can alter the original byte data and make recovery harder. Work on a copy and try the appropriate encoding first.

Consider the recipient’s environment

UTF-8 is common for web data exchange, but some business software requires another encoding. Follow the import rules of the destination system when they are specified.