Why Text Gets Garbled: ASCII, Unicode, UTF-8/16/32, EUC-KR, CP949 and BOM Explained
Key takeaways
Files and sockets store bytes, and garbled text means the reader used a different encoding from the writer. How ASCII, code pages, Unicode, UTF-8/16/32, EUC-KR and CP949 map characters to bytes, how to tell from the mojibake which mistake happened, where a BOM breaks things, and the defaults in Python, Excel, MySQL and requests that cause most Korean text corruption.
Introduction: Why Should You Know Character Encoding?
When developing, you experience Korean characters getting corrupted, files being unreadable, or API responses appearing strange. The root cause of all these problems is character encoding.
The idea that explains nearly every one of these bugs is simple: a file, a socket or a database column stores bytes, not text. An encoding is the rule for turning text into bytes and back, and in most places that rule is not stored with the data. A plain .txt file does not record whether it was written as UTF-8 or CP949; an HTTP body only says so if the Content-Type header includes a charset; a byte string in a program carries no label at all. Whenever the reader assumes a different rule from the writer, the text is garbled, and the garbling happens silently whenever the wrong rule happens to accept the bytes.
So debugging an encoding problem is always the same two questions: which encoding wrote these bytes, and which encoding is reading them? The rest of this article gives you the background to answer both, with examples in Python and a few other languages.
What This Article Covers:
- History of ASCII, ANSI, Unicode
- UTF-8, UTF-16, UTF-32 encoding methods
- Korean encoding (EUC-KR, CP949)
- BOM, Endian, encoding detection
- Practical problem solving
History of Character Encoding
Timeline
timeline
title Character Encoding Evolution
1963 : ASCII established\n7-bit, 128 chars
1987 : ISO-8859-1 (Latin-1)\n8-bit, 256 chars
1991 : Unicode 1.0\n16-bit unified charset
1992 : UTF-8 invented\nVariable length encoding
1996 : UTF-16\nSurrogate pairs
2003 : UTF-8 web standardization
2008 : UTF-8 most used\non the web
2020s : UTF-8 used by the\nvast majority of websites
Why Do Multiple Encodings Exist?
flowchart TB
Problem["Problem: Computers\nonly understand numbers"]
ASCII["ASCII\n128 English chars"]
Extended["Extended ASCII\n256 chars per language"]
Unicode["Unicode\nGlobal character integration"]
Problem --> ASCII
ASCII --> Extended
Extended --> Unicode
ASCII --> Issue1["Problem: Cannot express\nKorean, Chinese"]
Extended --> Issue2["Problem: Different\ncode pages per country"]
Unicode --> Solution["Solution: Assign unique\nnumber to all characters"]
ASCII: 7-bit Character Set
What is ASCII?
ASCII (American Standard Code for Information Interchange) represents English alphabet, numbers, and special characters with 7 bits (0-127).
ASCII Table
Dec Hex Char | Dec Hex Char | Dec Hex Char
-------------------------------------------------
32 20 Space | 64 40 @ | 96 60 `
33 21 ! | 65 41 A | 97 61 a
34 22 " | 66 42 B | 98 62 b
35 23 # | 67 43 C | 99 63 c
...
48 30 0 | 80 50 P | 112 70 p
49 31 1 | 81 51 Q | 113 71 q
...
57 39 9 | 90 5A Z | 122 7A z
ASCII Control Characters
# Main control characters
NUL = 0x00 # Null
LF = 0x0A # Line Feed (\n)
CR = 0x0D # Carriage Return (\r)
ESC = 0x1B # Escape
DEL = 0x7F # Delete
# Line break methods
# Unix/Linux: LF (\n)
# Windows: CR+LF (\r\n)
# Mac (Classic): CR (\r)
ASCII Examples
# Character → Code
ord('A') # 65
ord('a') # 97
ord('0') # 48
# Code → Character
chr(65) # 'A'
chr(97) # 'a'
# Check ASCII range
def is_ascii(text):
return all(ord(c) < 128 for c in text)
is_ascii("Hello") # True
is_ascii("안녕") # False
ANSI and Code Pages
What is ANSI?
ANSI extends to 8 bits (0-255) to support each country’s language. However, the meaning of the 128-255 range differs per Code Page.
“ANSI” is a Windows misnomer rather than a real encoding name. In Windows APIs and in Notepad’s old “Save as ANSI” option it means “the system’s current legacy code page”, which is Windows-1252 on a Western European or US installation and CP949 on a Korean one. The same “ANSI” file therefore contains different bytes depending on which machine saved it, which is exactly why files exchanged between Korean and non-Korean Windows machines used to arrive garbled. East Asian code pages like CP949, CP932 and CP936 are not single-byte at all: bytes above 127 start a two-byte sequence, which is why the single byte 0xC7 below cannot be decoded as CP949 on its own.
Major Code Pages
| Code Page | Name | Region | Features |
|---|---|---|---|
| CP437 | OEM-US | USA | DOS default |
| CP850 | Latin-1 | Western Europe | DOS multilingual |
| CP949 | Extended Complete | Korea | Windows Korean |
| CP932 | Shift-JIS | Japan | Windows Japanese |
| CP936 | GBK | China | Windows Chinese |
| ISO-8859-1 | Latin-1 | Western Europe | Unix/Web |
| ISO-8859-15 | Latin-9 | Western Europe | Euro (€) added |
Code Page Problems
# Same byte value, different meaning
byte_value = 0xC7
# CP949 (Korean): '한'
text_korean = byte_value.to_bytes(1, 'big').decode('cp949') # Error (needs 2 bytes)
# ISO-8859-1 (Latin-1): 'Ç'
text_latin = byte_value.to_bytes(1, 'big').decode('latin-1') # 'Ç'
# Reading same file with different encoding causes corruption!
Unicode: Global Character Integration
What is Unicode?
Unicode is a character set that assigns unique Code Points to all characters worldwide.
Unicode Structure
U+0000 ~ U+10FFFF (1,114,112 code points)
U+0000 ~ U+007F : ASCII (128 chars)
U+0080 ~ U+00FF : Latin-1 Supplement
U+0100 ~ U+017F : Latin Extended-A
U+0370 ~ U+03FF : Greek
U+0400 ~ U+04FF : Cyrillic
U+0600 ~ U+06FF : Arabic
U+0E00 ~ U+0E7F : Thai
U+3040 ~ U+309F : Hiragana (Japanese)
U+30A0 ~ U+30FF : Katakana (Japanese)
U+4E00 ~ U+9FFF : CJK Unified Ideographs (Chinese/Japanese/Korean)
U+AC00 ~ U+D7AF : Hangul Syllables (Korean 11,172 chars)
U+1F600 ~ U+1F64F : Emoticons (Emoji)
Korean Unicode Range
# Korean syllables (가-힣)
print(f"가: U+{ord('가'):04X}") # U+AC00
print(f"힣: U+{ord('힣'):04X}") # U+D7A3
# Korean letters (ㄱ-ㅎ, ㅏ-ㅣ)
print(f"ㄱ: U+{ord('ㄱ'):04X}") # U+3131
print(f"ㅎ: U+{ord('ㅎ'):04X}") # U+314E
print(f"ㅏ: U+{ord('ㅏ'):04X}") # U+314F
print(f"ㅣ: U+{ord('ㅣ'):04X}") # U+3163
# Emoji
print(f"😀: U+{ord('😀'):04X}") # U+1F600
Unicode vs Encoding
Unicode: Character Set
Assigns number (code point) to each character
Example: '한' = U+D55C
UTF-8/UTF-16/UTF-32: Encoding
Method to convert code points to bytes
Example: U+D55C → UTF-8: ED 95 9C (3 bytes)
→ UTF-16: D5 5C (2 bytes)
Three different units are easy to confuse, and most “string length” bugs come from mixing them up. A code point is the number Unicode assigns (U+D55C). A code unit is the fixed-size piece an encoding works in: one byte in UTF-8, two bytes in UTF-16, four in UTF-32. A grapheme cluster is what a user perceives as one character, which may be several code points: a flag emoji is two code points, a family emoji joined with zero-width joiners can be seven or more, and a Hangul syllable written in decomposed form (see Normalization below) is two or three. len() in Python counts code points, .length in JavaScript and Java counts UTF-16 code units, and strlen in C counts bytes. None of them counts what the user sees, which matters when you truncate a string for display or enforce a “maximum 20 characters” rule.
UTF-8: Variable Length Encoding
What is UTF-8?
UTF-8 encodes Unicode with 1-4 byte variable length. It’s the web standard and perfectly compatible with ASCII.
UTF-8 Encoding Rules
Code Point Range | Bytes | Encoding Pattern
U+0000 ~ U+007F | 1 | 0xxxxxxx
U+0080 ~ U+07FF | 2 | 110xxxxx 10xxxxxx
U+0800 ~ U+FFFF | 3 | 1110xxxx 10xxxxxx 10xxxxxx
U+10000 ~ U+10FFFF | 4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
UTF-8 Encoding Examples
English ‘A’ (U+0041)
Code Point: U+0041 (65)
Binary: 0100 0001
UTF-8 Encoding:
0100 0001 = 0x41 (1 byte)
Memory: 41
Korean ‘한’ (U+D55C)
Code Point: U+D55C (54,620)
Binary: 1101 0101 0101 1100
UTF-8 Encoding (3 bytes):
1110xxxx 10xxxxxx 10xxxxxx
1110 1101 10 010101 10 011100
E D 9 5 9 C
Memory: ED 95 9C
Emoji ’😀’ (U+1F600)
Code Point: U+1F600 (128,512)
Binary: 0001 1111 0110 0000 0000
UTF-8 Encoding (4 bytes):
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
11110 000 10 011111 10 011000 10 000000
F 0 9 F 9 8 8 0
Memory: F0 9F 98 80
The bit patterns are what give UTF-8 its useful properties. Every byte of a multi-byte sequence has its high bit set, so no byte of a Korean or emoji character can ever be mistaken for an ASCII character such as /, " or a NUL terminator. That is why UTF-8 text passes through C string functions, file paths and protocols designed for ASCII without breaking them. Lead bytes (110, 1110, 11110) and continuation bytes (10) are distinguishable, so a decoder that starts in the middle of a stream can skip forward to the next lead byte and resynchronize after at most three bytes.
The same structure makes UTF-8 easy to validate, and this is where the most useful error messages come from. A random byte sequence is very unlikely to be valid UTF-8, so when a file in CP949 or Latin-1 is decoded as UTF-8, the decoder usually fails quickly with a message such as 'utf-8' codec can't decode byte 0xc7 in position 0: invalid continuation byte. The reverse is not true: CP949 and Latin-1 accept most byte sequences, so decoding UTF-8 with them tends to “succeed” and produce garbage. When you have to guess, try strict UTF-8 first.
Two properties are sometimes overlooked. Some byte sequences are always invalid in UTF-8: an “overlong” encoding (such as C0 AF for /) or an encoded surrogate (ED A0 80) must be rejected, because accepting them has been used to smuggle characters like / past security filters. And cutting a UTF-8 string at an arbitrary byte offset can split a character, leaving an invalid sequence at the end; truncating to a byte limit (for a database column or a message size) must back up to a lead byte.
UTF-8 Advantages
flowchart TB
UTF8[UTF-8]
Adv1["✅ ASCII compatible\nEnglish is 1 byte"]
Adv2["✅ Self-synchronizing\nCan read from middle"]
Adv3["✅ Byte order independent\nNo Endian issues"]
Adv4["✅ Web standard\nDefault for HTML5 and JSON"]
UTF8 --> Adv1
UTF8 --> Adv2
UTF8 --> Adv3
UTF8 --> Adv4
UTF-8 Encoding with Python
# String → Bytes
text = "Hello 한글 😀"
# UTF-8 encoding
utf8_bytes = text.encode('utf-8')
print(utf8_bytes)
# b'Hello \xed\x95\x9c\xea\xb8\x80 \xf0\x9f\x98\x80'
# Byte analysis
for i, byte in enumerate(utf8_bytes):
print(f"{i:2d}: 0x{byte:02X} ({byte:3d}) {chr(byte) if byte < 128 else '?'}")
# Output:
# 0: 0x48 ( 72) H
# 1: 0x65 (101) e
# 2: 0x6C (108) l
# 3: 0x6C (108) l
# 4: 0x6F (111) o
# 5: 0x20 ( 32)
# 6: 0xED (237) ? ← '한' start
# 7: 0x95 (149) ?
# 8: 0x9C (156) ?
# 9: 0xEA (234) ? ← '글' start
# 10: 0xB8 (184) ?
# 11: 0x80 (128) ?
# 12: 0x20 ( 32)
# 13: 0xF0 (240) ? ← '😀' start
# 14: 0x9F (159) ?
# 15: 0x98 (152) ?
# 16: 0x80 (128) ?
# Bytes → String
decoded = utf8_bytes.decode('utf-8')
print(decoded) # "Hello 한글 😀"
UTF-16 and UTF-32
UTF-16
UTF-16 encodes with 2 or 4 bytes. Used internally in Windows, Java, and JavaScript.
UTF-16 Encoding Rules
Code Point Range | Bytes | Method
U+0000 ~ U+FFFF | 2 | Direct encoding
U+10000 ~ U+10FFFF | 4 | Surrogate pair
Surrogate Pair
# Encode emoji '😀' (U+1F600) to UTF-16
# 1. U+1F600 - 0x10000 = 0xF600
# 2. High 10 bits: 0x3D (61)
# 3. Low 10 bits: 0x200 (512)
# 4. High Surrogate: 0xD800 + 0x3D = 0xD83D
# 5. Low Surrogate: 0xDC00 + 0x200 = 0xDE00
text = "😀"
utf16_bytes = text.encode('utf-16-le')
print(utf16_bytes.hex()) # '3dd8 00de' (Little-Endian)
# UTF-16 BE (Big-Endian)
utf16_be = text.encode('utf-16-be')
print(utf16_be.hex()) # 'd83d de00'
UTF-16 Example
text = "Hello 한글"
# UTF-16 LE (Little-Endian)
utf16_le = text.encode('utf-16-le')
print(utf16_le.hex())
# 48 00 65 00 6c 00 6c 00 6f 00 20 00 5c d5 00 ae
# UTF-16 BE (Big-Endian)
utf16_be = text.encode('utf-16-be')
print(utf16_be.hex())
# 00 48 00 65 00 6c 00 6c 00 6f 00 20 d5 5c ae 00
UTF-16 was designed when Unicode was expected to fit in 16 bits, and Windows (wchar_t, the W APIs), Java and JavaScript adopted it in that era. When Unicode grew past U+FFFF, surrogate pairs were added, and code written with the “one character equals one 16-bit unit” assumption broke for emoji and rare CJK characters. In JavaScript, "😀".length is 2, "😀".charAt(0) is a lone high surrogate, and slicing a string in the middle of a pair produces text that some systems reject and others display as �. Use for...of, Array.from(str) or codePointAt() when you need code points, and Intl.Segmenter when you need user-visible characters.
UTF-16 also has two byte orders, which is why files saved as “Unicode” by older Windows tools begin with a BOM and why a UTF-16 file opened as UTF-8 shows every ASCII letter separated by NUL bytes (H\0e\0l\0l\0o\0). For storage and transmission there is rarely a reason to choose it today; it is mainly an in-memory representation you meet at API boundaries.
UTF-32
UTF-32 encodes all characters with fixed 4-byte length.
text = "A한😀"
# UTF-32 LE
utf32 = text.encode('utf-32-le')
print(utf32.hex())
# 41 00 00 00 5c d5 00 00 00 f6 01 00
# Each character is exactly 4 bytes
# 'A': 0x00000041
# '한': 0x0000D55C
# '😀': 0x0001F600
Encoding Comparison
text = "Hello 한글 😀"
encodings = ['utf-8', 'utf-16-le', 'utf-16-be', 'utf-32-le']
for enc in encodings:
encoded = text.encode(enc)
print(f"{enc:12s}: {len(encoded):2d} bytes | {encoded.hex()[:40]}...")
# Output:
# utf-8 : 17 bytes | 48656c6c6f20ed959ceab88020f09f9880...
# utf-16-le : 22 bytes | 480065006c006c006f0020005cd500ae20003dd8...
# utf-16-be : 22 bytes | 00480065006c006c006f0020d55cae000020d83d...
# utf-32-le : 40 bytes | 48000000650000006c0000006c0000006f000000...
Korean Encoding (EUC-KR, CP949)
Korean Encoding History
timeline
title Korean Encoding Evolution
1987 : KS X 1001\nComplete 2,350 chars
1992 : EUC-KR\nComplete standard
1996 : CP949 (MS)\nExtended complete 11,172 chars
2000s : UTF-8\nUnicode based
EUC-KR
EUC-KR represents 2,350 Korean characters with 2 bytes.
# EUC-KR encoding
text = "한글"
euckr_bytes = text.encode('euc-kr')
print(euckr_bytes.hex()) # c7d1 b1db
# '한': 0xC7D1
# '글': 0xB1DB
# Problem: Characters like '똠', '쀍' are not among the 2,350 syllables.
# Python does not raise here: it emits an 8-byte KS X 1001 "make-up
# sequence" (filler + jamo) that many other EUC-KR decoders do not understand.
print("똠".encode('euc-kr').hex()) # a4d4a4a8a4c7a4b1 (8 bytes)
print("똠".encode('cp949').hex()) # 8c63 (2 bytes, CP949 extension area)
EUC-KR is a byte-level wrapper around the KS X 1001 character set, which includes only 2,350 precomposed Hangul syllables chosen as the most common ones. The other 8,822 modern syllables, such as 똠 or 쀍, have no code. KS X 1001 defines an optional “make-up sequence” that spells such a syllable as eight bytes of filler and jamo, and Python’s euc-kr codec produces it, but most other software does not decode it, so the practical result is still broken text on the receiving side.
CP949 (Extended Complete)
CP949 extends EUC-KR to support all 11,172 characters.
Microsoft’s CP949 (also called Unified Hangul Code, or “extended Wansung”) keeps every EUC-KR code unchanged and places the missing 8,822 syllables in byte ranges EUC-KR does not use. That makes CP949 a strict superset: any valid EUC-KR text decodes identically as CP949, but not the other way round. In practice, text labeled euc-kr on the web or in old databases is very often CP949, which is why browsers and the WHATWG Encoding Standard treat the label euc-kr as CP949. When you decode Korean legacy data in your own code, use cp949 even if the source claims euc-kr; it decodes everything EUC-KR does and the extra syllables too.
# CP949 encoding
text = "똠방각하"
cp949_bytes = text.encode('cp949')
print(cp949_bytes.hex())
# Can represent characters not in EUC-KR
text2 = "쀍똠뙠"
print(text2.encode('cp949').hex())
UTF-8 vs EUC-KR Comparison
text = "Hello 한글"
# UTF-8: English 1 byte, Korean 3 bytes
utf8 = text.encode('utf-8')
print(f"UTF-8: {len(utf8)} bytes | {utf8.hex()}")
# UTF-8: 14 bytes | 48656c6c6f20ed959ceab880
# EUC-KR: English 1 byte, Korean 2 bytes
euckr = text.encode('euc-kr')
print(f"EUC-KR: {len(euckr)} bytes | {euckr.hex()}")
# EUC-KR: 10 bytes | 48656c6c6f20c7d1b1db
BOM and Endian
BOM (Byte Order Mark)
BOM is a special byte at the start of a file indicating encoding method and byte order.
Encoding | BOM (hex) | Size
UTF-8 | EF BB BF | 3 bytes
UTF-16 LE | FF FE | 2 bytes
UTF-16 BE | FE FF | 2 bytes
UTF-32 LE | FF FE 00 00 | 4 bytes
UTF-32 BE | 00 00 FE FF | 4 bytes
BOM Example
# UTF-8 with BOM
text = "Hello"
with open('file_with_bom.txt', 'wb') as f:
f.write(b'\xef\xbb\xbf') # BOM
f.write(text.encode('utf-8'))
# File content (hex):
# EF BB BF 48 65 6C 6C 6F
# ^^^^^^^^ BOM
# ^^^^^^^^^^^^^^ "Hello"
# UTF-8 without BOM (recommended)
with open('file_no_bom.txt', 'wb') as f:
f.write(text.encode('utf-8'))
# File content (hex):
# 48 65 6C 6C 6F
BOM Detection and Removal
def detect_and_remove_bom(data):
"""Detect and remove BOM"""
bom_signatures = [
(b'\xef\xbb\xbf', 'utf-8-sig'),
(b'\xff\xfe\x00\x00', 'utf-32-le'),
(b'\x00\x00\xfe\xff', 'utf-32-be'),
(b'\xff\xfe', 'utf-16-le'),
(b'\xfe\xff', 'utf-16-be'),
]
for bom, encoding in bom_signatures:
if data.startswith(bom):
return data[len(bom):], encoding
return data, None
# Usage
with open('file.txt', 'rb') as f:
data = f.read()
data, encoding = detect_and_remove_bom(data)
if encoding:
print(f"✅ BOM detected: {encoding}")
text = data.decode(encoding.replace('-sig', ''))
else:
print("ℹ️ No BOM, assuming UTF-8")
text = data.decode('utf-8')
For UTF-16 and UTF-32 the BOM carries real information (the byte order). For UTF-8 there is no byte order to indicate, and the BOM is only a signature saying “this is UTF-8”. The Unicode standard permits it but does not recommend it, and it causes a specific set of failures when a program does not expect it:
- A shell script saved with a BOM fails with an error like
#!/bin/bash: No such file or directory, because the first bytes are no longer#!. - PHP files with a BOM send three bytes of output before any code runs, which leads to
Cannot modify header information - headers already sent. - CSV readers include the BOM in the first column name, so a lookup of
row['id']fails because the key is actually'id'. - Some JSON parsers reject it (
JSON.parsein JavaScript fails withUnexpected tokenon a string starting with U+FEFF), and RFC 8259 says JSON must not be written with one.
The one place a UTF-8 BOM helps is Microsoft Excel, which uses it to recognize UTF-8 CSV files; without it, Excel on a Korean Windows opens UTF-8 CSV as CP949 and shows mojibake. A common compromise is to write CSV intended for Excel with utf-8-sig and everything else without a BOM, and to read with utf-8-sig so either form works.
Endian (Byte Order)
# Big-Endian: Large byte first
# Little-Endian: Small byte first
# Example: Store 0x1234 in memory
# Big-Endian: 12 34
# Little-Endian: 34 12
# Important in UTF-16
text = "한" # U+D55C
# UTF-16 BE (Big-Endian)
be = text.encode('utf-16-be')
print(be.hex()) # d5 5c
# UTF-16 LE (Little-Endian)
le = text.encode('utf-16-le')
print(le.hex()) # 5c d5
# UTF-8 is byte-based, so Endian independent
utf8 = text.encode('utf-8')
print(utf8.hex()) # ed 95 9c (always same)
Practical Problem Solving
Problem 1: Korean Character Corruption (���)
Cause
# ❌ Saved as UTF-8 but read as EUC-KR
with open('file.txt', 'w', encoding='utf-8') as f:
f.write("한글")
# Incorrect reading
with open('file.txt', 'r', encoding='euc-kr') as f:
text = f.read()
print(text) # UnicodeDecodeError: 'euc_kr' codec can't decode byte 0xed
# in position 0: illegal multibyte sequence
# With errors='replace' you get mojibake instead: '�븳湲�'
Solution
# ✅ Read with correct encoding
with open('file.txt', 'r', encoding='utf-8') as f:
text = f.read()
print(text) # '한글' (correct)
# ✅ Auto-detect encoding
import chardet
with open('file.txt', 'rb') as f:
raw_data = f.read()
result = chardet.detect(raw_data)
encoding = result['encoding']
confidence = result['confidence']
print(f"Detected: {encoding} ({confidence*100:.1f}% confidence)")
text = raw_data.decode(encoding)
print(text)
I’ve chased this exact ��� pattern to its root cause enough times that I can usually name the culprit before opening a hex editor: it’s a file written in one encoding and read by something that assumes another — an older CSV import tool defaulting to the OS locale, a text editor that remembered the wrong encoding from a previous file, or (most often, in my experience) a CP949 CSV exported from Excel on Korean Windows being fed to pandas.read_csv(), which assumes UTF-8 unless you pass encoding=. chardet’s statistical detection genuinely helps when the source encoding is unknown, but it’s a guess, not a certainty — for byte sequences short enough (a single word, a short filename), EUC-KR and UTF-8 can both look like plausible decodings with similar confidence scores, so treat a chardet result under roughly 80% confidence as a hint to verify manually, not as ground truth.
The shape of the garbage tells you which way the mistake went, which is often faster than any detection tool:
| What you see for “한글” | Bytes were | Read as |
|---|---|---|
한글 (accented Latin letters and symbols) | UTF-8 | Windows-1252 / Latin-1 |
�븳湲� (unrelated Hangul and Hanja, some �) | UTF-8 | CP949 / EUC-KR (with replacement) |
�ѱ� or only ���� | CP949 | UTF-8 (with replacement) |
H\0e\0l\0l\0o or letters with gaps | UTF-16 | UTF-8 or a single-byte encoding |
??? (literal question marks) | Lost during encoding | The text was already converted to a charset that lacks the characters |
The last row is the one that cannot be repaired. Real question marks mean an encoder replaced characters it could not represent (for example writing Korean through a Latin-1 connection to a database), and the original characters are gone from the stored data. Everything else in the table is reversible as long as you still have the original bytes: re-read them with the right encoding. What makes mojibake permanent is saving it again, because then the garbled text is encoded a second time.
Problem 2: UnicodeDecodeError
# ❌ Decode with wrong encoding
utf8_bytes = "한글".encode('utf-8')
try:
text = utf8_bytes.decode('ascii')
except UnicodeDecodeError as e:
print(f"❌ {e}")
# 'ascii' codec can't decode byte 0xed in position 0
# ✅ Error handling options
# 1. Ignore
text = utf8_bytes.decode('ascii', errors='ignore')
print(text) # "" (Korean removed)
# 2. Replace
text = utf8_bytes.decode('ascii', errors='replace')
print(text) # "������" (each bad byte becomes U+FFFD)
# 3. Keep the raw bytes visible (decode-side handler)
text = utf8_bytes.decode('ascii', errors='backslashreplace')
print(text) # "\xed\x95\x9c\xea\xb8\x80"
# Note: 'xmlcharrefreplace' only works when encoding; passing it to decode()
# raises TypeError: don't know how to handle UnicodeDecodeError in error callback
Problem 3: Korean Character Corruption on Web
import requests
# ❌ Wrong method
response = requests.get('https://example.com/korean-page')
print(response.text) # May be corrupted
# ✅ Check Content-Type header
response = requests.get('https://example.com/korean-page')
content_type = response.headers.get('Content-Type', '')
print(f"Content-Type: {content_type}")
# Content-Type: text/html; charset=euc-kr
# ✅ Decode with correct encoding
if 'euc-kr' in content_type.lower():
text = response.content.decode('euc-kr')
else:
text = response.text # requests auto-detects
# ✅ Or auto-detect with chardet
import chardet
detected = chardet.detect(response.content)
text = response.content.decode(detected['encoding'])
requests takes the encoding from the charset in Content-Type. If a text/* response has no charset, it falls back to ISO-8859-1 as the old HTTP/1.1 specification required, which is almost never right for Korean pages; response.apparent_encoding runs a detector on the body and is a better fallback, and a <meta charset> inside the HTML is not consulted by response.text at all.
Problem 4: CSV File Encoding
import csv
# ❌ CSV saved from Windows Excel (CP949)
with open('data.csv', 'r', encoding='utf-8') as f:
reader = csv.reader(f)
for row in reader:
print(row) # UnicodeDecodeError!
# ✅ Correct encoding
with open('data.csv', 'r', encoding='cp949') as f:
reader = csv.reader(f)
for row in reader:
print(row)
# ✅ Auto-detect encoding
import chardet
with open('data.csv', 'rb') as f:
raw_data = f.read()
detected = chardet.detect(raw_data)
encoding = detected['encoding']
with open('data.csv', 'r', encoding=encoding) as f:
reader = csv.reader(f)
for row in reader:
print(row)
Excel is the most common source of this problem. On Korean Windows, “CSV (Comma delimited)” is written in CP949, while “CSV UTF-8 (Comma delimited)” is written in UTF-8 with a BOM. A pipeline that accepts CSV uploads from users will receive both, so either tell users which format to choose or try utf-8-sig first and fall back to cp949 on UnicodeDecodeError. That order matters: UTF-8 is strict and fails fast on CP949 input, while CP949 would accept many UTF-8 byte sequences and silently produce mojibake.
Programming Language-specific Handling
Python
Python 3 strings are Unicode, and source files are UTF-8 by default, but open() without encoding= uses the locale’s encoding, not UTF-8. On a Korean Windows machine that is CP949, and the typical symptom is code that works on macOS or Linux and fails on a colleague’s Windows laptop with UnicodeDecodeError: 'cp949' codec can't decode byte 0xec in position 12: illegal multibyte sequence. Always pass encoding='utf-8' explicitly. Python 3.15 makes UTF-8 mode the default (PEP 686); until then, setting PYTHONUTF8=1 or running python -X utf8 gives the same behavior on older versions.
# Default encoding: UTF-8
text = "Hello 한글 😀"
# Encoding
utf8 = text.encode('utf-8')
utf16 = text.encode('utf-16')
euckr = text.encode('euc-kr') # Emoji causes error
# Decoding
text = utf8.decode('utf-8')
# File I/O
with open('file.txt', 'w', encoding='utf-8') as f:
f.write(text)
with open('file.txt', 'r', encoding='utf-8') as f:
text = f.read()
# Byte string literal
utf8_bytes = b'\xed\x95\x9c\xea\xb8\x80'
text = utf8_bytes.decode('utf-8') # "한글"
JavaScript/Node.js
// JavaScript internal: UTF-16
const text = "Hello 한글 😀";
// String length (caution: surrogate pairs)
console.log(text.length); // 11 (😀 counted as 2)
// Correct length
console.log([...text].length); // 10
// UTF-8 encoding (Node.js)
const buffer = Buffer.from(text, 'utf-8');
console.log(buffer); // <Buffer 48 65 6c 6c 6f 20 ...>
// Decoding
const decoded = buffer.toString('utf-8');
console.log(decoded); // "Hello 한글 😀"
// Supported encodings
// utf-8, utf-16le, latin1, base64, hex, ascii
Java
// Java internal: UTF-16
String text = "Hello 한글 😀";
// UTF-8 encoding
byte[] utf8Bytes = text.getBytes(StandardCharsets.UTF_8);
System.out.println(Arrays.toString(utf8Bytes));
// Decoding
String decoded = new String(utf8Bytes, StandardCharsets.UTF_8);
System.out.println(decoded);
// File I/O
// Write as UTF-8
Files.writeString(
Path.of("file.txt"),
text,
StandardCharsets.UTF_8
);
// Read as UTF-8
String content = Files.readString(
Path.of("file.txt"),
StandardCharsets.UTF_8
);
C++
#include <iostream>
#include <fstream>
#include <string>
#include <codecvt>
#include <locale>
int main() {
// UTF-8 string (C++11)
// u8"..." is a const char8_t[] in C++20 and no longer converts to std::string;
// a plain literal is UTF-8 if the source is UTF-8 (MSVC needs /utf-8)
std::string utf8_str = "Hello 한글 😀";
// UTF-16 string
std::u16string utf16_str = u"Hello 한글 😀";
// UTF-32 string
std::u32string utf32_str = U"Hello 한글 😀";
// UTF-8 → UTF-16 conversion (wstring_convert/codecvt_utf8_utf16 are
// deprecated since C++17; prefer ICU or a small library such as simdutf)
std::wstring_convert<std::codecvt_utf8_utf16<char16_t>, char16_t> converter;
std::u16string utf16 = converter.from_bytes(utf8_str);
// Write file (UTF-8)
std::ofstream file("file.txt", std::ios::binary);
file << utf8_str;
file.close();
// Read file
std::ifstream input("file.txt", std::ios::binary);
std::string content((std::istreambuf_iterator<char>(input)),
std::istreambuf_iterator<char>());
std::cout << content << std::endl;
return 0;
}
Go
package main
import (
"fmt"
"unicode/utf8"
"golang.org/x/text/encoding/korean"
"golang.org/x/text/transform"
"io"
"strings"
)
func main() {
// Go internal: UTF-8
text := "Hello 한글 😀"
// Byte length vs character (rune) length
fmt.Println("Bytes:", len(text)) // 17
fmt.Println("Runes:", utf8.RuneCountInString(text)) // 10
// UTF-8 → EUC-KR conversion
encoder := korean.EUCKR.NewEncoder()
euckrBytes, _, _ := transform.Bytes(encoder, []byte(text))
fmt.Printf("EUC-KR: %x\n", euckrBytes)
// EUC-KR → UTF-8 conversion
decoder := korean.EUCKR.NewDecoder()
utf8Text, _, _ := transform.String(decoder, string(euckrBytes))
fmt.Println(utf8Text)
}
Advanced Topics
Normalization
import unicodedata
# Two ways to represent Korean '가'
# 1. Composed (NFC): U+AC00
nfc = "가"
print(f"NFC: {len(nfc)} chars, {nfc.encode('utf-8').hex()}")
# NFC: 1 chars, eab080
# 2. Decomposed (NFD): U+1100 + U+1161 (ㄱ + ㅏ)
nfd = unicodedata.normalize('NFD', nfc)
print(f"NFD: {len(nfd)} chars, {nfd.encode('utf-8').hex()}")
# NFD: 2 chars, e18480e185a1
# Comparison
print(nfc == nfd) # False (different byte sequence)
# Compare after normalization
print(unicodedata.normalize('NFC', nfc) ==
unicodedata.normalize('NFC', nfd)) # True
Normalization is a real problem for Korean text, not a theoretical one. macOS file systems have historically stored file names in a decomposed form, so a file named 한글.txt created on a Mac and zipped or synced to Windows or Linux can arrive with a name that looks identical but consists of separate jamo. Some tools then display it as broken jamo (ㅎㅏㄴㄱㅡㄹ), a search for the name does not find it, and two “identical” names can coexist in one folder. Normalizing to NFC before comparing or storing user-supplied names avoids all three symptoms.
Encoding Detection
import chardet
def detect_encoding(file_path):
"""Auto-detect file encoding"""
with open(file_path, 'rb') as f:
raw_data = f.read()
result = chardet.detect(raw_data)
return {
'encoding': result['encoding'],
'confidence': result['confidence'],
'language': result.get('language', '')
}
# Usage
info = detect_encoding('unknown.txt')
print(f"Encoding: {info['encoding']}")
print(f"Confidence: {info['confidence']*100:.1f}%")
# Read with correct encoding
with open('unknown.txt', 'r', encoding=info['encoding']) as f:
content = f.read()
Encoding Conversion
def convert_file_encoding(input_file, output_file, from_enc, to_enc):
"""Convert file encoding"""
# Read original
with open(input_file, 'r', encoding=from_enc) as f:
content = f.read()
# Save with new encoding
with open(output_file, 'w', encoding=to_enc) as f:
f.write(content)
print(f"✅ Converted: {from_enc} → {to_enc}")
# EUC-KR → UTF-8 conversion
convert_file_encoding('old.txt', 'new.txt', 'euc-kr', 'utf-8')
Encoding in Web Development
HTML
<!DOCTYPE html>
<html>
<head>
<!-- ✅ UTF-8 declaration (required) -->
<meta charset="UTF-8">
<title>Korean Page</title>
</head>
<body>
<h1>안녕하세요</h1>
</body>
</html>
HTTP Headers
from flask import Flask, Response
app = Flask(__name__)
@app.route('/korean')
def korean_page():
content = "<h1>안녕하세요</h1>"
# ✅ Specify charset in Content-Type
return Response(
content,
mimetype='text/html; charset=utf-8'
)
# ❌ Without charset, browser guesses (may corrupt)
JSON
import json
data = {"name": "홍길동", "message": "안녕하세요"}
# JSON is UTF-8 by default
json_str = json.dumps(data, ensure_ascii=False)
print(json_str)
# {"name": "홍길동", "message": "안녕하세요"}
# ensure_ascii=True (default)
json_str_ascii = json.dumps(data, ensure_ascii=True)
print(json_str_ascii)
# {"name": "\ud64d\uae38\ub3d9", "message": "\uc548\ub155\ud558\uc138\uc694"}
URL Encoding
from urllib.parse import quote, unquote
# URL with Korean
text = "한글 검색"
# URL encoding (UTF-8 based)
encoded = quote(text)
print(encoded)
# %ED%95%9C%EA%B8%80%20%EA%B2%80%EC%83%89
# URL decoding
decoded = unquote(encoded)
print(decoded) # "한글 검색"
# Complete URL
url = f"https://example.com/search?q={encoded}"
print(url)
# https://example.com/search?q=%ED%95%9C%EA%B8%80%20%EA%B2%80%EC%83%89
Database Encoding
MySQL
-- Create database (UTF-8)
CREATE DATABASE mydb
CHARACTER SET utf8mb4
COLLATE utf8mb4_unicode_ci;
-- utf8mb4: 4-byte UTF-8 (emoji support)
-- utf8: 3-byte UTF-8 (no emoji, deprecated)
-- Create table
CREATE TABLE users (
id INT PRIMARY KEY,
name VARCHAR(100) CHARACTER SET utf8mb4
);
-- Set encoding on connection
SET NAMES utf8mb4;
PostgreSQL
-- Create database
CREATE DATABASE mydb
ENCODING 'UTF8'
LC_COLLATE 'ko_KR.UTF-8'
LC_CTYPE 'ko_KR.UTF-8';
-- Check client encoding
SHOW client_encoding;
-- Change encoding
SET client_encoding TO 'UTF8';
MySQL’s utf8 is a historical trap: it is an alias for utf8mb3, which stores at most three bytes per character, so emoji and some rare CJK characters cannot be stored. The failure is an error like Incorrect string value: '\xF0\x9F\x98\x80' for column 'name' at row 1 in strict mode, or silently truncated text in older non-strict configurations. The column, the table default, and the connection all need utf8mb4; a utf8mb4 column written through a connection set to latin1 still stores garbage, which is why the connection charset appears in both the SQL and the Python examples here.
In PostgreSQL, creating a database whose LC_COLLATE/LC_CTYPE differ from the template’s fails with an error such as new collation (ko_KR.UTF-8) is incompatible with the collation of the template database (C.UTF-8); add TEMPLATE template0 to the CREATE DATABASE statement. The locale must also exist on the server (locale -a), which is often not the case in minimal Docker images. Collation affects sorting and comparison only; the ENCODING is what determines which characters can be stored.
Python + DB
import psycopg2
# PostgreSQL connection
conn = psycopg2.connect(
host='localhost',
database='mydb',
user='user',
password='pass',
client_encoding='utf8' # ✅ Explicit specification
)
cursor = conn.cursor()
# Insert Korean data
cursor.execute(
"INSERT INTO users (name) VALUES (%s)",
("홍길동",)
)
# Query
cursor.execute("SELECT name FROM users")
name = cursor.fetchone()[0]
print(name) # "홍길동"
Practical Tools
Command Line Tools
# 1. Check encoding with file command
file -i file.txt
# file.txt: text/plain; charset=utf-8
# 2. Convert encoding with iconv
iconv -f EUC-KR -t UTF-8 old.txt > new.txt
# 3. Batch convert multiple files
find . -name "*.txt" -exec iconv -f EUC-KR -t UTF-8 {} -o {}.utf8 \;
# 4. Check bytes with hexdump
echo "한글" | hexdump -C
# 00000000 ed 95 9c ea b8 80 0a
# 5. Remove BOM
tail -c +4 file_with_bom.txt > file_no_bom.txt # UTF-8 BOM (3 bytes)
Python Script
#!/usr/bin/env python3
"""
Batch file encoding conversion tool
"""
import os
import sys
import chardet
from pathlib import Path
def convert_directory(directory, from_enc=None, to_enc='utf-8'):
"""Convert encoding of all text files in directory"""
for file_path in Path(directory).rglob('*.txt'):
try:
# Read original
with open(file_path, 'rb') as f:
raw_data = f.read()
# Detect encoding
if from_enc is None:
detected = chardet.detect(raw_data)
source_enc = detected['encoding']
confidence = detected['confidence']
if confidence < 0.7:
print(f"⚠️ {file_path}: Low confidence ({confidence:.2f})")
continue
else:
source_enc = from_enc
# Skip if already UTF-8
if source_enc.lower().replace('-', '') == 'utf8':
print(f"✓ {file_path}: Already UTF-8")
continue
# Convert
text = raw_data.decode(source_enc)
# Save
with open(file_path, 'w', encoding=to_enc) as f:
f.write(text)
print(f"✅ {file_path}: {source_enc} → {to_enc}")
except Exception as e:
print(f"❌ {file_path}: {e}")
if __name__ == '__main__':
if len(sys.argv) < 2:
print("Usage: python convert_encoding.py <directory>")
sys.exit(1)
convert_directory(sys.argv[1])
Two cautions about this kind of batch script. It rewrites files in place, so run it on a copy or under version control: a wrong detection result on a short file converts it into permanent mojibake. And chardet reports pure-ASCII files as ascii, which is harmless here (ASCII is valid UTF-8), but returns None for empty or binary files, and None.lower() would crash the loop without the surrounding try.
Encoding Comparison Table
Storage Space Comparison
text = "Hello 한글 😀"
encodings = {
'ASCII (English only)': 'ascii',
'UTF-8': 'utf-8',
'UTF-16 LE': 'utf-16-le',
'UTF-16 BE': 'utf-16-be',
'UTF-32 LE': 'utf-32-le',
'EUC-KR': 'euc-kr',
'CP949': 'cp949',
}
print(f"Original text: {text}\n")
print(f"{'Encoding':20s} | {'Bytes':6s} | Hex")
print("-" * 60)
for name, enc in encodings.items():
try:
encoded = text.encode(enc)
hex_str = encoded.hex()[:30] + ('...' if len(encoded) > 15 else '')
print(f"{name:20s} | {len(encoded):4d}B | {hex_str}")
except UnicodeEncodeError:
print(f"{name:20s} | {'N/A':6s} | (Cannot encode)")
# Output:
# Original text: Hello 한글 😀
#
# Encoding | Bytes | Hex
# ------------------------------------------------------------
# ASCII (English only) | N/A | (Cannot encode)
# UTF-8 | 17B | 48656c6c6f20ed959ceab88020f09f...
# UTF-16 LE | 22B | 480065006c006c006f0020005cd500...
# UTF-16 BE | 22B | 00480065006c006c006f0020d55cae...
# UTF-32 LE | 40B | 48000000650000006c0000006c0000...
# EUC-KR | N/A | (Cannot encode)
# CP949 | N/A | (Cannot encode)
Feature Comparison
| Encoding | Bytes/Char | ASCII Compatible | Korean Efficiency | Emoji | Main Usage |
|---|---|---|---|---|---|
| ASCII | 1 | ✅ | ❌ | ❌ | English only |
| EUC-KR | 1-2 | ✅ | ✅✅ | ❌ | Korean legacy |
| CP949 | 1-2 | ✅ | ✅✅ | ❌ | Windows Korean |
| UTF-8 | 1-4 | ✅ | ✅ | ✅ | Web, Linux, modern standard |
| UTF-16 | 2-4 | ❌ | ✅✅ | ✅ | Windows, Java internal |
| UTF-32 | 4 | ❌ | ❌ | ✅ | Internal processing |
Real-World Scenarios
Scenario 1: Legacy System Integration
# Problem: Bank API responds with EUC-KR
import requests
response = requests.get('http://legacy-bank-api.com/account')
# ❌ Auto-decode (assumes UTF-8)
# print(response.text) # Corrupted
# ✅ Correct handling
content = response.content # Bytes
text = content.decode('euc-kr')
print(text)
# ✅ Or provide hint to requests
response.encoding = 'euc-kr'
print(response.text)
Scenario 2: Multilingual Application
import locale
import sys
def setup_encoding():
"""Setup system encoding"""
# Check stdout encoding
print(f"stdout encoding: {sys.stdout.encoding}")
# System locale
print(f"System locale: {locale.getpreferredencoding()}")
# Force UTF-8 (Python 3.7+)
if sys.stdout.encoding != 'utf-8':
sys.stdout.reconfigure(encoding='utf-8')
# Handle multilingual text
texts = {
'en': "Hello",
'ko': "안녕하세요",
'ja': "こんにちは",
'zh': "你好",
'ar': "مرحبا",
'ru': "Здравствуйте",
'emoji': "👋🌍"
}
for lang, text in texts.items():
utf8 = text.encode('utf-8')
print(f"{lang:5s}: {text:15s} | {len(utf8):2d} bytes | {utf8.hex()[:30]}")
Scenario 3: File Upload Handling
from flask import Flask, request
import chardet
app = Flask(__name__)
@app.route('/upload', methods=['POST'])
def upload_file():
file = request.files['file']
# Read as binary
content = file.read()
# Detect encoding
detected = chardet.detect(content)
encoding = detected['encoding']
confidence = detected['confidence']
print(f"Detected: {encoding} ({confidence*100:.1f}%)")
# Convert to UTF-8
if encoding.lower() != 'utf-8':
try:
text = content.decode(encoding)
utf8_content = text.encode('utf-8')
return {
'status': 'converted',
'from': encoding,
'to': 'utf-8',
'content': text
}
except Exception as e:
return {'status': 'error', 'message': str(e)}, 400
return {
'status': 'ok',
'encoding': 'utf-8',
'content': content.decode('utf-8')
}
Declaring encodings in Python code and handling the BOM
Always pass encoding to open()
When you omit encoding, Python’s open() uses the locale encoding of the OS. On Linux and macOS that is usually UTF-8, so nothing looks wrong; on a Korean-language Windows machine the default is cp949, and the same code raises UnicodeDecodeError while reading a UTF-8 file. That is the usual cause of “it works on my machine but not on my colleague’s Windows laptop.” Running with python -X warn_default_encoding (3.10+) emits an EncodingWarning at every call site that omits the encoding, and setting PYTHONUTF8=1 switches on UTF-8 mode so the default becomes UTF-8. PEP 686 makes UTF-8 mode the default from Python 3.15, but until you can rely on that version, passing encoding='utf-8' explicitly is the only fix that works everywhere.
Error Handling
# ✅ Error handling strategy
def safe_decode(data, encodings=['utf-8', 'cp949', 'euc-kr', 'latin-1']):
"""Try multiple encodings"""
for enc in encodings:
try:
return data.decode(enc), enc
except UnicodeDecodeError:
continue
# If all fail, decode ignoring errors
return data.decode('utf-8', errors='replace'), 'utf-8'
# Usage
with open('unknown.txt', 'rb') as f:
data = f.read()
text, encoding = safe_decode(data)
print(f"Decoded as {encoding}: {text}")
The order of the list is doing all the work in this function, and it has two quirks worth knowing. euc-kr after cp949 never runs usefully, because anything EUC-KR can decode, CP949 already decoded. And latin-1 maps every byte to a character, so it never fails: the final errors='replace' line is unreachable, and any file that is neither UTF-8 nor CP949 comes back as Latin-1 mojibake labeled as success. If you would rather know that decoding failed, drop latin-1 from the list.
BOM Handling
# ✅ Reading: utf-8-sig strips a BOM if present, otherwise reads plain UTF-8
with open('file.txt', 'r', encoding='utf-8-sig') as f:
text = f.read()
# ✅ Ordinary files: save without a BOM
with open('file.txt', 'w', encoding='utf-8') as f:
f.write(text)
# ⚠️ Only when the consumer needs it, e.g. a CSV meant to be opened in Excel
with open('export.csv', 'w', encoding='utf-8-sig', newline='') as f:
f.write(text)
The reason to avoid a BOM by default is that tools unaware of it treat those three bytes as content: a BOM before #! stops a shell script from running, some JSON parsers fail on the first character, and PHP prints it, triggering “headers already sent.” Excel is the opposite case: without a BOM it opens a UTF-8 CSV using the system code page and Korean text comes out garbled. So the accurate rule is not “never write a BOM” but “write one only when the reader needs it.”
Problem Solving Checklist
When Korean Characters Are Corrupted
# 1. Check file encoding
import chardet
with open('file.txt', 'rb') as f:
result = chardet.detect(f.read())
print(result)
# 2. Read with correct encoding
with open('file.txt', 'r', encoding='cp949') as f:
text = f.read()
# 3. Re-save as UTF-8
with open('file.txt', 'w', encoding='utf-8') as f:
f.write(text)
When Korean Characters Are Corrupted on Web
# 1. Check HTTP header
import requests
response = requests.get('https://example.com')
print(response.encoding) # ISO-8859-1 (wrong guess)
# 2. Set correct encoding
response.encoding = 'utf-8'
print(response.text)
# 3. Check Content-Type header
print(response.headers.get('Content-Type'))
# text/html; charset=euc-kr
# 4. Explicit decoding
text = response.content.decode('euc-kr')
When Korean Characters Are Corrupted in Database
# 1. Check connection encoding
import pymysql
conn = pymysql.connect(
host='localhost',
user='user',
password='pass',
database='mydb',
charset='utf8mb4' # ✅ Explicit specification
)
# 2. Check table encoding
cursor = conn.cursor()
cursor.execute("SHOW CREATE TABLE users")
print(cursor.fetchone())
# 3. Convert encoding
# ALTER TABLE users CONVERT TO CHARACTER SET utf8mb4;
Choosing an encoding for new and legacy data
flowchart TD
Start[Start new project] --> Q1{Language?}
Q1 -->|English only| ASCII["ASCII\nor UTF-8"]
Q1 -->|Multilingual| UTF8["✅ UTF-8\nRecommended"]
Q1 -->|Legacy integration| Q2{System?}
Q2 -->|Windows Korean| CP949[CP949]
Q2 -->|Unix Korean| EUCKR[EUC-KR]
Q2 -->|Japanese| SJIS[Shift-JIS]
UTF8 --> Best["✅ Best choice\n- Web standard\n- All characters supported\n- ASCII compatible"]
For anything new, use UTF-8 everywhere: files, HTTP headers, database columns and connections. Legacy encodings belong only at the boundary where you read old data, and the goal is to convert them once and store the result as UTF-8.
The habit that destroys data is reaching for errors='replace' or errors='ignore' when decoding fails. It makes the exception go away by turning unknown bytes into � or dropping them, and the original text cannot be recovered afterwards. A decode error almost always means the data is in another encoding. For Korean text from Windows systems, try cp949 before euc-kr, because CP949 is a superset that covers syllables EUC-KR cannot represent. Keep replace for logs and diagnostics, never for data you store.
Debugging Tools
Python Encoding Debugger
def analyze_encoding(file_path):
"""Detailed file encoding analysis"""
with open(file_path, 'rb') as f:
raw_data = f.read()
print(f"📄 File: {file_path}")
print(f"📊 Size: {len(raw_data)} bytes\n")
# Check BOM
if raw_data.startswith(b'\xef\xbb\xbf'):
print("🔖 BOM: UTF-8")
elif raw_data.startswith(b'\xff\xfe'):
print("🔖 BOM: UTF-16 LE")
elif raw_data.startswith(b'\xfe\xff'):
print("🔖 BOM: UTF-16 BE")
else:
print("🔖 BOM: None")
# Detect encoding
detected = chardet.detect(raw_data)
print(f"\n🔍 Detected encoding: {detected['encoding']}")
print(f"📈 Confidence: {detected['confidence']*100:.1f}%")
# Try multiple encodings
print("\n🧪 Decoding test:")
encodings = ['utf-8', 'cp949', 'euc-kr', 'utf-16', 'latin-1']
for enc in encodings:
try:
text = raw_data.decode(enc)
preview = text[:50].replace('\n', '\\n')
print(f" ✅ {enc:10s}: {preview}")
except UnicodeDecodeError as e:
print(f" ❌ {enc:10s}: {e}")
# Hex dump (first 100 bytes)
print(f"\n🔢 Hex Dump (first 100 bytes):")
for i in range(0, min(100, len(raw_data)), 16):
hex_str = ' '.join(f'{b:02x}' for b in raw_data[i:i+16])
ascii_str = ''.join(chr(b) if 32 <= b < 127 else '.' for b in raw_data[i:i+16])
print(f" {i:04x}: {hex_str:48s} | {ascii_str}")
# Usage
analyze_encoding('mystery.txt')
References
- Unicode Standard
- UTF-8 Specification (RFC 3629)
- Character Encoding in Python
- The Absolute Minimum Every Software Developer Must Know About Unicode
Related Articles
- Linux and macOS Terminal Commands for Developers: GNU vs BSD Differences and Troubleshooting
- Debugging Guide: Common Errors in All Languages
- From HTTP/1.1 to HTTP/3