跳到主要内容

RFC 3629 - UTF-8, ISO 10646 的转换格式

  • 状态: Internet Standard
  • 发布日期: November 2003
  • Stream: IETF
  • 废弃了: RFC2279
  • 勘误: 无勘误

摘要​

ISO/IEC 10646-1 定义了一个称为 Universal Character Set (UCS) 的大型字符集, 它涵盖了世界上大多数书写系统. 然而, 最初提出的 UCS 编码与许多现有应用和协议并不兼容, 这促成了 UTF-8 的发展, 也就是本文档的主题. UTF-8 的特点是保留完整的 US-ASCII 范围, 从而兼容依赖 US-ASCII 值但对其他值透明的文件系统、解析器和其他软件. 本备忘录废弃并取代 RFC 2279.


本备忘录状态​

本文档为 Internet 社区规定一项 Internet 标准跟踪协议, 并请求讨论和改进建议. 请参考当前版本的 "Internet Official Protocol Standards" (STD 1), 了解本协议的标准化阶段和状态. 本备忘录可无限制分发.


版权声明​

Copyright (C) The Internet Society (2003). All Rights Reserved.


目录​


UTF-8 为什么重要?​

UTF-8 是现代 Internet 的标准字符编码. 几乎所有现代 Web 应用、API 和数据格式都使用 UTF-8.

核心优势​

特性描述重要性
ASCII 兼容ASCII 字符以相同字节编码极高
无字节序问题不存在 endian 差异极高
自同步可从任意位置恢复解码很高
空间效率较高英文 1 byte, CJK 通常 3 bytes很高
通用支持支持所有 Unicode 字符极高

UTF-8 编码规则速查​

编码表​

Unicode Range           Bytes  UTF-8 Byte Pattern
─────────────────────────────────────────────
U+0000 - U+007F 1 0xxxxxxx
U+0080 - U+07FF 2 110xxxxx 10xxxxxx
U+0800 - U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000 - U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

字符范围覆盖​

1 byte (ASCII):
- Latin letters, digits, basic punctuation
- Control characters
- Range: U+0000 - U+007F

2 bytes:
- Latin extensions
- Greek, Cyrillic, Arabic, Hebrew
- Range: U+0080 - U+07FF

3 bytes:
- CJK (Chinese, Japanese, Korean) characters
- Most other writing systems
- Range: U+0800 - U+FFFF

4 bytes:
- Emoji
- Historical scripts, rare CJK characters
- Range: U+10000 - U+10FFFF

编码示例​

ASCII 字符​

Character: 'A'
Unicode: U+0041
Binary: 0100 0001
UTF-8: 0x41
Bytes: 1

Encoding Process:
U+0041 < U+007F → use 1-byte template
0xxxxxxx → 01000001 → 0x41

汉字​

Character: '你'
Unicode: U+4F60
Binary: 0100 1111 0110 0000
UTF-8: 0xE4 0xBD 0xA0
Bytes: 3

Encoding Process:
U+4F60 in U+0800-U+FFFF range → use 3-byte template
1110xxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓
0100 111101 100000
↓ ↓ ↓
11100100 10111101 10100000
0xE4 0xBD 0xA0

Emoji​

Character: '😀'
Unicode: U+1F600
Binary: 0001 1111 0110 0000 0000 0000
UTF-8: 0xF0 0x9F 0x98 0x80
Bytes: 4

Encoding Process:
U+1F600 in U+10000-U+10FFFF range → use 4-byte template
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓ ↓
000011 111101 100000 000000
↓ ↓ ↓ ↓
11110000 10011111 10011000 10000000
0xF0 0x9F 0x98 0x80

UTF-8 的独特特性​

1. ASCII 兼容性​

ASCII file = Valid UTF-8 file

Example:
Hello World (ASCII)
is also valid UTF-8

Reason:
ASCII uses 7 bits (0xxxxxxx)
UTF-8's 1-byte form is ASCII

2. 自同步​

UTF-8 byte stream:
... E4 BD A0 E5 A5 BD ...
你 好

Starting from any position:
- First byte (1110xxxx or 110xxxxx or 11110xxx) marks character start
- Continuation bytes (10xxxxxx) never mistaken for first byte

Example:
E4 BD A0 E5 A5 BD
↑ ↑
Starting here identifies this as continuation byte
Starting here identifies new character

3. 无字节序问题​

UTF-16 needs BOM:
FE FF ... (Big Endian)
FF FE ... (Little Endian)

UTF-8 doesn't need it:
Byte order is fixed, high to low
No BOM needed to indicate byte order

常见应用​

Web 开发​

<!-- HTML file -->
<!DOCTYPE html>
<html>
<head>
<meta charset="UTF-8">
<title>UTF-8 Example</title>
</head>
<body>
<p>你好,世界! Hello, World! 😀</p>
</body>
</html>

HTTP 协议​

HTTP/1.1 200 OK
Content-Type: text/html; charset=UTF-8
Content-Length: 1234

<!DOCTYPE html>...

JSON 数据​

{
"name": "张三",
"message": "Hello 世界",
"emoji": "😀"
}

数据库​

CREATE DATABASE mydb CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;

安全考虑事项要点​

⚠️ 非最短形式攻击​

Prohibited overlong encodings:

Correct: 'A' → 0x41 (1 byte)
Wrong: 'A' → 0xC0 0x81 (2 bytes, overlong)
'A' → 0xE0 0x80 0x81 (3 bytes, overlong)

Danger:
Overlong encodings may bypass security checks
Example: path traversal "../" in overlong encoding

⚠️ 无效序列​

Must reject:
- Isolated continuation bytes (10xxxxxx)
- Beyond Unicode range (>U+10FFFF)
- UTF-16 surrogate pairs (U+D800-U+DFFF)
- Truncated multi-byte sequences

编程语言支持​

Python​

# Encoding
s = "你好世界"
b = s.encode('utf-8') # bytes object

# Decoding
s = b.decode('utf-8') # str object

# File operations
with open('file.txt', 'r', encoding='utf-8') as f:
content = f.read()

JavaScript​

// Encoding
const str = "你好世界";
const encoder = new TextEncoder();
const bytes = encoder.encode(str); // Uint8Array

// Decoding
const decoder = new TextDecoder('utf-8');
const text = decoder.decode(bytes); // string

Java​

// Encoding
String str = "你好世界";
byte[] bytes = str.getBytes(StandardCharsets.UTF_8);

// Decoding
String decoded = new String(bytes, StandardCharsets.UTF_8);

// File operations
Files.readString(path, StandardCharsets.UTF_8);

Go​

// Go's string is natively UTF-8
s := "你好世界"

// Convert to byte slice
b := []byte(s)

// Convert from byte slice
s = string(b)

性能特征​

空间效率对比​

文本类型UTF-8UTF-16UTF-32
英文1 byte2 bytes4 bytes
中文3 bytes2 bytes4 bytes
Emoji4 bytes4 bytes4 bytes
English-dominant text: UTF-8 optimal
CJK-dominant text: UTF-16 slightly better
Mixed text: UTF-8 usually optimal

相关资源​


快速诊断工具​

识别 UTF-8 编码​

def is_utf8(data):
"""Detect if data is valid UTF-8"""
try:
data.decode('utf-8')
return True
except UnicodeDecodeError:
return False

修复编码问题​

# Common problem: double encoding
# Original: "你好"
# Wrong display: "ä½ å¥½"

# Fix method:
text = "ä½ å¥½"
fixed = text.encode('latin1').decode('utf-8')
# Result: "你好"

重要提示: UTF-8 是现代 Internet 的默认标准. 请始终使用 UTF-8 编码, 避免使用 GBK、ISO-8859-1、Windows-1252 等遗留编码. 所有新项目都应将 UTF-8 作为唯一的字符编码.