RFC 3629 - UTF-8, ISO 10646 的转换格式
- 状态: Internet Standard
- 发布日期: November 2003
- Stream: IETF
- 废弃了: RFC2279
- 勘误: 无勘误
摘要
ISO/IEC 10646-1 定义了一个称为 Universal Character Set (UCS) 的大型字符集, 它涵盖了世界上大多数书写系统. 然而, 最初提出的 UCS 编码与许多现有应用和协议并不兼容, 这促成了 UTF-8 的发展, 也就是本文档的主题. UTF-8 的特点是保留完整的 US-ASCII 范围, 从而兼容依赖 US-ASCII 值但对其他值透明的文件系统、解析器和其他软件. 本备忘录废弃并取代 RFC 2279.
本备忘录状态
本文档为 Internet 社区规定一项 Internet 标准跟踪协议, 并请求讨论和改进建议. 请参考当前版本的 "Internet Official Protocol Standards" (STD 1), 了解本协议的标准化阶段和状态. 本备忘录可无限制分发.
版权声明
Copyright (C) The Internet Society (2003). All Rights Reserved.
目录
- 1. 引言
- 2. 记号约定
- 3. UTF-8 定义
- 4. UTF-8 字节序列语法
- 5. 标准版本 (Versions of the standards)
- 6. 字节顺序标记 (Byte order mark, BOM)
- 7. 示例 (Examples)
- 8. MIME 注册 (MIME registration)
- 9. IANA 考虑事项 (IANA Considerations)
- 10. 安全考虑事项 (Security Considerations)
- 11. 致谢
- 12. 相对于 RFC 2279 的变更
- 13. 规范性参考文献
- 14. 资料性参考文献
UTF-8 为什么重要?
UTF-8 是现代 Internet 的标准字符编码. 几乎所有现代 Web 应用、API 和数据格式都使用 UTF-8.
核心优势
| 特性 | 描述 | 重要性 |
|---|---|---|
| ASCII 兼容 | ASCII 字符以相同字节编码 | 极高 |
| 无字节序问题 | 不存在 endian 差异 | 极高 |
| 自同步 | 可从任意位置恢复解码 | 很高 |
| 空间效率较高 | 英文 1 byte, CJK 通常 3 bytes | 很高 |
| 通用支持 | 支持所有 Unicode 字符 | 极高 |
UTF-8 编码规则速查
编码表
Unicode Range Bytes UTF-8 Byte Pattern
─────────────────────────────────────────────
U+0000 - U+007F 1 0xxxxxxx
U+0080 - U+07FF 2 110xxxxx 10xxxxxx
U+0800 - U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000 - U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
字符范围覆盖
1 byte (ASCII):
- Latin letters, digits, basic punctuation
- Control characters
- Range: U+0000 - U+007F
2 bytes:
- Latin extensions
- Greek, Cyrillic, Arabic, Hebrew
- Range: U+0080 - U+07FF
3 bytes:
- CJK (Chinese, Japanese, Korean) characters
- Most other writing systems
- Range: U+0800 - U+FFFF
4 bytes:
- Emoji
- Historical scripts, rare CJK characters
- Range: U+10000 - U+10FFFF
编码示例
ASCII 字符
Character: 'A'
Unicode: U+0041
Binary: 0100 0001
UTF-8: 0x41
Bytes: 1
Encoding Process:
U+0041 < U+007F → use 1-byte template
0xxxxxxx → 01000001 → 0x41
汉字
Character: '你'
Unicode: U+4F60
Binary: 0100 1111 0110 0000
UTF-8: 0xE4 0xBD 0xA0
Bytes: 3
Encoding Process:
U+4F60 in U+0800-U+FFFF range → use 3-byte template
1110xxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓
0100 111101 100000
↓ ↓ ↓
11100100 10111101 10100000
0xE4 0xBD 0xA0
Emoji
Character: '😀'
Unicode: U+1F600
Binary: 0001 1111 0110 0000 0000 0000
UTF-8: 0xF0 0x9F 0x98 0x80
Bytes: 4
Encoding Process:
U+1F600 in U+10000-U+10FFFF range → use 4-byte template
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓ ↓
000011 111101 100000 000000
↓ ↓ ↓ ↓
11110000 10011111 10011000 10000000
0xF0 0x9F 0x98 0x80
UTF-8 的独特特性
1. ASCII 兼容性
ASCII file = Valid UTF-8 file
Example:
Hello World (ASCII)
is also valid UTF-8
Reason:
ASCII uses 7 bits (0xxxxxxx)
UTF-8's 1-byte form is ASCII
2. 自同步
UTF-8 byte stream:
... E4 BD A0 E5 A5 BD ...
你 好
Starting from any position:
- First byte (1110xxxx or 110xxxxx or 11110xxx) marks character start
- Continuation bytes (10xxxxxx) never mistaken for first byte
Example:
E4 BD A0 E5 A5 BD
↑ ↑
Starting here identifies this as continuation byte
Starting here identifies new character
3. 无字节序问题
UTF-16 needs BOM:
FE FF ... (Big Endian)
FF FE ... (Little Endian)
UTF-8 doesn't need it:
Byte order is fixed, high to low
No BOM needed to indicate byte order
常见应用
Web 开发
<!-- HTML file -->
<!DOCTYPE html>
<html>
<head>
<meta charset="UTF-8">
<title>UTF-8 Example</title>
</head>
<body>
<p>你好,世界! Hello, World! 😀</p>
</body>
</html>
HTTP 协议
HTTP/1.1 200 OK
Content-Type: text/html; charset=UTF-8
Content-Length: 1234
<!DOCTYPE html>...
JSON 数据
{
"name": "张三",
"message": "Hello 世界",
"emoji": "😀"
}
数据库
CREATE DATABASE mydb CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
安全考虑事项要点
⚠️ 非最短形式攻击
Prohibited overlong encodings:
Correct: 'A' → 0x41 (1 byte)
Wrong: 'A' → 0xC0 0x81 (2 bytes, overlong)
'A' → 0xE0 0x80 0x81 (3 bytes, overlong)
Danger:
Overlong encodings may bypass security checks
Example: path traversal "../" in overlong encoding
⚠️ 无效序列
Must reject:
- Isolated continuation bytes (10xxxxxx)
- Beyond Unicode range (>U+10FFFF)
- UTF-16 surrogate pairs (U+D800-U+DFFF)
- Truncated multi-byte sequences
编程语言支持
Python
# Encoding
s = "你好世界"
b = s.encode('utf-8') # bytes object
# Decoding
s = b.decode('utf-8') # str object
# File operations
with open('file.txt', 'r', encoding='utf-8') as f:
content = f.read()
JavaScript
// Encoding
const str = "你好世界";
const encoder = new TextEncoder();
const bytes = encoder.encode(str); // Uint8Array
// Decoding
const decoder = new TextDecoder('utf-8');
const text = decoder.decode(bytes); // string
Java
// Encoding
String str = "你好世界";
byte[] bytes = str.getBytes(StandardCharsets.UTF_8);
// Decoding
String decoded = new String(bytes, StandardCharsets.UTF_8);
// File operations
Files.readString(path, StandardCharsets.UTF_8);
Go
// Go's string is natively UTF-8
s := "你好世界"
// Convert to byte slice
b := []byte(s)
// Convert from byte slice
s = string(b)
性能特征
空间效率对比
| 文本类型 | UTF-8 | UTF-16 | UTF-32 |
|---|---|---|---|
| 英文 | 1 byte | 2 bytes | 4 bytes |
| 中文 | 3 bytes | 2 bytes | 4 bytes |
| Emoji | 4 bytes | 4 bytes | 4 bytes |
English-dominant text: UTF-8 optimal
CJK-dominant text: UTF-16 slightly better
Mixed text: UTF-8 usually optimal
相关资源
- 官方文本: RFC 3629 (TXT)
- 官方页面: RFC 3629 DataTracker
- 标准编号: STD 63
- 废弃了: RFC 2279
- Unicode 标准: Unicode.org
- ISO 10646: ISO/IEC 10646
快速诊断工具
识别 UTF-8 编码
def is_utf8(data):
"""Detect if data is valid UTF-8"""
try:
data.decode('utf-8')
return True
except UnicodeDecodeError:
return False
修复编码问题
# Common problem: double encoding
# Original: "你好"
# Wrong display: "ä½ å¥½"
# Fix method:
text = "ä½ å¥½"
fixed = text.encode('latin1').decode('utf-8')
# Result: "你好"
重要提示: UTF-8 是现代 Internet 的默认标准. 请始终使用 UTF-8 编码, 避免使用 GBK、ISO-8859-1、Windows-1252 等遗留编码. 所有新项目都应将 UTF-8 作为唯一的字符编码.