RFC 3629 - UTF-8, un formato di trasformazione di ISO 10646
- Stato: Internet Standard
- Pubblicato: November 2003
- Stream: IETF
- Sostituisce: RFC2279
- Errata: Nessun errata
Riassunto (Abstract)
ISO/IEC 10646-1 definisce un grande insieme di caratteri chiamato Universal Character Set (UCS) che comprende la maggior parte dei sistemi di scrittura del mondo. Tuttavia, le codifiche originariamente proposte per UCS non erano compatibili con molte applicazioni e protocolli attuali, il che ha portato allo sviluppo di UTF-8, oggetto di questo memo. UTF-8 ha la caratteristica di preservare l'intera gamma US-ASCII, fornendo compatibilità con file system, parser e altro software che si basa su valori US-ASCII ma è trasparente ad altri valori. Questo memo rende obsoleta e sostituisce RFC 2279.
Stato di questo memo (Status of this Memo)
Questo documento specifica un protocollo dello standard Internet per la comunità Internet e richiede discussione e suggerimenti per miglioramenti. Si prega di fare riferimento all'edizione corrente degli "Internet Official Protocol Standards" (STD 1) per lo stato di standardizzazione e lo stato di questo protocollo. La distribuzione di questo memo è illimitata.
Avviso di copyright (Copyright Notice)
Copyright (C) The Internet Society (2003). Tutti i diritti riservati.
Sommario (Table of Contents)
Sezioni principali
- 1. Introduzione (Introduction)
- 2. Convenzioni di notazione (Notational conventions)
- 3. Definizione UTF-8 (UTF-8 definition)
- 4. Sintassi delle sequenze di byte UTF-8 (Syntax of UTF-8 Byte Sequences)
- 5. Versioni degli standard (Versions of the standards)
- 6. Marcatore ordine byte (Byte order mark - BOM)
- 7. Esempi (Examples)
- 8. Registrazione MIME (MIME registration)
- 9. Considerazioni IANA (IANA Considerations)
- 10. Considerazioni sulla sicurezza (Security Considerations)
- 11. Ringraziamenti (Acknowledgements)
- 12. Modifiche rispetto a RFC 2279 (Changes from RFC 2279)
- 13. Riferimenti normativi (Normative References)
- 14. Riferimenti informativi (Informative References)
Perché UTF-8 è importante
UTF-8 è la codifica di caratteri standard per l'Internet moderno. Quasi tutte le moderne applicazioni web, API e formati di dati utilizzano UTF-8.
Vantaggi principali
| Caratteristica | Descrizione | Importanza |
|---|---|---|
| Compatibile ASCII | I caratteri ASCII sono codificati in modo identico | ⭐⭐⭐⭐⭐ |
| Nessun problema di ordine byte | Nessun problema di endianness | ⭐⭐⭐⭐⭐ |
| Auto-sincronizzante | Può decodificare da qualsiasi posizione | ⭐⭐⭐⭐ |
| Efficiente nello spazio | 1 byte per inglese, 3 byte per CJK | ⭐⭐⭐⭐ |
| Supporto globale | Supporta tutti i caratteri Unicode | ⭐⭐⭐⭐⭐ |
Riferimento rapido delle regole di codifica UTF-8
Tabella di codifica
Intervallo Unicode Byte Pattern di byte UTF-8
─────────────────────────────────────────────
U+0000 - U+007F 1 0xxxxxxx
U+0080 - U+07FF 2 110xxxxx 10xxxxxx
U+0800 - U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000 - U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
Copertura degli intervalli di caratteri
1 byte (ASCII):
- Lettere latine, numeri, punteggiatura di base
- Caratteri di controllo
- Intervallo: U+0000 - U+007F
2 byte:
- Latino esteso
- Greco, cirillico, arabo, ebraico
- Intervallo: U+0080 - U+07FF
3 byte:
- Caratteri CJK
- La maggior parte degli altri sistemi di scrittura
- Intervallo: U+0800 - U+FFFF
4 byte:
- Emoji
- Scritture storiche, caratteri Han rari
- Intervallo: U+10000 - U+10FFFF
Esempi di codifica
Carattere ASCII
Character: 'A'
Unicode: U+0041
Binary: 0100 0001
UTF-8: 0x41
Bytes: 1
Encoding Process:
U+0041 < U+007F → use 1-byte template
0xxxxxxx → 01000001 → 0x41
Carattere cinese
Character: '你'
Unicode: U+4F60
Binary: 0100 1111 0110 0000
UTF-8: 0xE4 0xBD 0xA0
Bytes: 3
Encoding Process:
U+4F60 in U+0800-U+FFFF range → use 3-byte template
1110xxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓
0100 111101 100000
↓ ↓ ↓
11100100 10111101 10100000
0xE4 0xBD 0xA0
Emoji
Character: '😀'
Unicode: U+1F600
Binary: 0001 1111 0110 0000 0000 0000
UTF-8: 0xF0 0x9F 0x98 0x80
Bytes: 4
Encoding Process:
U+1F600 in U+10000-U+10FFFF range → use 4-byte template
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓ ↓
000011 111101 100000 000000
↓ ↓ ↓ ↓
11110000 10011111 10011000 10000000
0xF0 0x9F 0x98 0x80
Caratteristiche distintive di UTF-8
1. Compatibilità ASCII
ASCII file = Valid UTF-8 file
Example:
Hello World (ASCII)
is also valid UTF-8
Reason:
ASCII uses 7 bits (0xxxxxxx)
UTF-8's 1-byte form is ASCII
2. Auto-sincronizzazione
UTF-8 byte stream:
... E4 BD A0 E5 A5 BD ...
你 好
Starting from any position:
- First byte (1110xxxx or 110xxxxx or 11110xxx) marks character start
- Continuation bytes (10xxxxxx) never mistaken for first byte
Example:
E4 BD A0 E5 A5 BD
↑ ↑
Starting here identifies this as continuation byte
Starting here identifies new character
3. Nessun problema di ordine dei byte
UTF-16 needs BOM:
FE FF ... (Big Endian)
FF FE ... (Little Endian)
UTF-8 doesn't need it:
Byte order is fixed, high to low
No BOM needed to indicate byte order
Applicazioni comuni
Sviluppo web
<!-- HTML file -->
<!DOCTYPE html>
<html>
<head>
<meta charset="UTF-8">
<title>UTF-8 Example</title>
</head>
<body>
<p>你好,世界! Hello, World! 😀</p>
</body>
</html>
Protocollo HTTP
HTTP/1.1 200 OK
Content-Type: text/html; charset=UTF-8
Content-Length: 1234
<!DOCTYPE html>...
Dati JSON
{
"name": "张三",
"message": "Hello 世界",
"emoji": "😀"
}
Database
CREATE DATABASE mydb CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;
Punti salienti delle considerazioni sulla sicurezza
⚠️ Attacco basato sulla forma non minima
Prohibited overlong encodings:
Correct: 'A' → 0x41 (1 byte)
Wrong: 'A' → 0xC0 0x81 (2 bytes, overlong)
'A' → 0xE0 0x80 0x81 (3 bytes, overlong)
Danger:
Overlong encodings may bypass security checks
Example: path traversal "../" in overlong encoding
⚠️ Sequenze non valide
Must reject:
- Isolated continuation bytes (10xxxxxx)
- Beyond Unicode range (>U+10FFFF)
- UTF-16 surrogate pairs (U+D800-U+DFFF)
- Truncated multi-byte sequences
Supporto nei linguaggi di programmazione
Python
# Encoding
s = "你好世界"
b = s.encode('utf-8') # bytes object
# Decoding
s = b.decode('utf-8') # str object
# File operations
with open('file.txt', 'r', encoding='utf-8') as f:
content = f.read()
JavaScript
// Encoding
const str = "你好世界";
const encoder = new TextEncoder();
const bytes = encoder.encode(str); // Uint8Array
// Decoding
const decoder = new TextDecoder('utf-8');
const text = decoder.decode(bytes); // string
Java
// Encoding
String str = "你好世界";
byte[] bytes = str.getBytes(StandardCharsets.UTF_8);
// Decoding
String decoded = new String(bytes, StandardCharsets.UTF_8);
// File operations
Files.readString(path, StandardCharsets.UTF_8);
Go
// Go's string is natively UTF-8
s := "你好世界"
// Convert to byte slice
b := []byte(s)
// Convert from byte slice
s = string(b)
Caratteristiche prestazionali
Confronto di efficienza nello spazio
| Tipo di testo | UTF-8 | UTF-16 | UTF-32 |
|---|---|---|---|
| Inglese | 1 byte | 2 byte | 4 byte |
| Cinese | 3 byte | 2 byte | 4 byte |
| Emoji | 4 byte | 4 byte | 4 byte |
English-dominant text: UTF-8 optimal
CJK-dominant text: UTF-16 slightly better
Mixed text: UTF-8 usually optimal
Risorse correlate (Related Resources)
- Testo ufficiale: RFC 3629 (TXT)
- Pagina ufficiale: RFC 3629 DataTracker
- Standard: STD 63
- Sostituisce: RFC 2279
- Standard Unicode: Unicode.org
- ISO 10646: ISO/IEC 10646
Strumenti diagnostici rapidi
Identificare la codifica UTF-8
def is_utf8(data):
"""Detect if data is valid UTF-8"""
try:
data.decode('utf-8')
return True
except UnicodeDecodeError:
return False
Correggere i problemi di codifica
# Common problem: double encoding
# Original: "你好"
# Wrong display: "ä½ å¥½"
# Fix method:
text = "ä½ å¥½"
fixed = text.encode('latin1').decode('utf-8')
# Result: "你好"
Nota importante: UTF-8 è lo standard predefinito per l'Internet moderno. Utilizzare sempre la codifica UTF-8 ed evitare codifiche legacy come GBK, ISO-8859-1, Windows-1252, ecc. Tutti i nuovi progetti dovrebbero utilizzare UTF-8 come unica codifica dei caratteri.