Passa al contenuto principale

RFC 3629 - UTF-8, un formato di trasformazione di ISO 10646

  • Stato: Internet Standard
  • Pubblicato: November 2003
  • Stream: IETF
  • Sostituisce: RFC2279
  • Errata: Nessun errata

Riassunto (Abstract)​

ISO/IEC 10646-1 definisce un grande insieme di caratteri chiamato Universal Character Set (UCS) che comprende la maggior parte dei sistemi di scrittura del mondo. Tuttavia, le codifiche originariamente proposte per UCS non erano compatibili con molte applicazioni e protocolli attuali, il che ha portato allo sviluppo di UTF-8, oggetto di questo memo. UTF-8 ha la caratteristica di preservare l'intera gamma US-ASCII, fornendo compatibilità con file system, parser e altro software che si basa su valori US-ASCII ma è trasparente ad altri valori. Questo memo rende obsoleta e sostituisce RFC 2279.


Stato di questo memo (Status of this Memo)​

Questo documento specifica un protocollo dello standard Internet per la comunità Internet e richiede discussione e suggerimenti per miglioramenti. Si prega di fare riferimento all'edizione corrente degli "Internet Official Protocol Standards" (STD 1) per lo stato di standardizzazione e lo stato di questo protocollo. La distribuzione di questo memo è illimitata.


Copyright (C) The Internet Society (2003). Tutti i diritti riservati.


Sommario (Table of Contents)​

Sezioni principali​


Perché UTF-8 è importante​

UTF-8 è la codifica di caratteri standard per l'Internet moderno. Quasi tutte le moderne applicazioni web, API e formati di dati utilizzano UTF-8.

Vantaggi principali​

CaratteristicaDescrizioneImportanza
Compatibile ASCIII caratteri ASCII sono codificati in modo identico⭐⭐⭐⭐⭐
Nessun problema di ordine byteNessun problema di endianness⭐⭐⭐⭐⭐
Auto-sincronizzantePuò decodificare da qualsiasi posizione⭐⭐⭐⭐
Efficiente nello spazio1 byte per inglese, 3 byte per CJK⭐⭐⭐⭐
Supporto globaleSupporta tutti i caratteri Unicode⭐⭐⭐⭐⭐

Riferimento rapido delle regole di codifica UTF-8​

Tabella di codifica​

Intervallo Unicode         Byte   Pattern di byte UTF-8
─────────────────────────────────────────────
U+0000 - U+007F 1 0xxxxxxx
U+0080 - U+07FF 2 110xxxxx 10xxxxxx
U+0800 - U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000 - U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

Copertura degli intervalli di caratteri​

1 byte (ASCII):
- Lettere latine, numeri, punteggiatura di base
- Caratteri di controllo
- Intervallo: U+0000 - U+007F

2 byte:
- Latino esteso
- Greco, cirillico, arabo, ebraico
- Intervallo: U+0080 - U+07FF

3 byte:
- Caratteri CJK
- La maggior parte degli altri sistemi di scrittura
- Intervallo: U+0800 - U+FFFF

4 byte:
- Emoji
- Scritture storiche, caratteri Han rari
- Intervallo: U+10000 - U+10FFFF

Esempi di codifica​

Carattere ASCII​

Character: 'A'
Unicode: U+0041
Binary: 0100 0001
UTF-8: 0x41
Bytes: 1

Encoding Process:
U+0041 < U+007F → use 1-byte template
0xxxxxxx → 01000001 → 0x41

Carattere cinese​

Character: '你'
Unicode: U+4F60
Binary: 0100 1111 0110 0000
UTF-8: 0xE4 0xBD 0xA0
Bytes: 3

Encoding Process:
U+4F60 in U+0800-U+FFFF range → use 3-byte template
1110xxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓
0100 111101 100000
↓ ↓ ↓
11100100 10111101 10100000
0xE4 0xBD 0xA0

Emoji​

Character: '😀'
Unicode: U+1F600
Binary: 0001 1111 0110 0000 0000 0000
UTF-8: 0xF0 0x9F 0x98 0x80
Bytes: 4

Encoding Process:
U+1F600 in U+10000-U+10FFFF range → use 4-byte template
11110xxx 10xxxxxx 10xxxxxx 10xxxxxx
↓ ↓ ↓ ↓
000011 111101 100000 000000
↓ ↓ ↓ ↓
11110000 10011111 10011000 10000000
0xF0 0x9F 0x98 0x80

Caratteristiche distintive di UTF-8​

1. Compatibilità ASCII​

ASCII file = Valid UTF-8 file

Example:
Hello World (ASCII)
is also valid UTF-8

Reason:
ASCII uses 7 bits (0xxxxxxx)
UTF-8's 1-byte form is ASCII

2. Auto-sincronizzazione​

UTF-8 byte stream:
... E4 BD A0 E5 A5 BD ...
你 好

Starting from any position:
- First byte (1110xxxx or 110xxxxx or 11110xxx) marks character start
- Continuation bytes (10xxxxxx) never mistaken for first byte

Example:
E4 BD A0 E5 A5 BD
↑ ↑
Starting here identifies this as continuation byte
Starting here identifies new character

3. Nessun problema di ordine dei byte​

UTF-16 needs BOM:
FE FF ... (Big Endian)
FF FE ... (Little Endian)

UTF-8 doesn't need it:
Byte order is fixed, high to low
No BOM needed to indicate byte order

Applicazioni comuni​

Sviluppo web​

<!-- HTML file -->
<!DOCTYPE html>
<html>
<head>
<meta charset="UTF-8">
<title>UTF-8 Example</title>
</head>
<body>
<p>你好,世界! Hello, World! 😀</p>
</body>
</html>

Protocollo HTTP​

HTTP/1.1 200 OK
Content-Type: text/html; charset=UTF-8
Content-Length: 1234

<!DOCTYPE html>...

Dati JSON​

{
"name": "张三",
"message": "Hello 世界",
"emoji": "😀"
}

Database​

CREATE DATABASE mydb CHARACTER SET utf8mb4 COLLATE utf8mb4_unicode_ci;

Punti salienti delle considerazioni sulla sicurezza​

⚠️ Attacco basato sulla forma non minima​

Prohibited overlong encodings:

Correct: 'A' → 0x41 (1 byte)
Wrong: 'A' → 0xC0 0x81 (2 bytes, overlong)
'A' → 0xE0 0x80 0x81 (3 bytes, overlong)

Danger:
Overlong encodings may bypass security checks
Example: path traversal "../" in overlong encoding

⚠️ Sequenze non valide​

Must reject:
- Isolated continuation bytes (10xxxxxx)
- Beyond Unicode range (>U+10FFFF)
- UTF-16 surrogate pairs (U+D800-U+DFFF)
- Truncated multi-byte sequences

Supporto nei linguaggi di programmazione​

Python​

# Encoding
s = "你好世界"
b = s.encode('utf-8') # bytes object

# Decoding
s = b.decode('utf-8') # str object

# File operations
with open('file.txt', 'r', encoding='utf-8') as f:
content = f.read()

JavaScript​

// Encoding
const str = "你好世界";
const encoder = new TextEncoder();
const bytes = encoder.encode(str); // Uint8Array

// Decoding
const decoder = new TextDecoder('utf-8');
const text = decoder.decode(bytes); // string

Java​

// Encoding
String str = "你好世界";
byte[] bytes = str.getBytes(StandardCharsets.UTF_8);

// Decoding
String decoded = new String(bytes, StandardCharsets.UTF_8);

// File operations
Files.readString(path, StandardCharsets.UTF_8);

Go​

// Go's string is natively UTF-8
s := "你好世界"

// Convert to byte slice
b := []byte(s)

// Convert from byte slice
s = string(b)

Caratteristiche prestazionali​

Confronto di efficienza nello spazio​

Tipo di testoUTF-8UTF-16UTF-32
Inglese1 byte2 byte4 byte
Cinese3 byte2 byte4 byte
Emoji4 byte4 byte4 byte
English-dominant text: UTF-8 optimal
CJK-dominant text: UTF-16 slightly better
Mixed text: UTF-8 usually optimal


Strumenti diagnostici rapidi​

Identificare la codifica UTF-8​

def is_utf8(data):
"""Detect if data is valid UTF-8"""
try:
data.decode('utf-8')
return True
except UnicodeDecodeError:
return False

Correggere i problemi di codifica​

# Common problem: double encoding
# Original: "你好"
# Wrong display: "ä½ å¥½"

# Fix method:
text = "ä½ å¥½"
fixed = text.encode('latin1').decode('utf-8')
# Result: "你好"

Nota importante: UTF-8 è lo standard predefinito per l'Internet moderno. Utilizzare sempre la codifica UTF-8 ed evitare codifiche legacy come GBK, ISO-8859-1, Windows-1252, ecc. Tutti i nuovi progetti dovrebbero utilizzare UTF-8 come unica codifica dei caratteri.