Skip to main content

UTF-8 select

AI-Generated Content
This page was generated with the assistance of AI and may contain inaccuracies. It is intended as a placeholder for future human verification. If you spot issues ahead of its initial review, please report them on GitHub!
Standard

UTF-8 select sequences switch the terminal between UTF-8 and legacy byte-oriented character encoding. These sequences originate from ISO 2022 (ECMA-35) code-switching mechanisms.

Syntax​

ESC % G       Select UTF-8 encoding
ESC % @ Return to ISO 8859-1 / default encoding
Formal syntax
UTF8_ON  = 0x1b, "%", "G" ;
UTF8_OFF = 0x1b, "%", "@" ;

Description​

ESC % G instructs the terminal to interpret subsequent input as UTF-8 encoded text. ESC % @ reverts to the default encoding, typically ISO 8859-1 (Latin-1).

These sequences are defined in the ISO 2022 framework for announcing code structures. In ISO 2022 terms, ESC % G announces a "UTF-8 Level 1" coding system.

Modern relevance​

Virtually all modern terminal emulators default to UTF-8 mode and many ignore these sequences entirely, since UTF-8 is now the universal standard. The sequences are primarily relevant for:

  • Legacy applications that need to explicitly switch a terminal into UTF-8 mode
  • Terminal emulators that support multiple encodings and allow runtime switching

XTerm is one of the few modern terminals that fully supports switching between UTF-8 and non-UTF-8 modes at runtime via these sequences.

Specifications​

SpecificationSection
XTerm ctlseqs—

Terminal support​

TerminalSupportVersionNotes
Terminal Emulators
Alacritty✗NoNo ESC % G charset selection found; UTF-8 is default and only encoding
Bobcat✓0.9.0SetCharset() and IsUtf8Mode() methods confirm UTF-8 support (Terminal.cpp:470)
contour✗NoESC % G (UTF-8 select) not implemented
foot✗NoNo handling for ESC % G found; foot only supports UTF-8 mode
Ghostty✗NoNo implementation found for ESC % G UTF-8 selection. Terminal uses UTF-8 by default but doesn't explicitly handle this sequence
iTerm2✓Yes
Kitty✗NoSequence recognized but ignored - Kitty is always UTF-8
Konsole✓Yes
mintty✓YesESC % G (and ESC % 8) at src/termout.c:2028-2030 enables UTF-8
mlterm✗NoNo ESC % G implementation found; UTF-8 handled via encoding system
PuTTY✓0.52ESC % G and ESC % @ for UTF-8 mode selection
Rio✗NoNo ESC % G handling found
rxvt-unicode✗NoNative UTF-8 support but no ESC % G/@ sequences
st✓0.8
terminology✗NoNot implemented
VT100✗No
VTE✗NoHandled by DOCS command which is not implemented
WezTerm✗NoNo evidence found for ESC % G sequence; WezTerm only supports UTF-8 (per performer.rs:460 comment)
Windows Terminal✓YesESC % G handled via DesignateCodingSystem() in OutputStateMachineEngine.cpp:295
xterm✓YesImplemented in charproc.c:6223. Controlled by wideChars resource and utf8_always setting
xterm.js✓3.5.0ESC % @ and ESC % G both select default charset
Multiplexers
cy✗NoESC % G not in EscDispatch
GNU Screen✗NoESC % G not implemented; UTF-8 set via encoding command only
tmux✗NoESC % G not needed; tmux uses UTF-8 internally
tuios✗No
Zellij✗No

See also​