Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Sequence

Sequence mode analyzes a string as an ordered list of code points, grouped into grapheme clusters. It is useful for examining combining marks, emoji sequences, and invisible characters.

Start in Sequence Mode

Pass text containing multiple code points:

sauva 'Á👩‍💻'

Use --text to prevent code point notation parsing:

sauva --text 41
sauva --text U+2192

Read text from a pipe or file:

printf 'Á👩‍💻' | sauva --text -
sauva --text - < input.txt

A single code point opens the Inspector. Sequence opens for multiple code points, even when they form a single visible character. Empty input is rejected.

See Command Line Options for the complete input rules.

Sequence and normalization demo

Code Points and Grapheme Clusters

A code point is one Unicode value, such as U+0041. A grapheme cluster groups code points into a unit that often corresponds to a user-perceived character. Sauva uses extended grapheme cluster boundaries as described in Unicode Text Segmentation.

For example, the string Á👩‍💻 contains five code points and two grapheme clusters:

ClusterCode points
ÁU+0041 LATIN CAPITAL LETTER A + U+0301 COMBINING ACUTE ACCENT
👩‍💻U+1F469 WOMAN + U+200D ZERO WIDTH JOINER + U+1F4BB PERSONAL COMPUTER

The accented A in the command above is U+0041 followed by U+0301.

Reading the Screen

Sequence

  • Each row shows the input position, code point, display representation, and Unicode name.
  • The left gutter numbers grapheme clusters and connects their member code points. A dot marks a cluster containing one code point.
  • The selected cluster is emphasized in the gutter.
  • The heading shows code point and grapheme counts, abbreviated as CP and GC in compact layouts.
  • When space is available, the selection panel shows the selected code point's properties and position within its cluster. The glyph image renders that individual code point.

Input order, repeated characters, spaces, and newlines are preserved. Labels and dotted circles make otherwise hard-to-see characters visible in the list.

Operations

KeysAction
j / k, Down / UpSelect a code point
g / GSelect the first / last code point
EnterOpen the selected code point in the Inspector
nCompare normalization forms
q, EscQuit

Sequence Inspector

Backspace in the Inspector returns to the sequence. From Sequence, press n to open Normalization.