ZuluSpell

The science of spelling

Letters that sound alike

What the literature suggests, what a recognizer measured, and where the two differ.

Letter confusion matrices

1. Expected confusions, from the literature

Each row is the letter that was said, and each column is the letter a listener might hear instead. Darker squares mean the pair is more likely to be mixed up. The shading here is qualitative — it comes from the confusion classes in the speech-recognition literature: the E-set (B, C, D, E, G, P, T, V, Z, whose names all end in the "ee" sound), the A-set (A, J, K, "ay"), the nasals M/N and the fricatives F/S. Pairs within a class share colour strength by class, not by measurement; the diagonal (letter heard correctly) is grey.

said↓ABCDEFGHIJKLMNOPQRSTUVWXYZ
A
B
C
D
E
F
G
H
I
J
K
L
M
N
O
P
Q
R
S
T
U
V
W
X
Y
Z

Point at a square to see the rate.

2. Measured confusions, from ISOLET

Same grid, but with measured rates. Each square shows how often that pair was confused, out of about 300 recordings of each letter. The data comes from ISOLET, a public set of 150 speakers saying every letter twice. A standard classifier was trained on it and always tested on speakers it had not heard before. It got 96.6% right overall, close to the 96% Fanty & Cole reported. N heard as M is the worst pair at 8.7%, then V→B, T→P, M→N and P→B. These are microphone recordings, so a phone line will do worse. Colour reaches full strength at 8.0%.

said↓ABCDEFGHIJKLMNOPQRSTUVWXYZ
A
B
C
D
E
F
G
H
I
J
K
L
M
N
O
P
Q
R
S
T
U
V
W
X
Y
Z

Point at a square to see the rate.

Data: Cole, Muthusamy & Fanty, ISOLET spoken letter database (UCI repository); Fanty & Cole, “Spoken Letter Recognition”, NIPS 1990 (paper). How this was measured, with a notebook →

Sources

The grouping above is taken from the speech-recognition literature on letters confusable over the telephone, not invented here.

  1. Munich & Lin, “Explicit Modelling of Common Acoustic Features for Character Recognition”, EUSIPCO 2004 — names the E-set (B, C, D, E, G, P, T, V, Z) sharing /iy/, the A-set (A, J, K) sharing /ey/, and the M/N and F/S pairs. eurasip.org
  2. Fanty, Cole & Roginski, “English Alphabet Recognition with Telephone Speech”, NIPS 1991 — human listeners on real telephone recordings; the confusions listed are BID/PIT/VIZ and F/S. papers.nips.cc
  3. Mitchell & Setlur, “Improved Spelling Recognition Using a Tree-Based Fast Lexical Match”, ICASSP 1999 — builds its search tree from confusion classes such as the e-set. congres.cran.univ-lorraine.fr

The classes those papers name, and the letters in them:

The two narrower modes differ only in how far down that list they go. “Very confusable” keeps the B, D, E, F, G, M, N, P, S, T, V, Z— the letters measured at 2.7% wrong or worse in the recordings. “Confusable” adds the rest of the classes just listed. “All letters” needs no classification at all: the phonetic alphabet was built to cover every letter.