How the measured matrix was made
The second grid on the confusion matrices page is not copied from a paper. I trained a letter recognizer on the public ISOLET recordings and counted its mistakes.
- Data
- ISOLET (Cole, Muthusamy & Fanty, OGI). 150 speakers, each letter said twice: 7,797 recordings (3 missing from the original).
- Input features
- The 617 features that ship with ISOLET, computed by its authors: spectral coefficients, contour, sonorant, pre-sonorant and post-sonorant features. No audio processing of my own.
- Recognizer
- Support vector machine, RBF kernel, C = 10, gamma = 'scale' (scikit-learn SVC), after standardising every feature to zero mean and unit variance. 26 classes, one-vs-one.
- Tuning
- None. C = 10 was picked once by hand; no search was run, so there is no tuning leak into the test folds.
- Evaluation
- 5-fold cross-validation, one fold per ISOLET speaker group of 30 speakers. Every recording is predicted by a model that never heard that speaker.
- Result
- 96.6% correct over all 7,797 recordings. Fanty & Cole (NIPS 1990) reported 96% with a neural network on the same task.
- Matrix
- Row = letter spoken, column = letter recognised. Each square is the count divided by that row's total (about 300).
Limits
- ISOLET is clean microphone speech. A telephone line cuts high frequencies, so F/S and the E-set will be worse than shown.
- This is one machine recognizer, not human listeners and not a modern dictation system. Its errors follow the same vowel groups people report, but the exact rates are specific to it.
- The first data file has groups 1–4 joined together. I split it into four equal blocks, and because a few recordings are missing, a handful of rows near the borders may sit in the neighbouring fold.
Run it yourself
The notebook downloads ISOLET, trains the same recognizer, and draws the matrix. It needs Python with numpy, scikit-learn and matplotlib, and runs in a few minutes on a laptop.
Download the Jupyter notebook