Pitch, Pitch Class, and Note Are Three Different Types
I've written before about eliminating primitive obsession with semantic types, using file paths and physical quantities as the examples. Those are good examples partly because the distinction is already intuitive: a path isn't a URL, and a temperature isn't a velocity. The compiler helps because the concepts were already distinct in everyone's head.
Music theory is a harder and better test. It's full of values that are the same integer, that a naive model would represent identically, and that mean genuinely different things. Building Semantics.Music made the case for the approach better than either of those did, and it also showed exactly where types stop helping.
Three Types Where Most Code Has One
Most code that touches music has a concept called "note" holding an integer, and that integer is doing at least three jobs.
A pitch class is one of the twelve names, octave-folded to 0 through 11. Every C is the same pitch class regardless of octave. This is the right type for asking whether a chord contains a third, because the answer doesn't depend on which octave anything is in.
A pitch is a concrete sounding frequency, modeled as a MIDI note number from 0 to 127, where 60 is middle C. Two pitches an octave apart share a pitch class and are not the same pitch. This is the right type for anything that will actually be played.
A note is a pitch with a rhythmic duration and a velocity. It's a sounding event, not a location on a keyboard. This is the right type for a sequencer, and the wrong type for harmonic analysis, which doesn't care how long anything lasted.
Collapsing those into one integer works right up until code asks a question at the wrong level. Transposing a melody up an octave should leave its harmony identical and its pitches all different. Asking whether two chords are the same voicing is a pitch question, and asking whether they're the same chord is a pitch-class question. With one type, both are the same call and one of them is silently wrong.
Spelling Carries Meaning
The part that convinced me the domain was worth modeling properly is enharmonic spelling.
C sharp and D flat are the same pitch. They are not the same note in the sense a musician means, because the spelling records what the note is doing harmonically. A raised fourth degree spelled as F sharp and a lowered fifth spelled as G flat sound identical and function differently, and a score that spells them wrong is harder to read even though it sounds the same.
So an accidental is modeled as a semitone offset from double flat through double sharp, carried alongside the letter rather than folded into the pitch. Double sharps exist in the type for the same reason they exist in notation: G double sharp is a real thing that appears in real music, and a model that normalizes it to A has thrown away the reason it was written that way.
This is the same argument as keeping an absolute path distinct from a relative one after both have been resolved. The resolved value is identical. The intent that produced it isn't, and the intent is what the next reader needs.
Direction Belongs in the Interval
An interval is signed semitones, positive ascending and negative descending.
That sounds obvious stated plainly, and it isn't how most quick implementations do it. The convenient version stores an unsigned distance and keeps direction somewhere else, or infers it from the order of the two pitches at the call site. Both work until an interval is stored, passed, or compared on its own, at which point the direction has been separated from the thing it describes.
Putting the sign in the value means a descending fifth and an ascending fourth stay distinguishable even though they land on the same pitch class. Inverting an interval is a negation rather than a lookup. And a function that takes an interval can't be handed one whose direction was left behind at the caller, because there's nowhere for it to have been left.
Where the Types Stop Helping
Key inference is where this stops being a type-system problem, and it's the part where the library refuses to guess.
Given a chord progression, what key is it in? Often there's no single correct answer. The same four chords can be a major key or its relative minor, and which one a listener hears depends on emphasis, on what came before, and sometimes on nothing that's present in the chords at all.
A type can't fix that. What it can do is refuse to pretend. So key inference returns all 24 candidates, the twelve tonics in major and natural minor, ranked best first, each paired with a fit score between zero and one. A score of one means every chord is diatonic with a matching triad quality.
The scoring is deliberately coarse and documented rather than clever: a chord scores one when its root is diatonic and its triad quality matches the diatonic triad at that degree, a half when the root is diatonic but the quality differs, and zero when the root isn't in the scale. Ties break toward a tonic-rooted last chord, then a tonic-rooted first chord, then major.
That tie-breaking rule is a musical opinion sitting in code, and writing it down as a rule rather than burying it in a comparison is the only way to ship it responsibly. Somebody will disagree with it, and when they do they'll be able to find it.
My opinion is that this is the right approach for any domain with genuine ambiguity. The type system's job isn't to make the ambiguous case correct. It's to make the ambiguity visible in the return type, so a caller can't accidentally treat a guess as a fact. A method returning a Key would be lying. A method returning ranked candidates with scores isn't.
Analysis as a Separate Layer
Harmony, cadence, and form ended up as separate concerns layered above the note types rather than mixed into them.
A cadence is classified by scale-degree motion into the final chord: authentic is five to one, plagal is four to one, half is anything arriving on five. That definition is about function, so it belongs with harmonic function and scale degrees rather than with pitches. The same three chords are a different cadence in a different key, which means a cadence can't be a property of the notes.
Form sits a layer above that again, with sections and section types and named forms. A twelve-bar blues is defined by harmonic function rather than by pitches, and a model that ties it to specific chords can't recognize the same form in a different key.
Splitting progression analysis into parsing, harmonic analysis, cadence detection, chromatic analysis, and key inference kept each of those to one job. Parsing is the biggest file at 175 lines, which is unremarkable, because a parser is long in proportion to the grammar it accepts and says nothing about how hard the problem is.
The informative one is key inference, at 106 lines and the largest of the four analysis passes, against 78 for chromatic analysis, 47 for harmonic analysis and 44 for cadence detection. It is the part with the least certainty and therefore the most rules. Cadence detection is short because a cadence is a pattern that either matches or does not. Key inference is long because the answer is a weighing of evidence, and every line is another thing that counts as evidence.
Practical Takeaways
- When one integer answers questions at several conceptual levels, that's the signal for separate types. Octave-folded, concrete, and sounding are three different questions about a note.
- Keep the spelling when the spelling carries intent, even when two spellings normalize to the same value. The normalized form is what gets played, not what gets read.
- Put direction, sign, and units inside the value rather than alongside it, so they can't be separated by being passed somewhere.
- For a genuinely ambiguous question, return ranked candidates with scores rather than an answer. A type system can't resolve the ambiguity, but it can stop a caller mistaking a guess for a fact.
- Write the tie-breaking rule down as a rule. Every domain opinion buried in a comparison function is an opinion nobody can find later to argue with.