The speech fixture classified as LIVE_HUMAN on its first 3s window and drifted
to MUSIC for every window after. I had validated only the first window and
reported the fixture as verified, which overstated it: any lab result resting
on that fixture — the hold-slayer scenarios above all — was proving less than
it appeared to.
The cause was one modelling error, not a tuning problem. `_detect_tonality`
looks for an autocorrelation peak above 0.5 in the 50-1000 Hz lag range, and
each syllable used a *constant* f0, which is perfectly periodic there. That
scored is_tonal=True, handing the music score a free 0.3 that speech could not
outrun — and the decision requires speech_score to strictly exceed
music_score, so ties went to music.
Real voices glide and jitter, so the periodicity never locks. The fundamental
now follows a per-syllable pitch contour (rise or fall, plus ~2% cycle-to-cycle
jitter), with the frequency integrated to phase rather than multiplied by t —
`2*pi*f*t` is only a chirp when f is the instantaneous rate, which it is not
once f0 itself moves. is_tonal is now False in every window.
Two smaller fixes fell out of that:
- Aspiration noise is high-passed rather than broadband. Flat noise puts
energy in every Goertzel bin, so the strongest DTMF row and column both
clear the detector's `total_power * 0.1` threshold and each syllable reads
as a keypress. A first-difference filter leaves the 697-1633 Hz bands
comparatively empty. The level is set for margin — spectral flatness lands
at ~0.46, mid-way through the 0.1-0.5 band, not on an edge.
- The music fixture gained two more harmonics and a recording-style noise
floor. Windows straddling a chord change had a momentarily sparse spectrum
and fell *below* the music score's 0.05 flatness floor, scoring as speech.
All three fixtures now classify correctly in 100% of windows (music 27/27,
speech 5/5, silence 2/2), and remain correct when the window is stepped by
half a window — a fixture that only works on aligned boundaries would still
be a trap in a live call, where the analysis window has no relationship to
where the audio began.
Confirmed on a real call through the lab: scenario 1003 now shows the whole
hold-slayer arc, speech -> sustained music -> speech, matching the dialplan.
tests/test_lab_fixtures.py guards this: it sweeps every window rather than
sampling the first, which is exactly what the original validation missed, and
checks the generator is byte-for-byte deterministic. It skips when the
fixtures have not been generated, since they are gitignored.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>