PJSUA2 media plane — audio finally reaches the classifier #8
Merged
r
merged 5 commits from 2026-07-29 21:34:10 +00:00
feat/pjsua2-media-tap into main
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| 3150f78552 |
fix(lab): make Asterisk logs usable — drop ANSI codes and healthcheck noise
Logging was configured correctly and shipping to Loki, but the stream was useless: measured at 100% healthcheck chatter, with every line wrapped in ANSI escape codes. The real SIP events were there and completely buried. Three causes, each masking the next: - asterisk.conf was written in the first lab commit with `nocolor = yes` and never mounted, so the setting had no effect. Now mounted. It was also overriding [directories] and runuser/rungroup, which the image sets up correctly itself — removed, since overriding them risks breaking the container for no gain. - The image's command is `-vvvdddf`: verbosity 3 and debug 3 forced on the command line, which overrides both asterisk.conf and logger.conf. Overridden in compose to drop -v and -d; warnings and errors still log, and verbosity is raisable at runtime when tracing a call. - The actual source: the image's healthcheck makes ~7 separate `asterisk -rx` connections every 30s, and Asterisk logs a connect/disconnect pair for each. Replaced with a single check on a 60s interval, and the check now runs `pjsip show transports` rather than `core show version` — that fails when Asterisk is up but unconfigured, which is exactly the state that produced a "healthy" container with no SIP stack on first deploy. logger.conf drops both `notice` and `verbose`, which is where those pairs arrive. Verified on galatea: noise down from ~48 to 8 lines per two minutes (-83%), zero ANSI codes in Loki, 90% of the stream now signal, container still healthy, transport and dialplan intact. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| c516f659cc |
feat(lab): add softphone endpoint so device registration can be tested
Asterisk is the registrar for devices, not Hold Slayer. A softphone REGISTERs to the lab and the gateway transfers a live call to it by dialling extension 2001. This is deliberate: Hold Slayer's own SIP listener answers 200 OK to any REGISTER with no digest challenge, so anything on the network could register as a device and receive transferred calls. Keeping registration in Asterisk means the lab does not exercise or depend on that path, and the device is authenticated. The pjsua CLI built alongside the Python bindings is the test device — same library stack as the gateway, so no new dependency. Verified end to end: a gateway call to 2001 produces two channels Up under one bridge id. Three things that cost time and are now written down: - `--realm=asterisk`, not `--realm='*'`: the wildcard fails against Asterisk's digest challenge with PJSIP_EFAILEDCREDENTIAL. - pjsua is an interactive console app and exits ~8s after start if stdin is closed or /dev/null. `script -qfc` and `setsid </dev/null` both appear to work — registration succeeds — and then the process dies, leaving a stale contact in Asterisk that routes INVITEs to a port nobody is listening on. Hold a fifo open on stdin instead, and verify the port is actually bound rather than trusting `pjsip show contacts`. - Qualify is off for this AOR: the pjsua console does not answer OPTIONS, so polling marks a working softphone Unavail and the dialplan refuses to ring it. The 2001 guard therefore tests PJSIP_AOR(softphone,contact) rather than DEVICE_STATE. A real hardphone answers OPTIONS and can have it re-enabled. The identify block now matches source address *and port*. A host-only match claims every packet from that address, so a co-located softphone's REGISTER was attributed to the gateway endpoint and checked against the gateway's password — surfacing as "Failed to authenticate" on a correct password. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 92c45e9c4d |
fix(lab): make the speech fixture actually classify as speech
The speech fixture classified as LIVE_HUMAN on its first 3s window and drifted to MUSIC for every window after. I had validated only the first window and reported the fixture as verified, which overstated it: any lab result resting on that fixture — the hold-slayer scenarios above all — was proving less than it appeared to. The cause was one modelling error, not a tuning problem. `_detect_tonality` looks for an autocorrelation peak above 0.5 in the 50-1000 Hz lag range, and each syllable used a *constant* f0, which is perfectly periodic there. That scored is_tonal=True, handing the music score a free 0.3 that speech could not outrun — and the decision requires speech_score to strictly exceed music_score, so ties went to music. Real voices glide and jitter, so the periodicity never locks. The fundamental now follows a per-syllable pitch contour (rise or fall, plus ~2% cycle-to-cycle jitter), with the frequency integrated to phase rather than multiplied by t — `2*pi*f*t` is only a chirp when f is the instantaneous rate, which it is not once f0 itself moves. is_tonal is now False in every window. Two smaller fixes fell out of that: - Aspiration noise is high-passed rather than broadband. Flat noise puts energy in every Goertzel bin, so the strongest DTMF row and column both clear the detector's `total_power * 0.1` threshold and each syllable reads as a keypress. A first-difference filter leaves the 697-1633 Hz bands comparatively empty. The level is set for margin — spectral flatness lands at ~0.46, mid-way through the 0.1-0.5 band, not on an edge. - The music fixture gained two more harmonics and a recording-style noise floor. Windows straddling a chord change had a momentarily sparse spectrum and fell *below* the music score's 0.05 flatness floor, scoring as speech. All three fixtures now classify correctly in 100% of windows (music 27/27, speech 5/5, silence 2/2), and remain correct when the window is stepped by half a window — a fixture that only works on aligned boundaries would still be a trap in a live call, where the analysis window has no relationship to where the audio began. Confirmed on a real call through the lab: scenario 1003 now shows the whole hold-slayer arc, speech -> sustained music -> speech, matching the dialplan. tests/test_lab_fixtures.py guards this: it sweeps every window rather than sampling the first, which is exactly what the original validation missed, and checks the generator is byte-for-byte deterministic. It skips when the fixtures have not been generated, since they are gitignored. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 2a05be27bf |
feat(sip): add PJSUA2 engine — audio finally reaches the classifier
The gateway can now hear. Verified end to end against the Asterisk lab: speech classifies as live_human, hold music as music, the speech→music transition tracks the dialplan, and DTMF reaches a real IVR (Asterisk logged "caller pressed 1 -> accounts" and branched). RTP stats show 0% packet loss. PJSUAEngine implements the existing SIPEngine interface, so the gateway, call manager and hold-slayer service are unchanged. It is selected with SIP_ENGINE=pjsua2; the default stays "sippy" while this is proven, and the mock remains opt-in as before. Why a new engine rather than fixing the old path: PJSUA2 exposes no standalone RTP media object, so it will not surface media for a dialog it does not own. Owning the dialog is the price of owning the media. Sippy keeps the SBC roles it is good at — device registration, routing, leg bridging — and SippyEngine remains fully functional for signalling; it simply cannot carry media, which its media branch now says plainly instead of calling a method that could never work. The safety invariants are untouched. This engine is reachable only through gateway.make_call, which refuses emergency numbers and enforces the concurrency cap before any SIP action. No new dial path was introduced. MediaPipeline.add_remote_stream(host, port) is replaced by attach_call_media(stream_id, audio_media), called from onCallMediaState — the one place PJSUA2 hands out RTP-backed media. Taps requested before media comes up are attached when it does, so the classifier never misses the start of a call. Three crash/lifetime bugs found by running against the real bindings, none of which any unit test would have caught: - pj.Call and pj.Account objects finalised after libDestroy() abort the process on a native assertion, exactly as media ports do. Both are now dropped and collected before the pipeline destroys the endpoint. - PJSUA2 keeps delivering callbacks during interpreter teardown, when module globals may already be cleared. The callbacks alias what they need locally and swallow everything: a raise there escapes into C++ and takes the worker thread with it. - hangup() only queues the BYE, so shutdown deleted the account with a call still active and left the far end on an unclosed dialog. stop() now waits briefly for the teardown to complete. Threading follows the established rule: PJSUA2 worker threads reach the loop only through _post_from_pj → run_coroutine_threadsafe, and any thread PJSUA2 did not create registers itself before touching a PJSUA2 object. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
|||
| 7979e70705 |
feat(media): implement PJSUA2 audio capture port; document the media-plane refactor
Implements the tap half of the media path and records why the other half requires moving call placement into PJSUA2. MediaPipeline.create_tap was a stub: it logged "🎤 Audio tap created" and returned a tap that nothing ever fed, so the classifier received no audio on a live call. It now builds a real pj.AudioMediaPort subclass whose onFrameReceived converts the SWIG ByteVector to PCM bytes and fans it out to every tap on the stream. One capture port per stream, shared by all taps: a second port on the same stream would be mixed back into the conference bridge and the call would echo. Thread safety is the constraint here. onFrameReceived runs on a PJSUA2 worker thread — a third execution context beside the asyncio loop and the Sippy ED thread — and touches nothing but AudioTap.feed, which hops to the owning loop via call_soon_threadsafe. An exception escaping into PJSUA2's C++ callback would tear down the worker thread and silently kill media for every call, so the handler catches and logs once per port rather than on every 20ms frame. Also fixes a hard crash found while testing this against real PJSUA2: a media port finalised after Endpoint.libDestroy() calls pjmedia_conf_remove_port against a freed conference bridge and aborts the process on a native assertion. Ports are now released in remove_stream while the bridge still exists, and stop() forces a collection before libDestroy — dropping the last Python reference is not sufficient on its own. Verified against the real bindings: frames fan out to multiple taps, cross the thread boundary intact, and shutdown is clean. add_remote_stream remains a stub, and deliberately so. PJSUA2 exposes no standalone RTP media object — every AudioMedia subclass in the Python bindings is a file player, recorder, tone generator or capture port, and RTP is reachable only via pj.Call.getAudioMedia() on a dialog PJSUA2 itself owns. A design where Sippy owns the dialog can never obtain media from PJSUA2, so that function cannot be written against this API. docs/architecture.md now explains this and records the resolution: PJSUA2 places the trunk call while Sippy keeps the SBC roles (device registration, routing, leg bridging), with the emergency guard and concurrency cap staying first in gateway.make_call regardless of which library dials. The architecture doc also had drift unrelated to media: it described the thread boundary as asyncio.run_in_executor() when the real mechanism is run_coroutine_threadsafe / ED2.callFromThread, claimed two execution contexts where there are three, and cited MediaPipeline.add_stream() and SippyEngine.bridge() — neither of which exists. Corrected, with the data flow now showing the emergency guard and concurrency cap in their real positions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |