Robert Helewka 7979e70705 feat(media): implement PJSUA2 audio capture port; document the media-plane refactor
Implements the tap half of the media path and records why the other half
requires moving call placement into PJSUA2.

MediaPipeline.create_tap was a stub: it logged "🎤 Audio tap created" and
returned a tap that nothing ever fed, so the classifier received no audio on
a live call. It now builds a real pj.AudioMediaPort subclass whose
onFrameReceived converts the SWIG ByteVector to PCM bytes and fans it out to
every tap on the stream.

One capture port per stream, shared by all taps: a second port on the same
stream would be mixed back into the conference bridge and the call would echo.

Thread safety is the constraint here. onFrameReceived runs on a PJSUA2 worker
thread — a third execution context beside the asyncio loop and the Sippy ED
thread — and touches nothing but AudioTap.feed, which hops to the owning loop
via call_soon_threadsafe. An exception escaping into PJSUA2's C++ callback
would tear down the worker thread and silently kill media for every call, so
the handler catches and logs once per port rather than on every 20ms frame.

Also fixes a hard crash found while testing this against real PJSUA2: a media
port finalised after Endpoint.libDestroy() calls pjmedia_conf_remove_port
against a freed conference bridge and aborts the process on a native
assertion. Ports are now released in remove_stream while the bridge still
exists, and stop() forces a collection before libDestroy — dropping the last
Python reference is not sufficient on its own.

Verified against the real bindings: frames fan out to multiple taps, cross the
thread boundary intact, and shutdown is clean.

add_remote_stream remains a stub, and deliberately so. PJSUA2 exposes no
standalone RTP media object — every AudioMedia subclass in the Python
bindings is a file player, recorder, tone generator or capture port, and RTP
is reachable only via pj.Call.getAudioMedia() on a dialog PJSUA2 itself owns.
A design where Sippy owns the dialog can never obtain media from PJSUA2, so
that function cannot be written against this API. docs/architecture.md now
explains this and records the resolution: PJSUA2 places the trunk call while
Sippy keeps the SBC roles (device registration, routing, leg bridging), with
the emergency guard and concurrency cap staying first in gateway.make_call
regardless of which library dials.

The architecture doc also had drift unrelated to media: it described the
thread boundary as asyncio.run_in_executor() when the real mechanism is
run_coroutine_threadsafe / ED2.callFromThread, claimed two execution contexts
where there are three, and cited MediaPipeline.add_stream() and
SippyEngine.bridge() — neither of which exists. Corrected, with the data flow
now showing the emergency guard and concurrency cap in their real positions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 06:56:50 -04:00
2026-03-21 19:21:33 +00:00

Hold Slayer 🔥

An AI-powered telephony gateway that calls companies, navigates IVR menus, waits on hold, and transfers you when a human picks up.

You give it a phone number and an intent ("dispute a charge on my December statement"). It dials the number through your SIP trunk, navigates the phone tree, sits through the hold music, and rings your desk phone the instant a live person answers. You never hear Vivaldi again.

Caution

Emergency calling — 911 Outbound calls to emergency numbers (911, 9911, 112) via the REST API or MCP tools are always refused — an AI agent must never place an emergency call, and API calls carry no E911 location data. Do not rely on this system as any part of your means of reaching emergency services; keep a phone with provider-registered E911 service available.

Architecture

┌─────────────────────────────────────────────────────────────────┐
│                        FastAPI Server                           │
│                                                                 │
│  ┌──────────┐  ┌──────────┐  ┌───────────┐  ┌──────────────┐  │
│  │ REST API │  │WebSocket │  │MCP Server │  │  Dashboard   │  │
│  │ /api/v1/*│  │ /ws/*    │  │ (HTTP)    │  │      /       │  │
│  └────┬─────┘  └────┬─────┘  └─────┬─────┘  └──────────────┘  │
│       │              │              │                            │
│  ┌────┴──────────────┴──────────────┴────┐                     │
│  │             Event Bus                  │                     │
│  │   (asyncio Queue pub/sub per client)   │                     │
│  └────┬──────────────┬──────────────┬────┘                     │
│       │              │              │                            │
│  ┌────┴─────┐  ┌─────┴─────┐  ┌────┴──────────┐               │
│  │   Call   │  │   Hold    │  │   Services    │               │
│  │ Manager  │  │  Slayer   │  │ (LLM, STT,   │               │
│  │          │  │           │  │  Recording,   │               │
│  │          │  │           │  │  Analytics,   │               │
│  │          │  │           │  │  Notify)      │               │
│  └────┬─────┘  └─────┬─────┘  └──────────────┘               │
│       │              │                                          │
│  ┌────┴──────────────┴───────────────────┐                     │
│  │         Sippy B2BUA Engine            │                     │
│  │  (SIP calls, DTMF, conference bridge) │                     │
│  └────┬──────────────────────────────────┘                     │
│       │                                                         │
└───────┼─────────────────────────────────────────────────────────┘
        │
   ┌────┴────┐
   │SIP Trunk│ ──→ PSTN
   └─────────┘

What's Implemented

Core Engine

  • Sippy B2BUA Engine (core/sippy_engine.py) — SIP call control, DTMF, bridging, conference, trunk registration
  • PJSUA2 Media Pipeline (core/media_pipeline.py) — Audio routing, recording ports, conference bridge, WAV playback (stub mode until the pjsua2 bindings are installed — see note below)
  • Call Manager (core/call_manager.py) — Active call state tracking, lifecycle management
  • Event Bus (core/event_bus.py) — Async pub/sub with per-subscriber queues, type filtering, history

Hold Slayer

  • IVR Navigation (services/hold_slayer.py) — Follows stored call flows step-by-step through phone menus, including SPEAK steps that synthesize speech via TTS
  • Audio Classifier (services/audio_classifier.py) — Real-time waveform analysis: silence, tones, DTMF, music, speech detection
  • Call Flow Learner (services/call_flow_learner.py) — Builds reusable call flows from exploration data, merges new discoveries
  • LLM Fallback — When a LISTEN step has no hardcoded DTMF, the LLM analyzes the transcript and picks the right menu option

AI Receptionist & Smart Routing

  • AI Receptionist (services/receptionist.py) — Answers inbound calls, greets via TTS, captures the caller's intent with STT + LLM, then routes to a device or takes a voicemail
  • Smart Routing (services/routing.py) — Caller-pattern (glob), DNIS, time-of-day (with tz + midnight wrap), per-device DND, and ring-chain priority. Rules win over the LLM on conflict.
  • TTS (services/tts.py) — Rhema (OpenAI-compatible /v1/audio/speech) — synthesizes Kokoro voices for the SPEAK step and receptionist prompts

Intelligence Layer

  • LLM Client (services/llm_client.py) — OpenAI-compatible API client (Ollama, vLLM, LM Studio, OpenAI) with JSON parsing, retry, stats
  • Transcription (services/transcription.py) — Speaches/Whisper STT integration for live call transcription
  • Recording (services/recording.py) — WAV recording with date-organized storage, dual-channel support, persisted to the recordings table
  • Call Persistence (services/call_persistence.py) — Writes completed calls + transcript chunks to the database on hangup
  • Notifications (services/notification.py) — WebSocket + SMS alerts for human detection, call failures, hold status

API Surface

  • REST API — Call management, call history, transcripts, recordings, routing rules, device DND, call flow CRUD
  • WebSocket — Real-time call events, transcripts, classification updates, receptionist state transitions
  • MCP Server — 15 tools + 3 resources for AI assistant integration (make calls, send DTMF, get transcripts, manage flows), served over streamable HTTP at /mcp/
  • Dashboard — SvelteKit UI served at / with live monitor, call history with transcript playback, and a routing-rules editor

Data Models

  • Call — Active call state with classification history, transcript chunks, hold time tracking
  • Call Flow — Stored IVR trees with steps (DTMF, LISTEN, HOLD, TRANSFER, SPEAK)
  • Routing Rule — Match (caller pattern, DNIS, time range) + action (ring_device, ring_chain, take_message, reject, dnd)
  • Transcript Chunk — Per-call STT segments with speaker tag and timestamp offset (for click-to-seek playback)
  • Recording — WAV file metadata (path, duration, size) per call
  • Events — 30+ typed events (call lifecycle, hold slayer, audio, device, system, receptionist, routing)
  • Device — SIP phone/softphone registration, priority, DND
  • Contact — Phone number management with routing preferences

Project Structure

hold-slayer/
├── main.py                      # FastAPI app + lifespan (service wiring)
├── config.py                    # Pydantic settings from .env
├── core/
│   ├── gateway.py               # Top-level gateway orchestrator
│   ├── sippy_engine.py          # Sippy B2BUA SIP engine
│   ├── media_pipeline.py        # PJSUA2 audio routing
│   ├── call_manager.py          # Active call state management
│   └── event_bus.py             # Async pub/sub event bus
├── services/
│   ├── hold_slayer.py           # IVR navigation + hold detection + SPEAK
│   ├── receptionist.py          # AI Receptionist state machine
│   ├── routing.py               # Smart routing (rules, DND, ring chain)
│   ├── tts.py                   # Rhema TTS client (OpenAI-compatible)
│   ├── audio_classifier.py      # Waveform analysis (music/speech/DTMF)
│   ├── call_flow_learner.py     # Auto-learns IVR trees from calls
│   ├── call_persistence.py      # Writes calls + transcript chunks on hangup
│   ├── llm_client.py            # OpenAI-compatible LLM client
│   ├── transcription.py         # Speaches/Whisper STT
│   ├── recording.py             # Call recording management
│   └── notification.py          # WebSocket + SMS notifications
├── api/
│   ├── calls.py                 # Call management endpoints
│   ├── call_history.py          # History, transcript, recording playback
│   ├── call_flows.py            # Call flow CRUD
│   ├── devices.py               # Device registration
│   ├── routing.py               # Routing rules CRUD + per-device DND
│   ├── websocket.py             # Real-time event stream
│   └── deps.py                  # FastAPI dependency injection
├── dashboard/                   # SvelteKit UI (built to dashboard/build)
│   └── src/routes/
│       ├── +page.svelte         # Live monitor
│       ├── history/             # Call history list
│       ├── calls/[call_id]/     # Detail page + transcript playback
│       └── routing/             # Rules editor + DND toggles
├── mcp_server/
│   └── server.py                # MCP tools + resources (15 tools)
├── models/
│   ├── call.py                  # Call state models
│   ├── call_flow.py             # IVR tree models
│   ├── routing.py               # Routing rule / match / action models
│   ├── events.py                # Event type definitions
│   ├── device.py                # Device models
│   └── contact.py               # Contact models
├── db/
│   └── database.py              # SQLAlchemy async (PostgreSQL + Alembic)
└── tests/
    ├── test_audio_classifier.py # 18 tests — waveform analysis
    ├── test_call_flows.py       # 10 tests — call flow models
    ├── test_hold_slayer.py      # 20 tests — IVR nav, EventBus, CallManager
    ├── test_services.py         # 27 tests — LLM, notifications, recording,
    │                            #             analytics, learner, EventBus
    ├── test_tts.py              #  4 tests — Rhema TTS client
    ├── test_routing.py          #  8 tests — rules evaluator
    └── test_receptionist.py     #  7 tests — receptionist decision logic

Quick Start

1. Install

python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

Note

The PJSUA2 media pipeline needs the pjsua2 Python bindings, which are not pip-installable — they're built from pjproject (./configure && make && make install with --enable-shared and the Python SWIG target). Without them the media layer runs in stub mode (signaling only): audio routing, recording and playback become no-ops that still return success.

See docs/pjsua2-build.md for the verified procedure (pjproject 2.17, no sudo required). It includes the patchelf RPATH step, without which the bindings compile and install but fail to import.

2. Configure

cp .env.example .env
# Edit .env with your SIP trunk credentials, LLM endpoint, etc.
# Required: DATABASE_URL, plus either the Casdoor SSO settings
# (CASDOOR_* + OWNER_NAME) or CASDOOR_ENABLED=false with HOST=127.0.0.1.

The gateway is owner-only. The browser dashboard signs in via Casdoor SSO (short-lived JWT); MCP and CLI clients use a Personal Access Token (hs_pat_…) minted from the dashboard's API Tokens menu. Both are presented as Authorization: Bearer <token> (WebSocket and <audio> recording downloads also accept ?token=…). Only the user whose Casdoor username matches OWNER_NAME may use any surface — everyone else gets 403. With CASDOOR_ENABLED=false the gateway runs in dev-owner mode, permitted only on a loopback bind.

cd dashboard
npm install
npm run build
cd ..

The gateway serves the built UI at / automatically when dashboard/build/ exists. Skip this step if you only need the REST/WS API.

4. Run

uvicorn main:app --host 0.0.0.0 --port 8000

5. Test

pytest tests/ -v

Docker

A single image bundles the FastAPI process and the built dashboard (the node stage compiles the SPA; pjsua2 is deliberately not built, so the media pipeline runs in stub mode — see the Dockerfile header). docker-compose.yaml brings up the app plus its own PostgreSQL:

cp .env.compose.example .env
# Fill in HS_DB_PASSWORD, and CASDOOR_CLIENT_ID/SECRET + OWNER_NAME.
docker compose up --build
# → http://localhost:21081

Because the published port binds the app to 0.0.0.0, the compose stack must run with Casdoor SSO enabled — dev-owner mode (CASDOOR_ENABLED=false) is loopback-only and is refused at startup here. Register a hold-slayer app in Casdoor (org heluca, redirect URI <PUBLIC_BASE_URL>/auth/callback) first. The image runs with USE_MOCK_SIP=true by default (a real trunk needs the SIP_TRUNK_* vars and USE_MOCK_SIP=false).

Usage

REST API

Launch Hold Slayer on a number:

curl -X POST http://localhost:8000/api/v1/calls/hold-slayer \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "number": "+18005551234",
    "intent": "dispute Amazon charge from December 15th",
    "call_flow_id": "chase_bank_main",
    "transfer_to": "sip_phone"
  }'

Check call status:

curl http://localhost:8000/api/v1/calls/call_abc123

Browse call history (persisted in the database):

curl http://localhost:8000/api/v1/calls/history?limit=50
curl http://localhost:8000/api/v1/calls/call_abc123/transcript
curl -O http://localhost:8000/api/v1/calls/call_abc123/recording   # WAV

Create a smart-routing rule:

curl -X POST http://localhost:8000/api/v1/routing/rules \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Block tollfree at night",
    "priority": 10,
    "enabled": true,
    "match": {
      "caller_pattern": "+1800*",
      "time_range": {"start": "22:00", "end": "06:00", "tz": "America/Toronto", "days": [0,1,2,3,4,5,6]}
    },
    "action": {"type": "reject", "message": "Office is closed."}
  }'

Toggle Do Not Disturb on a device:

curl -X PATCH http://localhost:8000/api/v1/routing/devices/dev_abc123/dnd \
  -H "Content-Type: application/json" \
  -d '{"enabled": true}'

WebSocket — Real-Time Events

const ws = new WebSocket(`ws://localhost:8000/ws/events?token=${token}`);
ws.onmessage = (msg) => {
  const event = JSON.parse(msg.data);
  // event.type: "human_detected", "hold_detected", "ivr_step", etc.
  // event.call_id: which call this is about
  // event.data: type-specific payload
};

MCP — AI Assistant Integration

The MCP server is served over streamable HTTP at /mcp/ (note the trailing slash) and authenticates with an owner-minted Personal Access Token (mint one from the dashboard's API Tokens menu — it starts with hs_pat_):

claude mcp add hold-slayer --transport http http://localhost:8000/mcp/ \
  --header "Authorization: Bearer hs_pat_..."

It exposes 15 tools and 3 resources (gateway://status, gateway://call-flows, gateway://active-calls):

Tool Description
make_call Dial a real number through the SIP trunk (emergency numbers refused)
hangup Hang up an active call
transfer_call Transfer an active call to a device
send_dtmf Send touch-tone digits to navigate menus
get_call_status Check current state of a call
get_call_transcript Get live transcript of a call
get_call_recording Get recording metadata and file path
list_active_calls List all calls in progress
list_devices List registered devices and status
gateway_status Trunk, devices, active calls, uptime
get_call_flow Look up a stored IVR flow for a number
create_call_flow Store a new IVR call flow
get_call_summary Stored summary and action items for a call
search_call_history Search past calls by number or intent
learn_call_flow Build/refine a reusable IVR flow from an exploration call

How It Works

Outbound (Hold Slayer)

  1. You request a call — via REST API, MCP tool, or dashboard
  2. Gateway dials out — Sippy B2BUA places the call through your SIP trunk
  3. Audio classifier listens — Real-time waveform analysis detects IVR prompts, hold music, ringing, silence, and live speech
  4. Transcription runs — Speaches/Whisper converts audio to text in real-time
  5. IVR navigator decides — If a stored call flow exists, it follows the steps (including SPEAK steps that synthesize speech via Rhema TTS). If not, the LLM analyzes the transcript and picks the right menu option
  6. Hold detection — When hold music is detected, the system waits patiently and monitors for transitions
  7. Human detection — The classifier detects the transition from music/silence to live speech
  8. Transfer — Your desk phone rings. Pick up and you're talking to the agent. Zero hold time.

Inbound (AI Receptionist + Smart Routing)

  1. SIP INVITE arrives — Sippy surfaces it to the gateway instead of auto-answering
  2. Routing rules evaluate — Caller pattern, DNIS, and time-of-day rules run in priority order. A reject or dnd action declines the call immediately.
  3. Receptionist answers — TTS plays the greeting; the call's audio tap captures the caller's response
  4. Intent capture — The utterance is transcribed and the LLM extracts intent, urgency, and a recommended action (ring / message / reject)
  5. Final decision — Routing rules win on conflict; otherwise the LLM's recommendation is followed
  6. Route or take a messagering_chain tries devices in priority order (skipping any in DND); if nobody picks up (or the action is take_message), the receptionist records up to 90s, transcribes it, and emits a RECEPTIONIST_MESSAGE_SAVED event

Configuration

All configuration is via environment variables (see .env.example):

Variable Description Default
DATABASE_URL PostgreSQL connection string — (required)
CASDOOR_ENABLED Enable Casdoor SSO (false → dev-owner, loopback only) false
CASDOOR_ENDPOINT Casdoor base URL https://id.ouranos.helu.ca
CASDOOR_CLIENT_ID Casdoor application client ID — (required if SSO on)
CASDOOR_CLIENT_SECRET Casdoor application client secret — (required if SSO on)
CASDOOR_ORG_NAME Casdoor organization heluca
CASDOOR_APP_NAME Casdoor application name
OWNER_NAME Casdoor username of the single operator (owner) — (required if SSO on)
PUBLIC_BASE_URL Public base URL for OAuth discovery (else derived from headers)
MAX_CONCURRENT_CALLS Cap on simultaneous outbound calls 4
SIP_TRUNK_HOST Your SIP provider hostname
SIP_TRUNK_USERNAME SIP auth username
SIP_TRUNK_PASSWORD SIP auth password
SIP_TRUNK_DID Your phone number (E.164)
GATEWAY_SIP_PORT Port for device registration 5080
SPEACHES_URL Speaches/Whisper STT endpoint http://localhost:22070
LLM_BASE_URL OpenAI-compatible LLM endpoint http://localhost:11434/v1
LLM_MODEL Model name for IVR analysis llama3
TTS_BASE_URL Rhema TTS endpoint (OpenAI-compatible) http://localhost:8000
TTS_MODEL TTS model ID speaches-ai/Kokoro-82M-v1.0-ONNX
TTS_VOICE Default Kokoro voice af_heart
TTS_API_KEY Optional bearer token for Rhema
RECEPTIONIST_ENABLED Answer inbound calls with the AI receptionist true
RECEPTIONIST_GREETING_TEMPLATE Spoken greeting "Hi, you've reached Robert's line. Who's calling, and what's this about?"
RECEPTIONIST_MESSAGE_MAX_SECONDS Voicemail cap 90

Tech Stack

  • Python 3.12+ + asyncio — Single-process async architecture
  • FastAPI — REST API + WebSocket server
  • SvelteKit — Dashboard UI (built static, served by FastAPI at /)
  • Sippy B2BUA — SIP call control and DTMF
  • PJSUA2 — Media pipeline, conference bridge, recording, WAV playback
  • Speaches (Whisper) — Speech-to-text
  • Rhema (Kokoro) — Text-to-speech (OpenAI-compatible /v1/audio/speech)
  • Ollama / vLLM / OpenAI — LLM for IVR menu analysis and receptionist intent capture
  • SQLAlchemy + Alembic — Async database (PostgreSQL; schema managed by migrations)
  • MCP (Model Context Protocol) — AI assistant integration

Documentation

Full documentation is in /docs:

  • Architecture — System design, data flow, threading model
  • Core Engine — SIP engine, media pipeline, call manager, event bus
  • Hold Slayer Service — IVR navigation, hold detection, human detection
  • Audio Classifier — Waveform analysis, feature extraction, classification
  • Services — LLM client, transcription, recording, analytics, notifications
  • Call Flows — Call flow model, step types, auto-learner
  • API Reference — REST endpoints, WebSocket, request/response schemas
  • MCP Server — MCP tools and resources for AI assistants
  • Configuration — All environment variables, deployment options
  • Development — Setup, testing, contributing

Build Phases

Phase 1: Core Engine

  • Extract EventBus to dedicated module with typed filtering
  • Implement Sippy B2BUA SIP engine (signaling, DTMF, bridging)
  • PJSUA2 media pipeline contract (conference bridge, audio tapping, recording) — runs in stub mode until pjsua2 bindings are installed
  • Call manager with active call state tracking
  • Gateway orchestrator wiring all components

Phase 2: Intelligence Layer

  • LLM client (OpenAI-compatible — Ollama, vLLM, LM Studio, OpenAI)
  • Hold Slayer IVR navigation with LLM fallback for LISTEN steps
  • Call Flow Learner — auto-builds reusable IVR trees from exploration
  • Recording service with date-organized WAV storage
  • Audio classifier with spectral analysis, DTMF detection, hold-to-human transition

Phase 3: API & Integration

  • REST API — calls, call flows, devices, DTMF
  • WebSocket real-time event streaming
  • MCP server with 15 tools + 3 resources, mounted at /mcp/ (streamable HTTP)
  • Notification service (WebSocket + SMS)
  • Service wiring in main.py lifespan

Phase 4: Production Hardening 🚧

  • Alembic database migrations (baseline + upgrade-on-boot)
  • API authentication — Casdoor SSO (browser JWT) + owner-minted PATs, owner-only across REST/WS/MCP
  • Emergency-number guard + concurrent-call cap on outbound calls
  • Rate limiting on API endpoints
  • Structured JSON logging
  • Honest /health — engine mode, DB ping, trunk registration, STT/TTS availability
  • Graceful degradation (classifier works without STT, etc.)
  • Docker Compose (Hold Slayer + PostgreSQL)

Phase 5: Additional Services 🚧

  • AI Receptionist — answer inbound calls, screen callers, take messages
  • Smart Routing — time-of-day rules, device priority, DND
  • TTS/Speech — play prompts into calls (SPEAK step support, Rhema/Kokoro)
  • Spam Filter — detect robocalls using caller ID + audio patterns
  • Noise Cancellation — RNNoise integration in media pipeline

Phase 6: Dashboard & UX 🚧

  • Web dashboard with real-time call monitor
  • Call history with transcript playback (click-to-seek)
  • Routing rules editor + per-device DND toggles
  • Call flow visual editor (drag-and-drop IVR tree builder)
  • Analytics dashboard with hold time graphs
  • Mobile app (or PWA) for on-the-go control

License

MIT

Description
No description provided
Readme MIT 1,008 KiB
Languages
Python 90.9%
Svelte 6.2%
TypeScript 2.2%
Dockerfile 0.5%
HTML 0.1%
Other 0.1%