docs: introduce Mercury Notebook Deliverable Pattern

This commit is contained in:
2026-07-08 13:43:12 -04:00
parent a991879061
commit c3260ae7b8
55 changed files with 12036 additions and 0 deletions

View File

@@ -0,0 +1,198 @@
The Genesys ROI documents claim 5 AI feature benefit categories:
Agent Copilot (voice + digital handle time + ACW)
Email AI (Auto-Respond + Auto-Suggest)
Speech & Text Analytics
Supervisor Copilot (AI Translate, AI Summary, Admin)
Predictive Routing
None of these are turn-key. Each requires configuration, tuning, and enablement effort. The original case has zero implementation cost.
Framework — four LoE dimensions per feature
Every AI feature carries four kinds of effort:
Dimension What it is Scales with
Fixed setup One-time base configuration — instance creation, settings, rules, permissions, security Roughly constant per feature
Variable configuration Per-scope effort — per queue, per language, per intent, per wrap-up code, per KB article Multipliers × unit count
Iterative tuning Test → measure → adjust cycles. Non-negotiable for AI features. Typical: 3-6 cycles before production stability Complexity of the feature
Enablement & change Agent/supervisor training, adoption support, communications, super-user network Headcount + geography
Plus a steady-state annual line that everyone forgets:
Steady-state What it is
Annual optimization Retraining, KB refresh, drift correction, model tuning as customer behavior shifts
Per-feature LoE — Genesys-claimed feature set
Below is my working LoE structure for the calculator. Hours are for a typical medium-complexity implementation. CTM-specific amplifiers follow in the next section.
1. Agent Copilot
Per the Genesys documentation you shared, the setup dimensions are: create the Copilot instance, configure settings, configure NLU (intents), configure rules, configure queues, configure per-language variants, wrap-up code configuration, AI Studio for custom summaries, testing, permissions, KB integration for answer highlighting.
Activity Unit Hours per unit Notes
Base Copilot instance setup Fixed 80-120 Per language variant (one instance per language)
Settings, rules, permissions config Fixed 40-60
NLU / intent modeling Per 10 intents 30-50 Includes utterance generation, training, validation
Wrap-up code mapping Per 20 wrap-ups 8-12 Includes utterance training per code
Queue configuration Per queue 1-3 Critical CTM scaling factor
Custom summary templates (AI Studio) Per template 20-40 If custom summaries wanted
Knowledge base article preparation Per 100 articles 20-40 Only if KB used for answer highlighting — separate from KB creation
Testing / tuning cycles Per cycle 80-120 Plan for 4-6 cycles Y1
Agent training Per 100 agents 8-12 Blended live/self-paced
Supervisor / admin enablement Per site 16-24
Typical medium implementation (10 queues, 1 language, 100 intents, 100 wrap-ups, 500 agents, 4 tuning cycles, 500 KB articles): ~1,500 hours.
2. Email AI (Auto-Suggest + Auto-Respond)
Activity Unit Hours per unit
Base Email AI setup Fixed 60-100
Intent library for email Per 10 intents 40-60 (higher than voice — more text nuance)
Response template library (Auto-Suggest) Per 20 templates 30-50
Auto-Respond flow design Per flow 60-100 (business rules, escalation logic, guardrails)
Integration to systems of record for response Per integration 80-200
Testing / tuning cycles Per cycle 100-160
Agent training on suggested/edit vs. auto Per 100 agents 6-10
Typical medium implementation: ~1,200-1,800 hours.
Critical note: Auto-Respond at any meaningful rate requires integration to case/order/account data — doesn't work without the ESB. Auto-Suggest is more forgiving. This shapes phasing.
3. Speech & Text Analytics
Activity Unit Hours per unit
STA topic/program setup Fixed 80-120
Program per language Per language 60-100
Topic library — compliance Per 20 topics 20-30
Topic library — CX / operational Per 20 topics 20-30
Category / phrase library tuning Per cycle 60-100 (plan for 3-5 cycles)
Dashboard / report configuration Per dashboard 20-30
Supervisor enablement Per site 8-16
Typical medium implementation: ~600-1,000 hours.
4. Supervisor Copilot
Activity Unit Hours per unit
Supervisor Copilot instance & settings Fixed 40-60
AI Translate configuration per language pair Per pair 8-16
AI Summary insight configuration Fixed 40-60
Alerting rules & thresholds Per rule set 20-40
Supervisor training Per 10 supervisors 8-16
Typical medium implementation: ~300-500 hours.
5. Predictive Routing
Activity Unit Hours per unit
PR model configuration Fixed 60-100
Data source setup and validation Fixed 40-80
Per-queue optimization Per queue 2-4
Baseline measurement & A/B Per cycle 80-120 (plan 2-3 cycles)
Model retraining automation Fixed 20-40
Typical medium implementation: ~500-800 hours.
Cross-cutting activities (allocate across features)
These are the ones that get missed and blow budgets:
Activity Unit Hours
KB curation & prep (source-of-truth for Copilot, Email AI, and STA) Per 100 articles 40-80
KB governance setup (versioning, ownership, refresh cadence) Fixed 100-200
AI governance framework (drift detection, model versioning, escalation paths) Fixed 120-200
Data pipeline / integration to systems of record Per SoR 200-500
Testing environment setup Fixed 80-160
Program management overhead Per month program duration 40-80
Regulatory / compliance review for AI features Per feature 20-60
CTM-specific amplifiers
Now the ugly part. Every parameter above gets multiplied at CTM scale:
Parameter Typical medium CTM
Tails 10-50 1,000+ (6-10× amplifier on queue-configuration line items)
Languages 1-3 7+ (English, French, Spanish, German, Mandarin, Cantonese, Japanese)
Sites 1-3 9 (change management overhead compounds)
Agent count 100-500 ~1,900 (training scales linearly)
Regions 1 4 (NAM, EMEA, AUZ, APAC) — program management overhead compounds
Systems-of-record integration 1-2 pre-built 0 today, ESB Nov 2026+
KB maturity Unknown Unknown — flag as major risk
Amplifier math for Agent Copilot at CTM scale
Using the LoE table above at CTM parameters, mid-range hours:
Activity CTM units Hours
Base Copilot instance × 7 languages 7 700
Settings/rules/permissions 1 50
NLU/intent modeling — 300 intents (large enterprise) 30 1,200
Wrap-up codes — 500 codes 25 250
Queue configuration — 1,000 queues at 2 hrs each 1,000 2,000
Custom summary templates — 15 templates 15 450
KB article preparation — 5,000 articles 50 1,500
Testing/tuning — 6 cycles 6 600
Agent training — 1,900 agents 19 190
Supervisor enablement — 9 sites 9 180
Agent Copilot subtotal ~7,100 hours
Full CTM AI implementation LoE
Feature Estimated hours
Agent Copilot 6,500 - 8,500
Email AI (Auto-Suggest + Auto-Respond) 3,000 - 4,500
Speech & Text Analytics 1,500 - 2,500
Supervisor Copilot 600 - 900
Predictive Routing 1,200 - 1,800
Feature subtotal 12,800 - 18,200
Cross-cutting (KB, governance, PM, integration) 4,000 - 7,000
Total Y1 implementation LoE 16,800 - 25,200 hours
Translating to dollars
I don't know your PS rate, but for context using industry-standard blended rates:
Blended rate Y1 implementation cost range
$175/hr (offshore-heavy blend) $2.9M - $4.4M
$225/hr (typical NTT DATA blended) $3.8 million - $5.7 million
$275/hr (onshore-heavy specialist) $4.6M - $6.9M
Plus annual steady-state at 15-20% of implementation = $430K - $1.4M/yr recurring for ongoing optimization, tuning, KB refresh, model retraining.
What this does to the case
Adding implementation costs to the model:
Component Y1 Y2 Y3 3-Year
Platform license $2.79M $2.79M $2.79M $8.37M
AI token costs (Realistic) $2.0M $3.5M $5.0M $10.5M
AI implementation LoE (new) $3.8 million-5.7 million $0.6M-1.1M $0.6M-1.1M $5.0M-7.9M
Legacy platform takeouts ($2.0M) ($2.0M) ($2.0M) ($6.0M)
Realistic AI benefits ($1.5M) ($4.5M) ($7.5M) ($13.5M)
NET +$5.1M to +$7.0M +$0.4M to +$0.9M -$1.1M to -$1.6M +$4.4M to +$6.3M
In the current model, program is net-negative $4-6M over 3 years even in Realistic scenario. Y1 is the ugly year because implementation cost front-loads. Y3 is when benefits catch up — barely.
And that's using Genesys's own claimed benefits, unadjusted. If we apply the realistic haircuts we discussed earlier (Y1 benefit realization at 30-50%), the picture gets worse.
Calculator amendment
Add to the spec:
New dataclass in inputs.py
@dataclass
class ImplementationEffort:
feature: str
fixed_setup_hours: float
per_queue_hours: float
per_language_hours: float
per_100_intents_hours: float
per_20_wrapups_hours: float
per_100_kb_articles_hours: float
tuning_cycles: int
hours_per_tuning_cycle: float
per_100_agents_training_hours: float
per_site_enablement_hours: float
@dataclass
class ImplementationCosts:
ps_blended_rate: float # $/hr — user-configurable
steady_state_pct_annual: float # default 0.175 (17.5%)
ramp_phasing: dict[str, float] # feature -> % of hours in Y1 vs Y2
New module tokencalc/implementation.py
def calculate_feature_loe(effort: ImplementationEffort,
sites: list[SiteInput],
feature_scope: FeatureScope,
complexity_params: ComplexityParams
) -> pd.DataFrame:
"""Returns hours breakdown by activity for one feature."""
def calculate_implementation_cost(loe_by_feature: dict,
ps_rate: float,
phasing: dict
) -> pd.DataFrame:
"""Distributes hours across Y1/Y2/Y3 with steady-state tail."""
New Streamlit page: "Implementation Effort"
LoE inputs per feature (adjustable, defaults from table above)
Complexity multipliers (queue count, language count, KB size — pulled from Inputs page)
PS rate slider
Steady-state % slider
Output table: hours × rate by feature, by year
Warning: "Implementation cost was NOT in the original Genesys business case."
Smell-test rule to add
If total AI implementation cost < 15% of AI benefit claim, flag as under-modeled. Industry benchmark: 20-40% of Y1 benefit claim goes to implementation.

View File

@@ -0,0 +1,64 @@
Correction #1 — Email Auto-Suggest is part of Copilot
You're correct, and I got this wrong. Per Genesys documentation, Agent Copilot handles both voice and digital channels (email, chat, messaging) — including Auto-Suggest for email. Email Auto-Suggest is not a separate product; it's a Copilot capability on the email channel.
Auto-Respond is the separate line — that's Email Bot / Digital Bot flow territory, distinct from Copilot's suggest-and-agent-approves model.
That collapses my LoE table. Corrected structure:
Feature (corrected) What it actually is
Agent Copilot Voice + digital channels (including email Auto-Suggest, chat suggest, messaging suggest) — one implementation
Email Auto-Respond Automated response without agent — separate flow, needs system-of-record integration
STA Standalone product
Supervisor Copilot Standalone product
Predictive Routing Standalone product
So I was double-counting Copilot email work. Email AI as I described it was inflated.
Correction #2 — The hours were wrong. Let me show my work.
You called it correctly. 16,800-25,200 hours is 8-12 person-years of labour for five features on a mature platform. That's absurd on its face. Genesys Cloud is not custom software; these features are configuration, not development. Let me diagnose where I inflated:
Where I over-estimated
Queue configuration at 2 hours × 1,000 queues = 2,000 hours for Copilot alone. Nonsense. Copilot doesn't require per-queue configuration in the way I implied. Enable Copilot on the account, configure once, apply to queues via policy. Real work: ~50-100 hours to define which queue types get Copilot, plus per-queue-type tuning, not per-individual-queue.
NLU / intent modeling at 40 hours per 10 intents = 1,200 hours for 300 intents. Overstated. Copilot leverages Genesys's pre-built LLM understanding — you're not training an NLU from scratch. You're configuring wrap-up code mappings and validating auto-summary quality. Real work: ~200-400 hours for a large enterprise intent library, not 1,200.
KB article prep at 40 hours per 100 articles = 1,500 hours for 5,000 articles. Double-counted. KB curation is either (a) already done in your KB platform and Copilot indexes it, or (b) a separate KB maturity project — not an AI implementation line item. Real work: ~200-400 hours for connector setup and indexing configuration, plus a separate flag that says "KB curation is a prerequisite project, not part of AI implementation."
Testing/tuning at 100 hours × 6 cycles. Overstated. Test cycles for a hosted AI feature aren't full regression cycles — they're validation of prompt output quality and adjustment. Real: ~40-60 hours per cycle, 3-4 cycles typical.
Training at 10 hours per 100 agents. Wildly high. This is a Copilot UI change — 30 minutes of orientation, some job aids, super-user support. Real: ~1-2 hours per 100 agents for train-the-trainer + material creation.
Cross-cutting at 4,000-7,000 hours. Inflated by carrying forward the double-counted items above.
Corrected LoE — Genesys AI features at CTM scale
Working from realistic effort, not vendor-services-inflation:
Feature Realistic hours (CTM scale)
Agent Copilot (voice + digital, all languages) 1,200 - 1,800
Email Auto-Respond (separate from Copilot; needs integration) 800 - 1,400
STA (topics, programs, tuning for 7 languages) 800 - 1,200
Supervisor Copilot 200 - 400
Predictive Routing 400 - 700
Feature subtotal 3,400 - 5,500
Cross-cutting (governance, PM, testing environment, integration coordination) 1,000 - 1,800
KB readiness project (separate line — prerequisite) 500 - 1,500 (flagged separately)
Total Y1 AI implementation 4,400 - 7,300 hours
Annual steady-state (Y2, Y3) 500 - 900 hours
That's 2-3.5 person-years of Y1 effort across 5 features. Still substantial — this is a real, multi-workstream program at CTM scale — but not the fantasy 12 person-years I had before.
Corrected cost impact
Blended rate Y1 implementation Annual steady-state
$175/hr $770K - $1.28M $88K - $158K
$225/hr $990K - $1.64M $113K - $203K
$275/hr $1.21M - $2.01M $138K - $248K
Updated combined case
Using $225/hr blended rate and Realistic scenario:
Component Y1 Y2 Y3 3-Year
Genesys CX 3 platform license $2.79M $2.79M $2.79M $8.37M
Base platform implementation $1.5M — — $1.5M
AI token costs $2.0M $3.5M $5.0M $10.5M
AI implementation (corrected) $1.3M $0.16M $0.16M $1.6M
Total future-state cost $7.6M $6.45M $7.95M $22.0M
Current-state takeout ($7.3M) ($7.3M) ($7.3M) ($21.9M)
AI benefits (realistic) ($1.5M) ($4.5M) ($7.5M) ($13.5M)
Program net -$1.2M +$5.35M +$6.85M +$11.0M