# Mercury Notebook Deliverable Pattern v1.0.0 Standardizes how Palladium studies ship business-case deliverables: a Mercury-served Jupyter notebook **is** the artifact β€” math in a self-contained study package, presentation in reactive notebooks, verification gates pinning every number, and LLM-readable exports. All new studies follow this pattern; the Streamlit app path is retired by it. ## 🐾 Red Panda Approvalβ„’ This pattern follows Red Panda Approval standards. **Audience note:** this document is written to be loaded whole as context by an LLM agent building or modifying a study. Rules are imperative (MUST/SHOULD/NEVER), each with a one-line *why*. Long code lives in the runnable template [`template/MercuryNotebook/`](../template/MercuryNotebook/) β€” copy it to start a study; snippets here are excerpts from it. The full-scale reference implementation is the CTM Genesys study, [`studies/202512_GenesysCX/ctm-token-calculator/`](../studies/202512_GenesysCX/ctm-token-calculator/). --- ## Why a Pattern, Not a Shared Implementation Every study's domain math is different, and studies are **frozen deliverables**: - The CTM Genesys study models per-feature AI token meters, WFM benefit scoping, and signed-contract mechanics (ramp, SOW milestones, managed services). - A TEI study reproduces a Forrester composite organization's benefits/costs and then personalizes them to a client. - A migration scenario strips capabilities out and prices the platform move alone. A shared calculation engine would make every signed business case a hostage of the next study's refactor β€” a change to shared NPV rounding could silently move a number the client already approved. So: - **Each study package owns ALL of its math**, including ~50 lines of finance primitives (NPV, payback). Duplication is the accepted price of immutability. - The pattern standardizes **structure and contracts** β€” layout, reactivity, staging, verification, export β€” not code. - The only shared plumbing is Athena access (`core/tei_client`) for client/opportunity context; it touches no study math. Instead, this pattern defines: - **Required structure & contracts** β€” every study must have - **Standard choice values** β€” conventions for interoperability (and LLM legibility) - **Recommended practices** β€” most studies should include - **Anti-patterns** β€” the mistakes that cost us debugging sessions, catalogued --- ## Repository Layout & Naming ``` palladium/ β”œβ”€β”€ docs/ # repo-wide docs (this pattern, brand.md) β”œβ”€β”€ template/ β”‚ └── MercuryNotebook/ # copy-me starting point (runnable) └── studies/ β”œβ”€β”€ YYYYMM_TEI_Vendor_Product/ # vendor TEI study, e.g. 202602_TEI_Amazon_Connect └── YYYYMM_Client_EngagementName/ # client study, e.g. 202512_CTM_GenesysCX ``` - Names MUST use **underscores, never dashes** β€” dashed directories can't be Python packages, and everything in a study is importable code. *(The CTM study's inner `ctm-token-calculator/` predates this rule; it gets renamed when the studies migrate.)* - Every study is self-contained with this layout (from the template): ``` / β”œβ”€β”€ / # THE study package β€” all math lives here, renamed per study β”‚ β”œβ”€β”€ model.py # domain model: anchors, schedules, case frame, KPIs β”‚ └── staging.py # stage/backstage detection (copy verbatim) β”œβ”€β”€ notebooks/ # the deliverables β€” presentation only β”œβ”€β”€ tests/ # hand-checked pinned acceptance numbers β”œβ”€β”€ scripts/export_report.py β”œβ”€β”€ docs/ # source material (vendor decks, contracts, LoE docs) β”œβ”€β”€ exports/ # generated .html + .md report sources β”œβ”€β”€ config.toml # Mercury app-shell theme (NTT DATA brand) └── pyproject.toml # full toolchain as core deps ``` --- ## Required Structure & Contracts The non-negotiables. Every study MUST satisfy all seven. ### 1 Β· The notebook is the artifact There is exactly **one** implementation of the study: the notebook(s), served interactively with `mercury --working-dir notebooks/` for stakeholders, opened in JupyterLab by analysts, executed headless by nbconvert for exports. NEVER build a parallel UI (Streamlit, Dash, a second "app" rendering the same model) β€” two surfaces over one model always diverge, and the notebook already *is* the interactive surface. ### 2 Β· Engine/presentation split Notebooks hold **no math**. Every number a stakeholder sees is computed in the study package and imported. *Why:* the package is testable and diffable; notebook cells are neither. The test suite pins the engine; the notebook only arranges its outputs. ```python # notebook cell β€” arrange and render, never compute from studylib.model import case_flows, case_kpis, money inc, net_by = case_flows(total_by_year, ben_by_year) kpi = case_kpis(inc, net_by, DISCOUNT_RATE) ``` ### 3 Β· The Mercury reactivity contract Mercury re-executes only the cells **below** a changed widget's cell β€” never the defining cell itself (upstream: `onWidgetUpdate` re-runs from `updatedIndex + 1`). Therefore, always a **widget-pair**: ```python # ── Cell A: widgets ONLY β€” no other statements, no output ─────────── _platform_w = mr.NumberInput(label="Platform run-rate ($/yr) β€” contracted", value=round(anchor("platform_annual")), min=0, max=10_000_000, step=25_000) ``` ```python # ── Cell B (below A): read .value, derive model state ─────────────── PLATFORM_ANNUAL = float(_platform_w.value) platform_by_year = platform_costs_by_year(RAMP_MONTHS, PLATFORM_ANNUAL) ``` - A `.value` read in the defining cell is **frozen at first render** β€” the classic "I changed the input and nothing happened" bug. - Widget cells MUST produce no other output β€” stray output in a widget cell leaks into the Mercury sidebar. - Always pass explicit `min=`/`max=` to `NumberInput` β€” Mercury clamps out-of-range seeds to a small default range (and older builds crash doing it). - Headless nbconvert takes every widget at its seed value, which is why defaults must form a coherent, gate-passing scenario. ### 4 Β· Verification gate cell Every deliverable notebook ends its analytical sections with a gate cell that MUST pass under headless `jupyter nbconvert --execute` β€” it is the study's smoke test and the reviewer's proof that the rendered numbers match the engine. ```python def _approx(got, want, tol=0.5): assert abs(got - want) <= tol, f"got {got:,.2f}, want {want:,.2f}" # Engine pins β€” EXPLICIT default arguments, independent of widget state _approx(anchor("platform_annual"), 500_000) # signed overlay _approx(platform_costs_by_year()[2026], 300_000) # hand-checked # Live-state checks β€” guarded, so a stakeholder moving a slider # doesn't crash the notebook _at_default = (PLATFORM_ANNUAL == round(anchor("platform_annual"))) if _at_default: _approx(kpi["net_3yr"], 400_000) for y in YEARS: # ties hold at ANY widget state _approx(net_by[y], ben_by_year[y] - (total_by_year[y] - BASELINE_ANNUAL)) backstage("All assertions passed.") ``` Pin *engine calls with explicit defaults* (widget-independent), guard *live-state value* checks with `_at_default`, and assert *structural ties* (stack sums, flow identities) unconditionally β€” they must hold at any widget setting. ### 5 Β· Machine-readable data appendix The last section dumps every model output as markdown tables plus **one JSON block of model state** (KPIs + assumptions). *Why:* plotly figures export as JavaScript an LLM cannot read; the appendix makes the exported `.md` complete LLM input β€” and it is the future Athena study-export payload. Renders backstage only (Β§6). See the appendix cell in [`template/MercuryNotebook/notebooks/business_case.ipynb`](../template/MercuryNotebook/notebooks/business_case.ipynb). ### 6 Β· Stage / backstage The Mercury app is **the stage** β€” stakeholder-facing. JupyterLab and nbconvert are **backstage** β€” analyst diagnostics and the data appendix belong there. Detection: the `mercury` CLI exports `MERCURY_CONFIG_DIR` into its server and every kernel inherits it; JupyterLab/nbconvert kernels don't have it. ```python # studylib/staging.py β€” copy verbatim def on_stage() -> bool: """True when running under the Mercury app (stakeholder-facing).""" return os.getenv("MERCURY_CONFIG_DIR") is not None def backstage(*args, **kwargs) -> None: """``print`` that renders only backstage (JupyterLab, nbconvert).""" if not on_stage(): print(*args, **kwargs) ``` Route through `backstage()`: dict echoes, seeds/cross-foots, gate chatter, analyst prose, the whole data appendix. Keep on stage (plain `print`): only curated one-line insights a stakeholder should read (e.g. *"the licence correction moves the 3-yr case by $2.2M"*). ### 7 Β· Packaging `pyproject.toml` declares the **whole toolchain as core dependencies** β€” mercury, jupyterlab, ipywidgets, nbconvert, tabulate, pandas, plotly, openpyxl. `pip install -e .` must be enough to serve, edit, test, and export. NEVER hide a runtime dep in an extra and NEVER keep a parallel `requirements.txt` β€” a stakeholder demo failing on a missing `tabulate` is how this rule was learned. Only `dev = ["pytest", "mypy"]` remains an extra. See [`template/MercuryNotebook/pyproject.toml`](../template/MercuryNotebook/pyproject.toml). --- ## Standard Choice Values Use these exact conventions β€” LLM agents and humans navigate studies by them. ### Confidence legend Every input and cost/benefit line carries one, in narrative and in tables: | Icon | Meaning | |---|---| | 🟒 | confirmed β€” published or contractual | | 🟑 | estimated β€” working assumption, stated | | πŸ”΄ | unknown β€” flagged, bounded by sensitivity if material | ### Naming | Convention | Example | Why | |---|---|---| | `_*_w` for widget objects | `_licence_w`, `_ramp_w` | underscore = presentation plumbing, `_w` = widget-pair member | | UPPERCASE for widget-derived state | `LICENCE_ANNUAL = float(_licence_w.value)` | reads as the notebook's scenario constants | | `*_by_year` dicts keyed by **calendar year** | `{2026: 1_600_000, ...}` | calendar years, never year-index 1/2/3, in anything displayed | | `*_VERBATIM` for source anchors | `TCO_VERBATIM`, `ANCHOR_VERBATIM` | the client/vendor record β€” NEVER edited (see Recommended) | | `money()` / `html_money()` | `$3.2M`, `$3.2M` | one house format; see the MathJax trap in Anti-Patterns | ### Section anchors Each section heading carries ``; a sidebar table of contents uses **onclick-JS** navigation (fragment `href`s don't scroll in Mercury's SPA, and python-markdown escapes any raw `<` inside handler attributes β€” keep handlers comparison-free). Recipe: the ToC cell in [`studies/202512_GenesysCX/ctm-token-calculator/notebooks/ctm_business_case_corrected.ipynb`](../studies/202512_GenesysCX/ctm-token-calculator/notebooks/ctm_business_case_corrected.ipynb). --- ## Recommended Practices Most studies should include these; deviate knowingly. ### Verbatim anchors + contracted overlay Keep the client's/vendor's source record **verbatim and immutable**; layer signed or corrected values over it; read through a helper: ```python ANCHOR_VERBATIM = {"platform_annual": 600_000, ...} # the pitch β€” do not edit ANCHOR_CONTRACTED = {"platform_annual": 500_000} # 🟒 signed def anchor(key): return ANCHOR_CONTRACTED.get(key, ANCHOR_VERBATIM[key]) ``` *Why:* the stakeholder walk starts from *their own numbers*. Corrections presented as overlays are auditable; corrections made by editing the source are arguments. Show a **vendor/deck-frame KPI column beside the contracted column** so the walk from the pitch to reality stays explicit. Reference: `TCO_VERBATIM`/`TCO_CONTRACTED`/`tco()` in [`studies/202512_GenesysCX/ctm-token-calculator/tokencalc/appendix4.py`](../studies/202512_GenesysCX/ctm-token-calculator/tokencalc/appendix4.py). ### Baseline-relative case frame Price against *doing nothing*: incremental cost = programme cost βˆ’ baseline run-rate; net = benefits βˆ’ incremental cost. One frame captures double-billing penalties and post-termination cost-avoidance credits without special cases. Where a case doesn't pay back in-window, add a **run-rate breakeven extrapolation** KPI (steady-state saving per year fills the end-of-window deficit). ### Chart chrome One `tei_layout()` helper per study (title top-left, recessive grid, ink text, house font stack); `bar()` with a 2px surface `marker_line` (the gap that separates stacked segments); fixed entityβ†’color maps (`COST_COLOR`, `REGION_COLOR`) so color follows the entity across every figure; diverging heatmaps centered with `zmid=0`; selective direct labels, not a number on every mark. Skeleton in the template's setup cell; palette rules in the repo dataviz reference and [`docs/brand.md`](brand.md). ### Widget placement Sidebar (default) for scenario knobs; `position="inline"` for data-collection tables that belong in the page flow (per-region costs, termination dates). Seed every widget from the engine's defaults (`value=round(anchor(...))`) so notebook and engine can't disagree about the starting scenario. ### Contract tables on stage Milestones, billing schedules, and scope tables render on stage as small DataFrames β€” stakeholders verify contracts by looking at them, not by trusting prose. --- ## Athena Integration - **Inbound (now):** pull client/opportunity context (company, industry, engagement metadata) via `core/tei_client` at study setup, instead of re-keying it. - **Outbound (roadmap):** on study completion, push the study export β€” the `.md` report source and the appendix JSON model state β€” to Athena, which becomes the **repository of the study**. Athena exposes an MCP server, so an LLM can pull a completed study export directly. This is why the data appendix (Required Β§5) is a contract, not a nicety: it is the export payload. --- ## Pattern Variants ### Variant 1 β€” Corrected / pressure-tested business case Keep the vendor's claimed benefits **verbatim**, add the costs the pitch omitted (consumption meters, implementation labour, double-billing), and bill contract mechanics as signed (ramp, milestones, managed services). The headline is the walk: *as-pitched β†’ corrected*. Reference: [`notebooks/ctm_business_case_corrected.ipynb`](../studies/202512_GenesysCX/ctm-token-calculator/notebooks/ctm_business_case_corrected.ipynb). ### Variant 2 β€” Scenario notebook on a thin module A second question over the same engine (e.g. "migration + WFM only, no AI") gets a **thin scenario module** that scopes and extrapolates but duplicates nothing, plus its own notebook and test pins. Reference: `tokencalc/migration_wfm.py` + [`notebooks/ctm_migration_wfm.ipynb`](../studies/202512_GenesysCX/ctm-token-calculator/notebooks/ctm_migration_wfm.ipynb). ### Variant 3 β€” Exploratory calculator Early-phase what-if surface: scenario selectors, tornado/break-even sweeps, no contract anchoring yet. Still engine-backed and gate-checked; it graduates into Variant 1 as facts arrive. Reference: [`notebooks/ctm_token_calculator.ipynb`](../studies/202512_GenesysCX/ctm-token-calculator/notebooks/ctm_token_calculator.ipynb). ### Variant 4 β€” TEI composite reproduction Reproduce a published TEI study's composite organization as the verbatim anchor (`ANCHOR_VERBATIM` = Forrester's tables), verify the reproduction against the published ROI/NPV/payback in the gate, then personalize with client inputs as the overlay. This is the target shape for `studies/202602_AmazonConnect/` when it migrates off Streamlit. --- ## Domain Extension Examples ### Adding domain math to a study New capability = new engine code + pins + a notebook section, in that order: 1. Model it in the study package (e.g. `studylib/model.py` grows `managed_services_by_year()`), keeping verbatim anchors separate from contracted overlays. 2. Pin it in `tests/` with hand-checked numbers **before** wiring the notebook. 3. Add the notebook section: markdown narrative β†’ (widget pair if tunable) β†’ engine call β†’ table/figure β†’ new gate pins β†’ appendix keys. ### Worked example (CTM, contract facts) The signed SOW added billing milestones and managed services. The change landed as: `PS_MILESTONES` + `managed_services_by_year()` in `tokencalc/appendix4.py`; pins in `tests/test_appendix4.py` (milestone shares cross-foot, proration hand-checked); a milestone table on stage; gate re-pinned; appendix JSON gained `ps_milestones`/`managed_services_annual`. The deck-frame column kept the vendor's original PS lump β€” anchors stay verbatim. --- ## Anti-Patterns Each of these cost a debugging session or a client-facing embarrassment. Don't. - ❌ **Don't read `widget.value` in the cell that defines the widget** β€” Mercury never re-runs the defining cell; the value freezes at first render. - ❌ **Don't put math in notebook cells** β€” untestable, undiffable; it will drift from the engine. - ❌ **Don't edit verbatim anchors** β€” layer a contracted/corrected overlay and read through a helper; the source record is the client's trust anchor. - ❌ **Don't build a parallel UI** (Streamlit twin, second app) β€” two surfaces over one model always diverge; the notebook is the surface. - ❌ **Don't hide runtime deps in extras or a requirements.txt** β€” `pip install -e .` must fully provision serve/test/export. - ❌ **Don't put two bare `$` in one plotly annotation** β€” MathJax eats the text; use `html_money()` (`$`). - ❌ **Don't emit any output from a widget cell** β€” it leaks into the Mercury sidebar. - ❌ **Don't gate on live widget state without an `_at_default` guard** β€” a stakeholder moving a slider must not crash the notebook. - ❌ **Don't hardcode KPI numbers in markdown** β€” compute them into prints/annotations, or the gate can't catch the drift when inputs change. - ❌ **Don't use kebab-case directories** β€” they can't be imported. - ❌ **Don't use fragment-`href` ToC links** β€” they don't scroll in Mercury's SPA; use onclick-JS (and no raw `<` inside handler attributes β€” python-markdown escapes the tag). - ❌ **Don't construct `NumberInput` without `min=`/`max=`** β€” Mercury clamps to a small default range and out-of-range seeds break. --- ## Settings ```toml # config.toml β€” loaded from the directory where you launch `mercury`; # restart the server to apply. Template: template/MercuryNotebook/config.toml [main] # title, favicon_emoji, footer, notebooks_button_label [welcome] # gallery header + message (name every notebook and the export command) [theme] # NTT DATA brand palette β€” see docs/brand.md ``` - `MERCURY_CONFIG_DIR` β€” set by the `mercury` CLI for its server; kernels inherit it. The pattern uses its **presence** as the stage signal (`studylib/staging.py`). Do not set it manually except to simulate the stage in tests. - Serve from the project root: `mercury --working-dir notebooks/` (so `config.toml` loads). Analyst view: `jupyter lab`. Exports: `python scripts/export_report.py`. --- ## Testing Layers, from inner to outer β€” every study ships all four: 1. **Engine pins** (`tests/`) β€” hand-checked acceptance numbers for every model function; the contracted overlay AND the verbatim record both pinned, so neither can drift. Compute by hand first, then pin (see [`template/MercuryNotebook/tests/test_model.py`](../template/MercuryNotebook/tests/test_model.py)). 2. **Stage/backstage test** β€” `monkeypatch` `MERCURY_CONFIG_DIR`, assert `backstage()` prints only off stage (copy `tests/test_staging.py` from the template). 3. **In-notebook gate** (Required Β§4) β€” proves the *rendered* notebook matches the engine; runs on every execution, interactive or headless. 4. **Headless execution + export** β€” the CI-style check: ```bash pytest jupyter nbconvert --to notebook --execute --inplace notebooks/*.ipynb # gates green? python scripts/export_report.py # exports carry appendix + JSON? MERCURY_CONFIG_DIR=$(mktemp -d) jupyter nbconvert --to notebook --execute \ --inplace notebooks/_stage_sim.ipynb # stage sim: gate+appendix silent on stage ``` After any engine change: recompute expected numbers first, update test pins and gate pins together, re-execute all notebooks, regenerate exports β€” and report the KPI moves honestly.