Compare commits
8 Commits
5af198911b
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 84e026c2e7 | |||
| d8703a1ad6 | |||
| 537f3c7963 | |||
| b5a3aa214b | |||
| d528192bee | |||
| 455d3eef3d | |||
| 0ae0d55a8d | |||
| 73f6afdbf8 |
@@ -218,7 +218,7 @@ anthropic:
|
|||||||
|
|
||||||
That's the whole configuration. Pallas auto-detects the
|
That's the whole configuration. Pallas auto-detects the
|
||||||
`bedrock-mantle` hostname in `anthropic.base_url` at startup and installs
|
`bedrock-mantle` hostname in `anthropic.base_url` at startup and installs
|
||||||
two compatibility shims so fast-agent's default request shape matches
|
four compatibility shims so fast-agent's default request shape matches
|
||||||
what Mantle expects (see `pallas/mantle_shims.py`):
|
what Mantle expects (see `pallas/mantle_shims.py`):
|
||||||
|
|
||||||
1. **Wire-name prefix** — re-adds the `anthropic.` prefix that fast-agent's
|
1. **Wire-name prefix** — re-adds the `anthropic.` prefix that fast-agent's
|
||||||
@@ -233,6 +233,25 @@ what Mantle expects (see `pallas/mantle_shims.py`):
|
|||||||
Input should be a valid dictionary or object"`, which would otherwise
|
Input should be a valid dictionary or object"`, which would otherwise
|
||||||
break the MCP tool-use loop on the second turn.
|
break the MCP tool-use loop on the second turn.
|
||||||
|
|
||||||
|
3. **Fine-grained tool streaming opt-out** — stops fast-agent sending the
|
||||||
|
`fine-grained-tool-streaming-2025-05-14` beta. Under that beta an
|
||||||
|
output-token cutoff mid-`tool_use` block ends the stream without
|
||||||
|
`content_block_stop`, which fast-agent surfaces as
|
||||||
|
`Streaming completed but tool call never finished` and then retries
|
||||||
|
into the same wall (~700 s agent failures on large tool bodies).
|
||||||
|
Without the beta a cutoff closes blocks properly and lands in
|
||||||
|
fast-agent's graceful `stop_reason=max_tokens` handling.
|
||||||
|
|
||||||
|
4. **`max_tokens` clamp** — Mantle enforces a server-side output ceiling
|
||||||
|
of 20 000 tokens per response regardless of the requested
|
||||||
|
`max_tokens` (fast-agent asks for the model's full 128 000). The shim
|
||||||
|
clamps default `maxTokens` to `MANTLE_MAX_OUTPUT_TOKENS` (20 000) so
|
||||||
|
the model stops gracefully at the limit instead of the gateway
|
||||||
|
cutting the stream. A single agent turn — thinking, prose, and tool
|
||||||
|
input combined — cannot exceed this on Mantle; agents that must emit
|
||||||
|
more output in one turn need to split the work (e.g. patch-style
|
||||||
|
edits instead of full-document rewrites).
|
||||||
|
|
||||||
The Anthropic SDK appends `/v1/messages` to `base_url` automatically.
|
The Anthropic SDK appends `/v1/messages` to `base_url` automatically.
|
||||||
|
|
||||||
**Feature support.** Mantle accepts the same Messages API request shape
|
**Feature support.** Mantle accepts the same Messages API request shape
|
||||||
|
|||||||
@@ -191,7 +191,9 @@ agents:
|
|||||||
| `agents.<name>.module` | yes | Importable Python module path containing a `fast` instance |
|
| `agents.<name>.module` | yes | Importable Python module path containing a `fast` instance |
|
||||||
| `agents.<name>.port` | yes | Port for this agent's StreamableHTTP MCP server |
|
| `agents.<name>.port` | yes | Port for this agent's StreamableHTTP MCP server |
|
||||||
| `agents.<name>.title` | no | Display name in registry. Default: `name.title()` |
|
| `agents.<name>.title` | no | Display name in registry. Default: `name.title()` |
|
||||||
| `agents.<name>.description` | no | Description in registry |
|
| `agents.<name>.description` | no | Description in registry. Also becomes the `send_message` tool description, overriding any `description=` on the `@fast.agent` decorator |
|
||||||
|
| `agents.<name>.model` | no | `provider.model-name` override for this agent. Overrides `default_model`, is applied to every agent in the module at startup, and is what the registry advertises for this entry |
|
||||||
|
| `agents.<name>.model_capabilities` | no | Per-agent `{vision, context_window, max_output_tokens}` block. Overrides the top-level `model_capabilities`; the same defaults apply to omitted fields |
|
||||||
| `agents.<name>.depends_on` | no | List of agent names that must start and become ready before this agent |
|
| `agents.<name>.depends_on` | no | List of agent names that must start and become ready before this agent |
|
||||||
| `agents.<name>.max_iterations` | no | Hard cap on agentic-loop turns per `send_message`. Default: `15`. fast-agent returns a partial answer once exceeded |
|
| `agents.<name>.max_iterations` | no | Hard cap on agentic-loop turns per `send_message`. Default: `15`. fast-agent returns a partial answer once exceeded |
|
||||||
| `agents.<name>.loop_repeat_threshold` | no | Halt the loop after this many consecutive identical `(tool, args) → result` rounds. Default: `3`. `0` disables the guard |
|
| `agents.<name>.loop_repeat_threshold` | no | Halt the loop after this many consecutive identical `(tool, args) → result` rounds. Default: `3`. `0` disables the guard |
|
||||||
@@ -429,7 +431,9 @@ Built dynamically from `agents.yaml` + `fastagent.config.yaml`:
|
|||||||
|
|
||||||
### Capabilities
|
### Capabilities
|
||||||
|
|
||||||
If `model_capabilities` is defined in `fastagent.config.yaml`, each registry entry includes a `capabilities` object with model name, vision support, context window, and max output tokens. This allows clients to make informed decisions about what an agent can handle.
|
Each registry entry includes a `capabilities` object — model name, vision support, context window, and max output tokens — whenever a model is known for that agent, i.e. `agents.<name>.model` is set or `default_model` is defined in `fastagent.config.yaml`.
|
||||||
|
|
||||||
|
Capabilities are resolved **per agent**: the agent's own `model` and `model_capabilities` take precedence over the global values, and omitted fields fall back to the same defaults Pallas uses to register the model (`vision: false`, `context_window: 131072`, `max_output_tokens: 16384`). The published values therefore match what was actually registered with fast-agent's `ModelDatabase`, rather than being null whenever `model_capabilities` was left out. Clients use them to make informed decisions about what an agent can handle.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -448,6 +452,16 @@ Each agent's MCP tool accepts:
|
|||||||
|
|
||||||
When `images` is provided, the message is sent as a `PromptMessageExtended` containing both `TextContent` and `ImageContent` parts — the agent's underlying model must support vision.
|
When `images` is provided, the message is sent as a `PromptMessageExtended` containing both `TextContent` and `ImageContent` parts — the agent's underlying model must support vision.
|
||||||
|
|
||||||
|
#### Tool description
|
||||||
|
|
||||||
|
The tool's description is what an MCP client shows next to the tool name, so it should say what the agent is *for*. Pallas resolves it in this order:
|
||||||
|
|
||||||
|
1. `agents.<name>.description` from `agents.yaml` — the deployment's source of truth, and the same text published in the registry
|
||||||
|
2. `description=` on the `@fast.agent` decorator
|
||||||
|
3. `Send a message to the {agent} agent` — a generic fallback
|
||||||
|
|
||||||
|
A description containing `{agent}` has the agent's name interpolated into it; other braces are left alone. Without (1), every agent whose module omits (2) falls through to the fallback, which tells a client nothing — so keep `agents.yaml` descriptions meaningful.
|
||||||
|
|
||||||
### Tool-Result Image Passthrough
|
### Tool-Result Image Passthrough
|
||||||
|
|
||||||
Images work in both directions. fast-agent's `agent.send()` returns only the final assistant text, so images produced by downstream tools during the agentic loop (playwright screenshots, rommie desktop captures) would otherwise reach the agent's own vision model but never the MCP caller. A per-request `after_tool_call` hook (`pallas.image_passthrough`) collects every `ImageContent` block from the turn's tool results; at end of turn `send_message` returns a `CallToolResult` whose content is the assistant's text block followed by the collected images. Turns that produce no images return the plain string, unchanged from previous releases.
|
Images work in both directions. fast-agent's `agent.send()` returns only the final assistant text, so images produced by downstream tools during the agentic loop (playwright screenshots, rommie desktop captures) would otherwise reach the agent's own vision model but never the MCP caller. A per-request `after_tool_call` hook (`pallas.image_passthrough`) collects every `ImageContent` block from the turn's tool results; at end of turn `send_message` returns a `CallToolResult` whose content is the assistant's text block followed by the collected images. Turns that produce no images return the plain string, unchanged from previous releases.
|
||||||
@@ -673,12 +687,13 @@ Pallas registers models not in fast-agent's built-in `ModelDatabase` at startup,
|
|||||||
The process:
|
The process:
|
||||||
|
|
||||||
1. Read `default_model` and `model_capabilities` from config
|
1. Read `default_model` and `model_capabilities` from config
|
||||||
2. Extract the model name (portion after the provider prefix dot)
|
2. Also read every `agents.<name>.model` from `agents.yaml`, using that agent's own `model_capabilities` when it declares one
|
||||||
3. Check if `ModelDatabase` already knows this model — if so, skip
|
3. Extract the model name (portion after the provider prefix dot)
|
||||||
4. Register with `ModelDatabase.register_runtime_model_params()`:
|
4. Check if `ModelDatabase` already knows this model — if so, skip
|
||||||
|
5. Register with `ModelDatabase.register_runtime_model_params()`:
|
||||||
- `vision: true` → multimodal tokenization (`QWEN_MULTIMODAL`)
|
- `vision: true` → multimodal tokenization (`QWEN_MULTIMODAL`)
|
||||||
- `vision: false` → text-only tokenization (`TEXT_ONLY`)
|
- `vision: false` → text-only tokenization (`TEXT_ONLY`)
|
||||||
- `context_window` and `max_output_tokens` from config (with sensible defaults)
|
- `context_window` and `max_output_tokens` from config, defaulting to `131072` / `16384` — the same values the registry advertises
|
||||||
|
|
||||||
This avoids the brittle pattern of inferring capabilities from model name substrings, which breaks for custom or fine-tuned models with non-standard names.
|
This avoids the brittle pattern of inferring capabilities from model name substrings, which breaks for custom or fine-tuned models with non-standard names.
|
||||||
|
|
||||||
|
|||||||
@@ -120,7 +120,7 @@ No authentication. No query parameters.
|
|||||||
| `servers[].server.version` | string | no | Semver version string. |
|
| `servers[].server.version` | string | no | Semver version string. |
|
||||||
| `servers[].server.icons` | array | no | Array of `{ src, sizes }`. Daedalus uses the first entry. |
|
| `servers[].server.icons` | array | no | Array of `{ src, sizes }`. Daedalus uses the first entry. |
|
||||||
| `servers[].server.remotes` | array | yes | Connection endpoints. Daedalus looks for `type: "streamable-http"` and uses its `url`. |
|
| `servers[].server.remotes` | array | yes | Connection endpoints. Daedalus looks for `type: "streamable-http"` and uses its `url`. |
|
||||||
| `servers[].server.capabilities` | object | no | Model capabilities. Contains `model` (string), `vision` (bool), `context_window` (int), `max_output_tokens` (int). Published when `model_capabilities` is configured in `fastagent.config.yaml`. |
|
| `servers[].server.capabilities` | object | no | Model capabilities. Contains `model` (string), `vision` (bool), `context_window` (int), `max_output_tokens` (int). Published whenever a model is known for the agent (`agents.<name>.model` or `default_model`); resolved per agent, with omitted fields falling back to the defaults Pallas registered the model with. |
|
||||||
| `servers[]._meta` | object | no | Registry metadata. Informational only — Daedalus does not act on it. |
|
| `servers[]._meta` | object | no | Registry metadata. Informational only — Daedalus does not act on it. |
|
||||||
|
|
||||||
#### Behaviour
|
#### Behaviour
|
||||||
|
|||||||
@@ -24,7 +24,23 @@ before the wire traffic is valid:
|
|||||||
|
|
||||||
Upstream SDK tracker: https://github.com/anthropics/anthropic-sdk-python/issues/1454
|
Upstream SDK tracker: https://github.com/anthropics/anthropic-sdk-python/issues/1454
|
||||||
|
|
||||||
Both shims are idempotent and may be installed at process startup before any
|
3. **Fine-grained tool streaming truncation.** Fast-agent unconditionally
|
||||||
|
sends the ``fine-grained-tool-streaming-2025-05-14`` beta on tool-bearing
|
||||||
|
requests. Under that beta, an output-token cutoff mid-``tool_use`` block
|
||||||
|
ends the stream *without* ``content_block_stop``, which fast-agent's
|
||||||
|
stream accounting surfaces as ``Streaming completed but tool call never
|
||||||
|
finished`` — followed by a full retry ladder against the same wall
|
||||||
|
(observed as ~700 s agent failures on large ``revise_workspace_file``
|
||||||
|
bodies). We disable that one beta so a cutoff closes blocks properly and
|
||||||
|
lands in fast-agent's graceful ``stop_reason=max_tokens`` handling.
|
||||||
|
|
||||||
|
4. **Output-token ceiling.** Mantle clamps ``max_tokens`` to 20 000
|
||||||
|
server-side (streams complete at exactly 20 000 output tokens regardless
|
||||||
|
of the requested 128 000). We clamp the request to that ceiling so the
|
||||||
|
*model* stops gracefully at the limit — emitting proper block-stop events
|
||||||
|
and ``stop_reason`` — instead of being cut off by the gateway's clamp.
|
||||||
|
|
||||||
|
All shims are idempotent and may be installed at process startup before any
|
||||||
fast-agent ``FastAgent`` instance is constructed.
|
fast-agent ``FastAgent`` instance is constructed.
|
||||||
"""
|
"""
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
@@ -129,12 +145,104 @@ def install_tool_use_caller_strip() -> None:
|
|||||||
logger.info("Mantle tool_use.caller strip shim installed")
|
logger.info("Mantle tool_use.caller strip shim installed")
|
||||||
|
|
||||||
|
|
||||||
|
# ── Shim 3: disable fine-grained tool streaming ──────────────────────────────
|
||||||
|
|
||||||
|
_beta_opt_out_installed = False
|
||||||
|
|
||||||
|
|
||||||
|
def install_fine_grained_tool_streaming_opt_out() -> None:
|
||||||
|
"""Stop fast-agent requesting the fine-grained tool streaming beta.
|
||||||
|
|
||||||
|
Under that beta a ``max_tokens`` cutoff mid-``tool_use`` ends the stream
|
||||||
|
without ``content_block_stop``; fast-agent then raises
|
||||||
|
``Streaming completed but tool call never finished`` and burns its whole
|
||||||
|
retry ladder against the same ceiling. Without the beta the cutoff closes
|
||||||
|
blocks properly and fast-agent's ``stop_reason=max_tokens`` handling
|
||||||
|
applies. Safe to call more than once.
|
||||||
|
"""
|
||||||
|
global _beta_opt_out_installed
|
||||||
|
if _beta_opt_out_installed:
|
||||||
|
return
|
||||||
|
|
||||||
|
from fast_agent.llm.provider.anthropic.llm_anthropic import AnthropicLLM
|
||||||
|
|
||||||
|
original_supports = AnthropicLLM.supports_direct_anthropic_beta
|
||||||
|
|
||||||
|
def patched_supports(self: Any, feature: str) -> bool:
|
||||||
|
if feature == "fine_grained_tool_streaming":
|
||||||
|
return False
|
||||||
|
return original_supports(self, feature)
|
||||||
|
|
||||||
|
AnthropicLLM.supports_direct_anthropic_beta = patched_supports # type: ignore[method-assign]
|
||||||
|
|
||||||
|
_beta_opt_out_installed = True
|
||||||
|
logger.info("Mantle fine-grained tool streaming opt-out shim installed")
|
||||||
|
|
||||||
|
|
||||||
|
# ── Shim 4: clamp max_tokens to Mantle's output ceiling ──────────────────────
|
||||||
|
|
||||||
|
# Observed server-side clamp: Mantle streams stop at exactly 20 000 output
|
||||||
|
# tokens however large the requested max_tokens. Requesting the ceiling
|
||||||
|
# explicitly makes the model stop gracefully (proper block close + stop_reason)
|
||||||
|
# instead of the gateway cutting the stream at its own limit.
|
||||||
|
MANTLE_MAX_OUTPUT_TOKENS = 20_000
|
||||||
|
|
||||||
|
_max_tokens_clamp_installed = False
|
||||||
|
|
||||||
|
|
||||||
|
def install_max_tokens_clamp() -> None:
|
||||||
|
"""Clamp default ``maxTokens`` to Mantle's output ceiling.
|
||||||
|
|
||||||
|
Wraps ``AnthropicLLM._initialize_default_params`` so every agent's
|
||||||
|
default request params carry an explicit ``maxTokens`` no higher than the
|
||||||
|
ceiling. Per-request ``RequestParams`` overrides bypass this — none of
|
||||||
|
our agent modules set one. Safe to call more than once.
|
||||||
|
"""
|
||||||
|
global _max_tokens_clamp_installed
|
||||||
|
if _max_tokens_clamp_installed:
|
||||||
|
return
|
||||||
|
|
||||||
|
from fast_agent.llm.provider.anthropic.llm_anthropic import AnthropicLLM
|
||||||
|
|
||||||
|
original_init = AnthropicLLM._initialize_default_params # noqa: SLF001
|
||||||
|
|
||||||
|
def patched_init(self: Any, kwargs: dict) -> Any:
|
||||||
|
params = original_init(self, kwargs)
|
||||||
|
if params.maxTokens is None or params.maxTokens > MANTLE_MAX_OUTPUT_TOKENS:
|
||||||
|
params.maxTokens = MANTLE_MAX_OUTPUT_TOKENS
|
||||||
|
return params
|
||||||
|
|
||||||
|
AnthropicLLM._initialize_default_params = patched_init # noqa: SLF001
|
||||||
|
|
||||||
|
_max_tokens_clamp_installed = True
|
||||||
|
logger.info(
|
||||||
|
"Mantle max_tokens clamp shim installed (ceiling %d)",
|
||||||
|
MANTLE_MAX_OUTPUT_TOKENS,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
# ── Orchestrator ─────────────────────────────────────────────────────────────
|
# ── Orchestrator ─────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
def install_all() -> None:
|
def install_all() -> None:
|
||||||
"""Install all Mantle shims. Call once at process startup."""
|
"""Install all Mantle shims. Call once at process startup.
|
||||||
|
|
||||||
|
``install_max_tokens_clamp`` is deliberately NOT installed. Its 20 000
|
||||||
|
ceiling was an empirical observation, never a documented Mantle limit,
|
||||||
|
and re-investigation could not establish what enforces it: fast-agent
|
||||||
|
carries no 20 000 default anywhere, no model overlay is configured, and
|
||||||
|
``ModelDatabase`` reports ``max_output_tokens=128000`` for opus-4-8.
|
||||||
|
Hardcoding the constant would cement a ceiling we cannot prove and would
|
||||||
|
silently truncate turns that might otherwise complete. The shim is kept
|
||||||
|
below so it can be re-enabled if the limit is ever confirmed.
|
||||||
|
|
||||||
|
The fine-grained-tool-streaming opt-out is what actually fixes the
|
||||||
|
``Streaming completed but tool call never finished`` crash loop: it costs
|
||||||
|
no output length, it only lets a cutoff close its blocks properly and
|
||||||
|
land in fast-agent's graceful ``stop_reason=max_tokens`` handling.
|
||||||
|
"""
|
||||||
install_wire_name_prefix()
|
install_wire_name_prefix()
|
||||||
install_tool_use_caller_strip()
|
install_tool_use_caller_strip()
|
||||||
|
install_fine_grained_tool_streaming_opt_out()
|
||||||
|
|
||||||
|
|
||||||
def maybe_install(anthropic_base_url: str | None) -> bool:
|
def maybe_install(anthropic_base_url: str | None) -> bool:
|
||||||
|
|||||||
@@ -22,11 +22,13 @@ import yaml
|
|||||||
from prometheus_client import CONTENT_TYPE_LATEST, generate_latest
|
from prometheus_client import CONTENT_TYPE_LATEST, generate_latest
|
||||||
from starlette.applications import Starlette
|
from starlette.applications import Starlette
|
||||||
|
|
||||||
from pallas.metrics import REGISTRY as _metrics_registry, set_agent_info
|
|
||||||
from starlette.requests import Request
|
from starlette.requests import Request
|
||||||
from starlette.responses import JSONResponse, PlainTextResponse, Response
|
from starlette.responses import JSONResponse, PlainTextResponse, Response
|
||||||
from starlette.routing import Route
|
from starlette.routing import Route
|
||||||
|
|
||||||
|
from pallas.metrics import REGISTRY as _metrics_registry, set_agent_info
|
||||||
|
from pallas.server import DEFAULT_CONTEXT_WINDOW, DEFAULT_MAX_OUTPUT_TOKENS
|
||||||
|
|
||||||
logger = logging.getLogger(__name__)
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
|
||||||
@@ -45,37 +47,53 @@ def _load_deployment_config() -> dict:
|
|||||||
return yaml.safe_load(f) or {}
|
return yaml.safe_load(f) or {}
|
||||||
|
|
||||||
|
|
||||||
def _load_model_capabilities() -> dict:
|
def _load_fastagent_defaults() -> tuple[str, dict]:
|
||||||
"""Read model info and capabilities from the active fastagent.config.yaml."""
|
"""Read ``(default_model, model_capabilities)`` from fastagent.config.yaml."""
|
||||||
config_path = _config_root() / "fastagent.config.yaml"
|
config_path = _config_root() / "fastagent.config.yaml"
|
||||||
if not config_path.exists():
|
if not config_path.exists():
|
||||||
return {}
|
return "", {}
|
||||||
|
|
||||||
with open(config_path) as f:
|
with open(config_path) as f:
|
||||||
config = yaml.safe_load(f) or {}
|
config = yaml.safe_load(f) or {}
|
||||||
|
|
||||||
default_model = config.get("default_model", "")
|
return (
|
||||||
capabilities = config.get("model_capabilities", {})
|
config.get("default_model", "") or "",
|
||||||
|
config.get("model_capabilities", {}) or {},
|
||||||
if not default_model and not capabilities:
|
|
||||||
return {}
|
|
||||||
|
|
||||||
model_name = (
|
|
||||||
default_model.split(".", 1)[-1] if "." in default_model else default_model
|
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve_capabilities(
|
||||||
|
agent: dict, default_model: str, default_capabilities: dict
|
||||||
|
) -> dict | None:
|
||||||
|
"""Resolve the capabilities Pallas actually registered for one agent.
|
||||||
|
|
||||||
|
Mirrors ``server._register_unknown_models``: an agent's own ``model``
|
||||||
|
in agents.yaml wins over the global ``default_model``, its own
|
||||||
|
``model_capabilities`` block wins over the global one, and the same
|
||||||
|
defaults apply — so the registry advertises the effective values rather
|
||||||
|
than nulls. Returns ``None`` when no model is known for this agent.
|
||||||
|
"""
|
||||||
|
model_spec = agent.get("model") or default_model
|
||||||
|
if not model_spec:
|
||||||
|
return None
|
||||||
|
|
||||||
|
capabilities = agent.get("model_capabilities") or default_capabilities
|
||||||
|
model_name = model_spec.split(".", 1)[-1] if "." in model_spec else model_spec
|
||||||
|
|
||||||
return {
|
return {
|
||||||
"model": model_name or None,
|
"model": model_name,
|
||||||
"vision": capabilities.get("vision", False),
|
"vision": capabilities.get("vision", False),
|
||||||
"context_window": capabilities.get("context_window", None),
|
"context_window": capabilities.get("context_window", DEFAULT_CONTEXT_WINDOW),
|
||||||
"max_output_tokens": capabilities.get("max_output_tokens", None),
|
"max_output_tokens": capabilities.get(
|
||||||
|
"max_output_tokens", DEFAULT_MAX_OUTPUT_TOKENS
|
||||||
|
),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
def _build_registry(config: dict) -> dict:
|
def _build_registry(config: dict) -> dict:
|
||||||
"""Build the registry JSON from agents.yaml + fastagent.config.yaml."""
|
"""Build the registry JSON from agents.yaml + fastagent.config.yaml."""
|
||||||
now = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
now = datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ")
|
||||||
model_caps = _load_model_capabilities()
|
default_model, default_capabilities = _load_fastagent_defaults()
|
||||||
|
|
||||||
host = config.get("host", "localhost")
|
host = config.get("host", "localhost")
|
||||||
namespace = config.get("namespace", "")
|
namespace = config.get("namespace", "")
|
||||||
@@ -101,8 +119,9 @@ def _build_registry(config: dict) -> dict:
|
|||||||
}
|
}
|
||||||
],
|
],
|
||||||
}
|
}
|
||||||
if model_caps:
|
capabilities = _resolve_capabilities(agent, default_model, default_capabilities)
|
||||||
server_entry["capabilities"] = model_caps
|
if capabilities:
|
||||||
|
server_entry["capabilities"] = capabilities
|
||||||
|
|
||||||
entries.append(
|
entries.append(
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -24,6 +24,12 @@ from pallas.multimodal_server import MultimodalAgentMCPServer
|
|||||||
|
|
||||||
logger = logging.getLogger(__name__)
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
# Effective model capability defaults, applied when `model_capabilities` omits
|
||||||
|
# them. registry.py mirrors these so the published registry advertises what
|
||||||
|
# was actually registered with fast-agent's ModelDatabase.
|
||||||
|
DEFAULT_CONTEXT_WINDOW = 131072
|
||||||
|
DEFAULT_MAX_OUTPUT_TOKENS = 16384
|
||||||
|
|
||||||
|
|
||||||
def _config_root() -> Path:
|
def _config_root() -> Path:
|
||||||
"""Return the working directory where agents.yaml and fastagent configs live."""
|
"""Return the working directory where agents.yaml and fastagent configs live."""
|
||||||
@@ -143,8 +149,8 @@ def _register_one_model(model_spec: str, capabilities: dict) -> None:
|
|||||||
return
|
return
|
||||||
|
|
||||||
is_vision = capabilities.get("vision", False)
|
is_vision = capabilities.get("vision", False)
|
||||||
context_window = capabilities.get("context_window", 131072)
|
context_window = capabilities.get("context_window", DEFAULT_CONTEXT_WINDOW)
|
||||||
max_output_tokens = capabilities.get("max_output_tokens", 16384)
|
max_output_tokens = capabilities.get("max_output_tokens", DEFAULT_MAX_OUTPUT_TOKENS)
|
||||||
|
|
||||||
if is_vision:
|
if is_vision:
|
||||||
tokenizes = list(ModelDatabase.QWEN_MULTIMODAL)
|
tokenizes = list(ModelDatabase.QWEN_MULTIMODAL)
|
||||||
@@ -277,12 +283,26 @@ async def _start_agent(name: str, agents: dict[str, dict]) -> None:
|
|||||||
if entry.get(k) is not None
|
if entry.get(k) is not None
|
||||||
}
|
}
|
||||||
|
|
||||||
|
# The agents.yaml description is the one an operator actually wrote,
|
||||||
|
# and it is already published in the registry. Reuse it as the
|
||||||
|
# `send_message` tool description so MCP clients see something
|
||||||
|
# meaningful instead of the "Send a message to the {agent} agent"
|
||||||
|
# fallback.
|
||||||
|
#
|
||||||
|
# Precedence, per register_agent_tools: this value wins over a
|
||||||
|
# `description=` on the @fast.agent decorator, which in turn wins over
|
||||||
|
# the generic fallback. agents.yaml is the deployment's source of
|
||||||
|
# truth for agent metadata, so an operator editing it should not be
|
||||||
|
# silently overridden by a value buried in the agent module.
|
||||||
|
tool_description = entry.get("description") or None
|
||||||
|
|
||||||
server = MultimodalAgentMCPServer(
|
server = MultimodalAgentMCPServer(
|
||||||
primary_instance=primary_instance,
|
primary_instance=primary_instance,
|
||||||
create_instance=fast_instance._server_instance_factory,
|
create_instance=fast_instance._server_instance_factory,
|
||||||
dispose_instance=fast_instance._server_instance_dispose,
|
dispose_instance=fast_instance._server_instance_dispose,
|
||||||
instance_scope="request",
|
instance_scope="request",
|
||||||
server_name=f"{fast_instance.name}-MCP-Server",
|
server_name=f"{fast_instance.name}-MCP-Server",
|
||||||
|
tool_description=tool_description,
|
||||||
host="0.0.0.0",
|
host="0.0.0.0",
|
||||||
get_registry_version=fast_instance._get_registry_version,
|
get_registry_version=fast_instance._get_registry_version,
|
||||||
request_limits=request_limits,
|
request_limits=request_limits,
|
||||||
|
|||||||
@@ -114,34 +114,99 @@ def test_install_tool_use_caller_strip_is_idempotent() -> None:
|
|||||||
mantle_shims.install_tool_use_caller_strip() # must not raise or re-wrap
|
mantle_shims.install_tool_use_caller_strip() # must not raise or re-wrap
|
||||||
|
|
||||||
|
|
||||||
|
# ── install_fine_grained_tool_streaming_opt_out ──────────────────────────────
|
||||||
|
|
||||||
|
def test_fine_grained_beta_opt_out() -> None:
|
||||||
|
from fast_agent.llm.provider.anthropic.llm_anthropic import AnthropicLLM
|
||||||
|
|
||||||
|
mantle_shims.install_fine_grained_tool_streaming_opt_out()
|
||||||
|
|
||||||
|
# The patched method never touches self, so a bare object suffices.
|
||||||
|
stub = object()
|
||||||
|
assert AnthropicLLM.supports_direct_anthropic_beta(stub, "fine_grained_tool_streaming") is False
|
||||||
|
# Every other beta keeps the base-class answer (True).
|
||||||
|
assert AnthropicLLM.supports_direct_anthropic_beta(stub, "interleaved_thinking") is True
|
||||||
|
assert AnthropicLLM.supports_direct_anthropic_beta(stub, "long_context") is True
|
||||||
|
|
||||||
|
|
||||||
|
def test_fine_grained_beta_opt_out_is_idempotent() -> None:
|
||||||
|
mantle_shims.install_fine_grained_tool_streaming_opt_out()
|
||||||
|
mantle_shims.install_fine_grained_tool_streaming_opt_out() # must not re-wrap
|
||||||
|
|
||||||
|
from fast_agent.llm.provider.anthropic.llm_anthropic import AnthropicLLM
|
||||||
|
|
||||||
|
assert (
|
||||||
|
AnthropicLLM.supports_direct_anthropic_beta(object(), "fine_grained_tool_streaming")
|
||||||
|
is False
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ── install_max_tokens_clamp ─────────────────────────────────────────────────
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"initial,expected",
|
||||||
|
[
|
||||||
|
(128000, mantle_shims.MANTLE_MAX_OUTPUT_TOKENS), # over the ceiling → clamped
|
||||||
|
(None, mantle_shims.MANTLE_MAX_OUTPUT_TOKENS), # unset → pinned to ceiling
|
||||||
|
(4096, 4096), # under the ceiling → untouched
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_max_tokens_clamp(
|
||||||
|
monkeypatch: pytest.MonkeyPatch, initial: int | None, expected: int
|
||||||
|
) -> None:
|
||||||
|
from fast_agent.llm.provider.anthropic.llm_anthropic import AnthropicLLM
|
||||||
|
from fast_agent.types import RequestParams
|
||||||
|
|
||||||
|
# Stub the underlying initializer, then force a fresh wrap around it.
|
||||||
|
monkeypatch.setattr(
|
||||||
|
AnthropicLLM,
|
||||||
|
"_initialize_default_params",
|
||||||
|
lambda self, kwargs: RequestParams(maxTokens=initial),
|
||||||
|
)
|
||||||
|
monkeypatch.setattr(mantle_shims, "_max_tokens_clamp_installed", False)
|
||||||
|
mantle_shims.install_max_tokens_clamp()
|
||||||
|
|
||||||
|
params = AnthropicLLM._initialize_default_params(object(), {})
|
||||||
|
assert params.maxTokens == expected
|
||||||
|
|
||||||
|
|
||||||
|
def test_max_tokens_clamp_is_idempotent() -> None:
|
||||||
|
mantle_shims.install_max_tokens_clamp()
|
||||||
|
mantle_shims.install_max_tokens_clamp() # must not raise or re-wrap
|
||||||
|
|
||||||
|
|
||||||
# ── maybe_install ────────────────────────────────────────────────────────────
|
# ── maybe_install ────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
def test_maybe_install_installs_when_mantle(monkeypatch: pytest.MonkeyPatch) -> None:
|
_INSTALLER_NAMES = [
|
||||||
|
("install_wire_name_prefix", "wire"),
|
||||||
|
("install_tool_use_caller_strip", "tool_use"),
|
||||||
|
("install_fine_grained_tool_streaming_opt_out", "beta_opt_out"),
|
||||||
|
("install_max_tokens_clamp", "max_tokens"),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def _patch_installers(monkeypatch: pytest.MonkeyPatch) -> list[str]:
|
||||||
calls: list[str] = []
|
calls: list[str] = []
|
||||||
monkeypatch.setattr(
|
for attr, label in _INSTALLER_NAMES:
|
||||||
mantle_shims, "install_wire_name_prefix",
|
monkeypatch.setattr(
|
||||||
lambda: calls.append("wire"),
|
mantle_shims, attr,
|
||||||
)
|
lambda label=label: calls.append(label),
|
||||||
monkeypatch.setattr(
|
)
|
||||||
mantle_shims, "install_tool_use_caller_strip",
|
return calls
|
||||||
lambda: calls.append("tool_use"),
|
|
||||||
)
|
|
||||||
|
def test_maybe_install_installs_when_mantle(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||||
|
calls = _patch_installers(monkeypatch)
|
||||||
|
|
||||||
installed = mantle_shims.maybe_install("https://bedrock-mantle.us-east-1.api.aws/anthropic")
|
installed = mantle_shims.maybe_install("https://bedrock-mantle.us-east-1.api.aws/anthropic")
|
||||||
assert installed is True
|
assert installed is True
|
||||||
assert calls == ["wire", "tool_use"]
|
# "max_tokens" is intentionally absent: the 20 000 clamp is not installed
|
||||||
|
# by default because the ceiling was never confirmed. See install_all().
|
||||||
|
assert calls == ["wire", "tool_use", "beta_opt_out"]
|
||||||
|
|
||||||
|
|
||||||
def test_maybe_install_noop_for_non_mantle(monkeypatch: pytest.MonkeyPatch) -> None:
|
def test_maybe_install_noop_for_non_mantle(monkeypatch: pytest.MonkeyPatch) -> None:
|
||||||
calls: list[str] = []
|
calls = _patch_installers(monkeypatch)
|
||||||
monkeypatch.setattr(
|
|
||||||
mantle_shims, "install_wire_name_prefix",
|
|
||||||
lambda: calls.append("wire"),
|
|
||||||
)
|
|
||||||
monkeypatch.setattr(
|
|
||||||
mantle_shims, "install_tool_use_caller_strip",
|
|
||||||
lambda: calls.append("tool_use"),
|
|
||||||
)
|
|
||||||
|
|
||||||
assert mantle_shims.maybe_install("https://api.anthropic.com") is False
|
assert mantle_shims.maybe_install("https://api.anthropic.com") is False
|
||||||
assert mantle_shims.maybe_install(None) is False
|
assert mantle_shims.maybe_install(None) is False
|
||||||
|
|||||||
202
tests/test_registry.py
Normal file
202
tests/test_registry.py
Normal file
@@ -0,0 +1,202 @@
|
|||||||
|
"""Tests for pallas.registry — per-agent model capability resolution.
|
||||||
|
|
||||||
|
The registry advertises the capabilities Pallas actually registered with
|
||||||
|
fast-agent, resolved *per agent*: an agent's own ``model`` /
|
||||||
|
``model_capabilities`` in agents.yaml override the global ``default_model`` /
|
||||||
|
``model_capabilities`` in fastagent.config.yaml, and the same effective
|
||||||
|
defaults as ``server._register_one_model`` apply when a field is omitted.
|
||||||
|
|
||||||
|
Regression cover for two bugs: a single capabilities dict was previously
|
||||||
|
built from ``default_model`` and attached to *every* agent entry (so an
|
||||||
|
agent with a ``model:`` override was advertised under the wrong model), and
|
||||||
|
``context_window`` / ``max_output_tokens`` were published as ``null`` when
|
||||||
|
``model_capabilities`` was absent even though the model had been registered
|
||||||
|
with 131072 / 16384.
|
||||||
|
|
||||||
|
``_build_registry`` takes the deployment config as an argument but reads
|
||||||
|
fastagent.config.yaml from the working directory on every call, so each test
|
||||||
|
chdirs into a clean temp workspace and writes only the config it needs.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
from pallas import registry
|
||||||
|
from pallas.server import DEFAULT_CONTEXT_WINDOW, DEFAULT_MAX_OUTPUT_TOKENS
|
||||||
|
|
||||||
|
|
||||||
|
# ── Helpers ──────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def workspace(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
|
||||||
|
"""Chdir into a clean temp workspace — _build_registry reads cwd."""
|
||||||
|
monkeypatch.chdir(tmp_path)
|
||||||
|
return tmp_path
|
||||||
|
|
||||||
|
|
||||||
|
def _write_fastagent(workspace: Path, **config) -> None:
|
||||||
|
(workspace / "fastagent.config.yaml").write_text(yaml.safe_dump(config))
|
||||||
|
|
||||||
|
|
||||||
|
def _deployment(**agents) -> dict:
|
||||||
|
"""Minimal agents.yaml-shaped config; each agent needs at least a port."""
|
||||||
|
return {
|
||||||
|
"name": "test-project",
|
||||||
|
"namespace": "ca.helu.test",
|
||||||
|
"host": "test-host",
|
||||||
|
"agents": {
|
||||||
|
name: {"module": f"agents.{name}", "port": 9000 + i, **overrides}
|
||||||
|
for i, (name, overrides) in enumerate(agents.items())
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _entries(config: dict) -> dict[str, dict]:
|
||||||
|
"""Build the registry and return ``{agent slug: server entry}``."""
|
||||||
|
servers = registry._build_registry(config)["servers"]
|
||||||
|
return {e["server"]["name"].rsplit("/", 1)[-1]: e["server"] for e in servers}
|
||||||
|
|
||||||
|
|
||||||
|
def _capabilities(config: dict, agent: str = "solo") -> dict | None:
|
||||||
|
return _entries(config)[agent].get("capabilities")
|
||||||
|
|
||||||
|
|
||||||
|
# ── Model resolution ─────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_model_overrides_default_model(workspace: Path) -> None:
|
||||||
|
"""An agents.yaml ``model:`` wins over the global default_model."""
|
||||||
|
_write_fastagent(workspace, default_model="openai.global-model")
|
||||||
|
|
||||||
|
caps = _capabilities(_deployment(solo={"model": "anthropic.agent-model"}))
|
||||||
|
|
||||||
|
assert caps is not None
|
||||||
|
assert caps["model"] == "agent-model"
|
||||||
|
|
||||||
|
|
||||||
|
def test_falls_back_to_default_model(workspace: Path) -> None:
|
||||||
|
"""With no per-agent model, the global default_model is advertised."""
|
||||||
|
_write_fastagent(workspace, default_model="anthropic.claude-opus-4-7")
|
||||||
|
|
||||||
|
caps = _capabilities(_deployment(solo={}))
|
||||||
|
|
||||||
|
assert caps is not None
|
||||||
|
# Provider prefix is stripped — clients get the bare model name.
|
||||||
|
assert caps["model"] == "claude-opus-4-7"
|
||||||
|
|
||||||
|
|
||||||
|
def test_model_without_provider_prefix_passes_through(workspace: Path) -> None:
|
||||||
|
"""A default_model with no ``provider.`` prefix is emitted verbatim."""
|
||||||
|
_write_fastagent(workspace, default_model="bare-model")
|
||||||
|
|
||||||
|
assert _capabilities(_deployment(solo={}))["model"] == "bare-model"
|
||||||
|
|
||||||
|
|
||||||
|
# ── Capability resolution ────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_capabilities_override_global(workspace: Path) -> None:
|
||||||
|
"""A per-agent model_capabilities block replaces the global one."""
|
||||||
|
_write_fastagent(
|
||||||
|
workspace,
|
||||||
|
default_model="openai.global-model",
|
||||||
|
model_capabilities={"vision": False, "context_window": 200000},
|
||||||
|
)
|
||||||
|
|
||||||
|
caps = _capabilities(
|
||||||
|
_deployment(
|
||||||
|
solo={
|
||||||
|
"model": "anthropic.agent-model",
|
||||||
|
"model_capabilities": {"vision": True, "context_window": 400000},
|
||||||
|
}
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert caps["vision"] is True
|
||||||
|
assert caps["context_window"] == 400000
|
||||||
|
|
||||||
|
|
||||||
|
def test_effective_defaults_when_capabilities_absent(workspace: Path) -> None:
|
||||||
|
"""Absent model_capabilities publishes the values actually registered.
|
||||||
|
|
||||||
|
Regression: these were previously advertised as ``null`` while
|
||||||
|
``server._register_one_model`` registered 131072 / 16384.
|
||||||
|
"""
|
||||||
|
_write_fastagent(workspace, default_model="openai.some-model")
|
||||||
|
|
||||||
|
caps = _capabilities(_deployment(solo={}))
|
||||||
|
|
||||||
|
assert caps["context_window"] == DEFAULT_CONTEXT_WINDOW
|
||||||
|
assert caps["max_output_tokens"] == DEFAULT_MAX_OUTPUT_TOKENS
|
||||||
|
assert caps["vision"] is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_global_capabilities_are_published(workspace: Path) -> None:
|
||||||
|
"""The ordinary case: one global model + capabilities for every agent."""
|
||||||
|
_write_fastagent(
|
||||||
|
workspace,
|
||||||
|
default_model="anthropic.claude-opus-4-7",
|
||||||
|
model_capabilities={
|
||||||
|
"vision": True,
|
||||||
|
"context_window": 200000,
|
||||||
|
"max_output_tokens": 50000,
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
caps = _capabilities(_deployment(solo={}))
|
||||||
|
|
||||||
|
assert caps == {
|
||||||
|
"model": "claude-opus-4-7",
|
||||||
|
"vision": True,
|
||||||
|
"context_window": 200000,
|
||||||
|
"max_output_tokens": 50000,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
# ── Per-agent independence ───────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def test_agents_resolve_independently(workspace: Path) -> None:
|
||||||
|
"""Each entry gets its own capabilities — the original shared-dict bug."""
|
||||||
|
_write_fastagent(
|
||||||
|
workspace,
|
||||||
|
default_model="openai.global-model",
|
||||||
|
model_capabilities={"vision": False, "context_window": 128000},
|
||||||
|
)
|
||||||
|
|
||||||
|
entries = _entries(
|
||||||
|
_deployment(
|
||||||
|
inherits={},
|
||||||
|
overrides={
|
||||||
|
"model": "anthropic.special-model",
|
||||||
|
"model_capabilities": {"vision": True, "context_window": 500000},
|
||||||
|
},
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
assert entries["inherits"]["capabilities"]["model"] == "global-model"
|
||||||
|
assert entries["inherits"]["capabilities"]["context_window"] == 128000
|
||||||
|
assert entries["overrides"]["capabilities"]["model"] == "special-model"
|
||||||
|
assert entries["overrides"]["capabilities"]["context_window"] == 500000
|
||||||
|
|
||||||
|
|
||||||
|
# ── Omission ─────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def test_capabilities_omitted_without_any_model(workspace: Path) -> None:
|
||||||
|
"""No fastagent.config.yaml and no per-agent model → no capabilities key."""
|
||||||
|
entry = _entries(_deployment(solo={}))["solo"]
|
||||||
|
|
||||||
|
assert "capabilities" not in entry
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_model_published_without_fastagent_config(workspace: Path) -> None:
|
||||||
|
"""An agent's own model is advertised even with no fastagent.config.yaml."""
|
||||||
|
caps = _capabilities(_deployment(solo={"model": "anthropic.agent-model"}))
|
||||||
|
|
||||||
|
assert caps["model"] == "agent-model"
|
||||||
|
assert caps["context_window"] == DEFAULT_CONTEXT_WINDOW
|
||||||
88
tests/test_tool_description.py
Normal file
88
tests/test_tool_description.py
Normal file
@@ -0,0 +1,88 @@
|
|||||||
|
"""Tests for send_message tool-description resolution.
|
||||||
|
|
||||||
|
``server._start_agent`` passes the agents.yaml ``description`` through as
|
||||||
|
``tool_description``, so MCP clients see the description an operator actually
|
||||||
|
wrote instead of the generic "Send a message to the {agent} agent" fallback.
|
||||||
|
|
||||||
|
These pin the precedence implemented in
|
||||||
|
``MultimodalAgentMCPServer.register_agent_tools``:
|
||||||
|
|
||||||
|
tool_description (agents.yaml) > @fast.agent description > fallback
|
||||||
|
|
||||||
|
Constructing a real ``MultimodalAgentMCPServer`` needs a live FastAgent
|
||||||
|
instance, so these exercise the resolution expression directly — it is the
|
||||||
|
part that carries the logic, and the part that would silently regress.
|
||||||
|
"""
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
|
||||||
|
# ── Helpers ──────────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def _resolve(
|
||||||
|
tool_description: str | None,
|
||||||
|
agent_description: str | None = None,
|
||||||
|
agent_name: str = "scotty",
|
||||||
|
) -> str:
|
||||||
|
"""Mirror register_agent_tools' description resolution."""
|
||||||
|
resolved = (
|
||||||
|
tool_description.format(agent=agent_name)
|
||||||
|
if tool_description and "{agent}" in tool_description
|
||||||
|
else tool_description
|
||||||
|
)
|
||||||
|
return (
|
||||||
|
resolved
|
||||||
|
or agent_description
|
||||||
|
or f"Send a message to the {agent_name} agent"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# ── Precedence ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def test_agents_yaml_description_is_used() -> None:
|
||||||
|
"""The agents.yaml description reaches the tool, not the fallback."""
|
||||||
|
assert _resolve("Systems administration expert", None) == (
|
||||||
|
"Systems administration expert"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_agents_yaml_wins_over_decorator() -> None:
|
||||||
|
"""agents.yaml is the deployment's source of truth for agent metadata."""
|
||||||
|
assert _resolve("From agents.yaml", "From the decorator") == "From agents.yaml"
|
||||||
|
|
||||||
|
|
||||||
|
def test_decorator_used_when_no_yaml_description() -> None:
|
||||||
|
"""An agent module's own description still beats the generic fallback."""
|
||||||
|
assert _resolve(None, "From the decorator") == "From the decorator"
|
||||||
|
|
||||||
|
|
||||||
|
def test_falls_back_when_nothing_configured() -> None:
|
||||||
|
"""With neither source set, the generic fallback stands."""
|
||||||
|
assert _resolve(None, None) == "Send a message to the scotty agent"
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("empty", ["", None])
|
||||||
|
def test_empty_description_falls_through(empty: str | None) -> None:
|
||||||
|
"""An empty agents.yaml description must not shadow the other sources."""
|
||||||
|
assert _resolve(empty, "From the decorator") == "From the decorator"
|
||||||
|
|
||||||
|
|
||||||
|
# ── Templating ───────────────────────────────────────────────────────────────
|
||||||
|
|
||||||
|
|
||||||
|
def test_agent_placeholder_is_interpolated() -> None:
|
||||||
|
"""``{agent}`` in a description is substituted with the agent name."""
|
||||||
|
assert _resolve("Talk to {agent} about ops") == "Talk to scotty about ops"
|
||||||
|
|
||||||
|
|
||||||
|
def test_other_braces_are_left_alone() -> None:
|
||||||
|
"""Prose containing braces is only formatted when it holds ``{agent}``.
|
||||||
|
|
||||||
|
Guards the ``"{agent}" in ...`` check — an unconditional ``.format()``
|
||||||
|
would raise KeyError on a description mentioning JSON.
|
||||||
|
"""
|
||||||
|
description = 'Handles JSON like {"a": 1} safely'
|
||||||
|
assert _resolve(description, None) == description
|
||||||
Reference in New Issue
Block a user