15 Commits

Author SHA1 Message Date
f6e8e9bba5 🐾 fix(api): treat a NULL owner as unclaimed, not foreign
The previous commit's guard rejected any library whose owner_username didn't
match the caller — including NULL. On the live instance 12 of 16 workspace
libraries have no owner (created before owner stamping, or via the plain
/libraries/ endpoint, which never sets one), so that turned a silent false
204 into a hard 409 and made them undeletable. Strictly worse: the orphan
became permanent rather than merely unrecorded.

A NULL owner is unclaimed. Only a *different* non-null owner is foreign, and
only that returns 409.

Verified against live Neo4j on umbriel with throwaway libraries: NULL owner
deletes (204, node gone), caller-owned deletes (204, node gone), other-owned
is refused (409, node survives), absent stays 204. Test data cleaned up; the
30 real libraries were untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 18:32:23 -04:00
e10331dd44 🐾 fix(api): workspace DELETE must not report success it didn't perform
DELETE returned 204 for a library owned by another user, having deleted
nothing. The caller (Daedalus) recorded the workspace as cleaned up and
dropped its own row, while the Library survived holding its globally-unique
name — so that name could never be reused, and nothing anywhere recorded why.

An unowned library now returns 409 owner_conflict, reusing the code the create
path already emits for the same condition. Ownership stays opaque on GET (404
as before): the disclosure concern is about reads, and a delete that silently
does nothing is the worse failure.

A genuinely absent library still returns 204 — that idempotency is relied upon
and is correct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 13:03:32 -04:00
b5fba26d4c Merge pull request '🐾 fix(api): cascade collection and item deletes too' (#10) from fix/collection-item-cascade into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 4m24s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m12s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m38s
Reviewed-on: #10
2026-08-03 17:13:46 +00:00
15a517e389 🐾 fix(api): cascade collection and item deletes too
Collection and Item DELETE — both the REST endpoints and the HTML
views — called bare .delete(), leaking each collection's Items and every
item's Chunks/Images/ImageEmbeddings. New delete_collection_cascade and
delete_item_cascade in the shared library_delete service now back all
four paths. No orphan-Concept GC in either (that invariant stays
library-delete-only, matching the existing supersede path); the ingest
supersede helper in tasks.py now delegates to delete_item_cascade
instead of duplicating its Cypher.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 13:12:25 -04:00
c040e6f3dd Merge pull request '🐾 fix(api): cascade plain library delete; case-insensitive name conflict' (#9) from fix/library-api-delete-and-name-case into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 4m30s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m16s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m46s
Reviewed-on: #9
2026-08-03 16:45:48 +00:00
f1639dbdec 🐾 fix(api): cascade plain library delete; case-insensitive name conflict
DELETE /library/api/libraries/{uid}/ called bare lib.delete(), orphaning
Collections/Items/Chunks/Images and skipping Concept GC — the only delete
path not using the shared cascade. It now delegates to
delete_library_cascade like the HTML and workspace delete paths.

Name-conflict checks on both create endpoints are now case-insensitive
via find_library_by_name_ci (parameterised Cypher toLower comparison —
not neomodel iexact, which embeds the value in a regex and breaks on
names like "C++ Notes"). The Neo4j unique index is case-sensitive, so
"amazon connect" previously coexisted silently with "Amazon Connect",
confusing every name-matching client (Spelunker matches names
case-insensitively). The 409 reports the existing spelling. The plain
create also gains a UniqueProperty catch for the pre-check/save race,
mirroring the workspace path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 12:41:55 -04:00
7237468e3c Merge pull request '🐾 fix(tasks): unsupported file types fail ingest terminally, no retries' (#8) from fix/unsupported-file-type-terminal into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 4m11s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m15s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m49s
Reviewed-on: #8
2026-08-03 16:28:58 +00:00
e1f128659e 🐾 fix(tasks): unsupported file types fail ingest terminally, no retries
Unsupported file types raised a bare ValueError from the parser, which
ingest_from_daedalus treated like any transient fault — ERROR logs and
pointless Celery retries for input that can never parse. A new
UnsupportedFileTypeError (ValueError subclass, so existing callers still
catch it) is classified in the task as a terminal client-data failure:
logged at WARNING, job marked failed with reason
"unsupported_file_type", never retried.

Recovered from uncommitted work predating PR #6 (which deliberately
excluded it); test fixes on top: IngestJob's pk is id not job_id, and
.run() needs push_request() since the task persists self.request.id
into the NOT NULL celery_task_id column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 12:27:24 -04:00
acf829e5ca Merge pull request '🐾 feat(library): per-app managed_by replaces hardcoded Daedalus badge' (#7) from feature/managed-by into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 4m34s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m29s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m42s
Reviewed-on: #7
2026-08-03 01:55:16 +00:00
0a14cf00c5 🐾 feat(library): per-app managed_by replaces hardcoded Daedalus badge
Libraries created through the API are now stamped with the name of the
UserToken that created them (Library.managed_by), on both the workspace
and plain create endpoints; web-session creates stay null/unmanaged. The
idempotent workspace re-POST lazily backfills null managed_by, and a
one-off backfill_managed_by command labels pre-existing rows (workspace
inference: kairos-mail-* → Kairos, else Daedalus; Spelunker via ingest
job provenance).

UI badges and warnings now render "Managed by <app>" via
managed_by_display (inference fallback keeps legacy rows accurate before
backfill). The list scope filter becomes all/managed/unmanaged with the
old daedalus/global values aliased for bookmarks. The plain create
endpoint also gains an explicit 409 name_conflict (previously a raw
UniqueProperty 500) reporting the existing library's uid and manager.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 12:39:10 -04:00
9f5df20d2b Merge pull request '🐾 fix(parsers): rasterize SVG instead of failing ingest' (#6) from fix/svg-ingest-rasterize into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 3m16s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m10s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m56s
Reviewed-on: #6
2026-07-26 12:03:42 +00:00
224541a4ce 🐾 fix(parsers): rasterize SVG instead of failing ingest
"svg" was in both IMAGE_EXTENSIONS and PYMUPDF_EXTENSIONS, and the image
check ran first, so every SVG reached PIL.Image.open — which cannot decode
vector XML. Ingest raised UnidentifiedImageError, logged "Failed to read
image file_type=svg", and counted documents_parsed_total{status="error"}.
SVG ingest has never worked.

Rasterize to PNG instead, via PyMuPDF (already a dependency). This also
unblocks the vision stage downstream: it sends images to a vision LLM as a
data: URI, which cannot carry image/svg+xml either, so the description and
OCR fields were unreachable for SVG regardless of the parse fix.

Route SVG explicitly before both extension sets. Falling through to
PYMUPDF_EXTENSIONS would parse but not split, and a multi-page Write note
rendered as one document is either blank (no root width/height, so MuPDF
falls back to US-Letter and emits the top-left corner) or an illegible tall
strip (root sized across every stacked page). Splitting per write-page
element is correct for both formats, and since Write never rewrites
existing files both persist indefinitely.

svg_raster is a deliberate twin of daedalus's extraction/svg.py — the repos
share no common package, so the duplication is noted in both docstrings and
fixes belong in both.

lxml is a new dependency: recover=True is needed for the real files that
aren't well-formed XML, and stdlib ElementTree has no equivalent.
2026-07-26 07:41:20 -04:00
3ae5adebed Merge pull request '🐾 feat(library): add 'email' library type' (#4) from feat/email-library-type into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 4m27s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m27s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m30s
Reviewed-on: #4
2026-07-12 12:20:48 +00:00
9f20110f56 🐾 feat(library): add 'email' library type
Kairos is about to provision per-user app-managed (workspace) libraries for
synced mail, and none of the existing types fit correspondence. New entry in
LIBRARY_TYPE_DEFAULTS mirrors journal's entry-level chunking (one message ≈
one entry) with instructions tuned for email: correspondents, dates,
requests/commitments, and quoted-reply context. Registered in the Library
node choices and the API serializer choice list; forms and
load_library_types derive from the registry, and the compose init sidecar
re-runs the seeder on deploy.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 08:03:28 -04:00
d6b541636a Merge pull request '🐾 feat(settings): register kairos-mail source bucket for ingest' (#3) from feat/kairos-mail-bucket into main
All checks were successful
CVE Scan & Docker Build / security-scan (push) Successful in 4m22s
Build & Deploy Docs / build-and-deploy (push) Successful in 1m26s
CVE Scan & Docker Build / build-and-push (push) Successful in 2m53s
Reviewed-on: #3
2026-07-12 11:50:53 +00:00
27 changed files with 1748 additions and 105 deletions

View File

@@ -85,26 +85,21 @@ an explicit `when: mnemosyne_first_deploy` flag.
```bash ```bash
# Apply Django ORM migrations (PostgreSQL schema) # Apply Django ORM migrations (PostgreSQL schema)
docker compose -f /srv/mnemosyne/docker-compose.yaml run --rm app migrate docker compose run --rm app migrate
# Create Neo4j vector + full-text indexes and load library-type defaults # Create Neo4j vector + full-text indexes and load library-type defaults
docker compose -f /srv/mnemosyne/docker-compose.yaml \ docker compose run --rm app setup
run --rm app setup
# Seed the MCPSigningKey used to sign long-lived Pallas team JWTs. # Seed the MCPSigningKey used to sign long-lived Pallas team JWTs.
# --retire-other deactivates any previously-active key. The hex # --retire-other deactivates any previously-active key. The hex
# emitted to stdout is persisted in Mnemosyne's database and is # emitted to stdout is persisted in Mnemosyne's database and is
# not re-injected from the vault — no operator action required # not re-injected from the vault — no operator action required
# beyond running this command once per fresh deployment. # beyond running this command once per fresh deployment.
docker compose -f /srv/mnemosyne/docker-compose.yaml \ docker compose run --rm app python manage.py seed_signing_key --kid daedalus-1 --retire-other
run --rm app \
python manage.py seed_signing_key --kid daedalus-1 --retire-other
# Create Django groups for SSO role mapping (View Only / Staff / SME / Admin). # Create Django groups for SSO role mapping (View Only / Staff / SME / Admin).
# Safe to re-run — idempotent. # Safe to re-run — idempotent.
docker compose -f /srv/mnemosyne/docker-compose.yaml \ docker compose run --rm app python manage.py create_sso_groups
run --rm app \
python manage.py create_sso_groups
``` ```
The `seed_signing_key` command prints the generated secret once to stdout — it The `seed_signing_key` command prints the generated secret once to stdout — it

View File

@@ -15,6 +15,7 @@ LIBRARY_TYPE_CHOICES = [
"film", "film",
"art", "art",
"journal", "journal",
"email",
"business", "business",
"finance", "finance",
] ]
@@ -36,6 +37,7 @@ class LibrarySerializer(serializers.Serializer):
required=False, allow_blank=True, default="" required=False, allow_blank=True, default=""
) )
workspace_id = serializers.CharField(read_only=True) workspace_id = serializers.CharField(read_only=True)
managed_by = serializers.CharField(read_only=True)
created_at = serializers.DateTimeField(read_only=True) created_at = serializers.DateTimeField(read_only=True)
@@ -192,6 +194,7 @@ class WorkspaceStatusSerializer(serializers.Serializer):
name = serializers.CharField() name = serializers.CharField()
library_type = serializers.CharField() library_type = serializers.CharField()
description = serializers.CharField(allow_blank=True) description = serializers.CharField(allow_blank=True)
managed_by = serializers.CharField(allow_null=True, required=False)
item_count = serializers.IntegerField() item_count = serializers.IntegerField()
chunk_count = serializers.IntegerField() chunk_count = serializers.IntegerField()
created_at = serializers.DateTimeField() created_at = serializers.DateTimeField()

View File

@@ -10,6 +10,7 @@ import os
from django.core.files.base import ContentFile from django.core.files.base import ContentFile
from django.core.files.storage import default_storage from django.core.files.storage import default_storage
from neomodel.exceptions import UniqueProperty
from rest_framework import status from rest_framework import status
from rest_framework.decorators import api_view, parser_classes, permission_classes from rest_framework.decorators import api_view, parser_classes, permission_classes
from rest_framework.parsers import FormParser, JSONParser, MultiPartParser from rest_framework.parsers import FormParser, JSONParser, MultiPartParser
@@ -17,6 +18,12 @@ from rest_framework.permissions import IsAuthenticated
from rest_framework.response import Response from rest_framework.response import Response
from library.content_types import get_library_type_config from library.content_types import get_library_type_config
from library.services.library_delete import (
delete_collection_cascade,
delete_item_cascade,
delete_library_cascade,
)
from mcp_server.drf_auth import request_token_label
from .serializers import ( from .serializers import (
CollectionSerializer, CollectionSerializer,
@@ -50,7 +57,7 @@ def library_list_create(request):
a per-library ``item_count``. Off by default because the count is a a per-library ``item_count``. Off by default because the count is a
Cypher aggregate; on for the Daedalus-side registry poll. Cypher aggregate; on for the Daedalus-side registry poll.
""" """
from library.models import Library from library.models import Library, find_library_by_name_ci
if request.method == "GET": if request.method == "GET":
include_workspace = request.GET.get("include_workspace", "true").lower() != "false" include_workspace = request.GET.get("include_workspace", "true").lower() != "false"
@@ -84,6 +91,28 @@ def library_list_create(request):
serializer.is_valid(raise_exception=True) serializer.is_valid(raise_exception=True)
data = serializer.validated_data data = serializer.validated_data
# Library names are unique. The Neo4j index is case-sensitive, but
# clients (Spelunker, humans) treat names case-insensitively, so the
# create-time check is case-insensitive too — otherwise "amazon connect"
# silently creates a near-duplicate of "Amazon Connect". Reject with a
# clean 409 (uid + managed_by included so the caller can say who owns
# the name) instead of letting the unique-index save raise a 500.
existing = find_library_by_name_ci(data["name"])
if existing is not None:
logger.warning(
"library_create name_conflict name=%s existing_uid=%s caller=%s",
data["name"], existing.uid, request.user.username,
)
return Response(
{
"detail": f"A library named '{existing.name}' already exists.",
"code": "name_conflict",
"uid": existing.uid,
"managed_by": existing.managed_by_display or None,
},
status=status.HTTP_409_CONFLICT,
)
# Populate defaults from content-type config if not provided # Populate defaults from content-type config if not provided
library_type = data["library_type"] library_type = data["library_type"]
defaults = get_library_type_config(library_type) defaults = get_library_type_config(library_type)
@@ -92,6 +121,7 @@ def library_list_create(request):
name=data["name"], name=data["name"],
library_type=library_type, library_type=library_type,
description=data.get("description", ""), description=data.get("description", ""),
managed_by=request_token_label(request),
chunking_config=data.get("chunking_config") or defaults["chunking_config"], chunking_config=data.get("chunking_config") or defaults["chunking_config"],
embedding_instruction=( embedding_instruction=(
data.get("embedding_instruction") or defaults["embedding_instruction"] data.get("embedding_instruction") or defaults["embedding_instruction"]
@@ -103,7 +133,22 @@ def library_list_create(request):
data.get("llm_context_prompt") or defaults["llm_context_prompt"] data.get("llm_context_prompt") or defaults["llm_context_prompt"]
), ),
) )
lib.save() try:
lib.save()
except UniqueProperty:
# Race between the pre-check and save — an exact-name twin landed
# in between. Same 409 shape, minus the loser's uid/manager.
logger.warning(
"library_create name_conflict (save race) name=%s caller=%s",
data["name"], request.user.username,
)
return Response(
{
"detail": f"A library named '{data['name']}' already exists.",
"code": "name_conflict",
},
status=status.HTTP_409_CONFLICT,
)
return Response(LibrarySerializer(lib).data, status=status.HTTP_201_CREATED) return Response(LibrarySerializer(lib).data, status=status.HTTP_201_CREATED)
@@ -141,8 +186,16 @@ def library_detail(request, uid):
lib.save() lib.save()
return Response(LibrarySerializer(lib).data) return Response(LibrarySerializer(lib).data)
# DELETE # DELETE — use the shared cascade so child nodes (Collections/Items/
lib.delete() # Chunks/Images) and orphan Concepts are removed too; a bare
# lib.delete() would leak them all.
result = delete_library_cascade(lib)
logger.info(
"Library deleted via API library_uid=%s name=%s items=%d "
"orphans_deleted=%d caller=%s",
result["library_uid"], result["name"], result["item_count"],
result["orphans_deleted"], request.user.username,
)
return Response(status=status.HTTP_204_NO_CONTENT) return Response(status=status.HTTP_204_NO_CONTENT)
@@ -224,7 +277,13 @@ def collection_detail(request, uid):
col.save() col.save()
return Response(CollectionSerializer(col).data) return Response(CollectionSerializer(col).data)
col.delete() # DELETE — cascade Items/Chunks/Images too; a bare col.delete() leaks them.
result = delete_collection_cascade(col)
logger.info(
"Collection deleted via API collection_uid=%s name=%s items=%d caller=%s",
result["collection_uid"], result["name"], result["item_count"],
request.user.username,
)
return Response(status=status.HTTP_204_NO_CONTENT) return Response(status=status.HTTP_204_NO_CONTENT)
@@ -305,7 +364,13 @@ def item_detail(request, uid):
item.save() item.save()
return Response(ItemSerializer(item).data) return Response(ItemSerializer(item).data)
item.delete() # DELETE — cascade Chunks/Images/embeddings too; a bare item.delete()
# leaks them.
delete_item_cascade(item.uid)
logger.info(
"Item deleted via API item_uid=%s caller=%s",
item.uid, request.user.username,
)
return Response(status=status.HTTP_204_NO_CONTENT) return Response(status=status.HTTP_204_NO_CONTENT)

View File

@@ -25,6 +25,7 @@ from rest_framework.response import Response
from library.content_types import get_library_type_config from library.content_types import get_library_type_config
from library.services.library_delete import delete_library_cascade from library.services.library_delete import delete_library_cascade
from mcp_server.drf_auth import request_token_label
from .serializers import WorkspaceCreateSerializer, WorkspaceStatusSerializer from .serializers import WorkspaceCreateSerializer, WorkspaceStatusSerializer
@@ -49,6 +50,7 @@ def _serialize_workspace(lib):
"name": lib.name, "name": lib.name,
"library_type": lib.library_type, "library_type": lib.library_type,
"description": lib.description or "", "description": lib.description or "",
"managed_by": lib.managed_by,
"item_count": item_count, "item_count": item_count,
"chunk_count": chunk_count, "chunk_count": chunk_count,
"created_at": lib.created_at, "created_at": lib.created_at,
@@ -65,7 +67,7 @@ def workspace_create(request):
workspace (200) — not an error. The library_type is frozen at first workspace (200) — not an error. The library_type is frozen at first
create; subsequent calls are not allowed to change it. create; subsequent calls are not allowed to change it.
""" """
from library.models import Library from library.models import Library, find_library_by_name_ci
serializer = WorkspaceCreateSerializer(data=request.data) serializer = WorkspaceCreateSerializer(data=request.data)
serializer.is_valid(raise_exception=True) serializer.is_valid(raise_exception=True)
@@ -104,6 +106,18 @@ def workspace_create(request):
}, },
status=status.HTTP_409_CONFLICT, status=status.HTTP_409_CONFLICT,
) )
# Lazy backfill: pre-managed_by libraries pick up the label from
# the first idempotent re-POST. Null-only — an already-stamped
# library never changes manager.
if not existing.managed_by:
label = request_token_label(request)
if label:
existing.managed_by = label
existing.save()
logger.info(
"Backfilled managed_by=%s workspace_id=%s library_uid=%s",
label, existing.workspace_id, existing.uid,
)
logger.info( logger.info(
"Workspace already exists workspace_id=%s library_uid=%s", "Workspace already exists workspace_id=%s library_uid=%s",
data["workspace_id"], existing.uid, data["workspace_id"], existing.uid,
@@ -113,6 +127,28 @@ def workspace_create(request):
status=status.HTTP_200_OK, status=status.HTTP_200_OK,
) )
# New workspace: reject a name already taken by any other library,
# case-insensitively — the Neo4j index is case-sensitive, so without
# this check "amazon connect" would silently coexist with an existing
# "Amazon Connect" and confuse every name-matching client.
name_taken = find_library_by_name_ci(data["name"])
if name_taken is not None:
logger.warning(
"workspace_create name_conflict workspace_id=%s name=%s "
"existing_uid=%s",
data["workspace_id"], data["name"], name_taken.uid,
)
return Response(
{
"detail": (
f"A library named '{name_taken.name}' already exists in "
"Mnemosyne."
),
"code": "name_conflict",
},
status=status.HTTP_409_CONFLICT,
)
defaults = get_library_type_config(data["library_type"]) defaults = get_library_type_config(data["library_type"])
lib = Library( lib = Library(
name=data["name"], name=data["name"],
@@ -120,6 +156,7 @@ def workspace_create(request):
description=data.get("description", ""), description=data.get("description", ""),
workspace_id=data["workspace_id"], workspace_id=data["workspace_id"],
owner_username=request.user.username, owner_username=request.user.username,
managed_by=request_token_label(request),
chunking_config=defaults["chunking_config"], chunking_config=defaults["chunking_config"],
embedding_instruction=defaults["embedding_instruction"], embedding_instruction=defaults["embedding_instruction"],
reranker_instruction=defaults["reranker_instruction"], reranker_instruction=defaults["reranker_instruction"],
@@ -176,23 +213,50 @@ def workspace_detail_or_delete(request, workspace_id):
except Library.DoesNotExist: except Library.DoesNotExist:
lib = None lib = None
# Cross-user reads/writes look like "not found" — don't disclose # A NULL owner_username is *unclaimed*, not someone else's: workspace
# existence across users. # libraries created before owner stamping (and any created through the
if lib is not None and lib.owner_username != request.user.username: # plain /libraries/ endpoint, which never sets it) have no owner. Treating
lib = None # those as foreign would make them permanently undeletable through this
# endpoint — the very orphans this view is trying not to create.
#
# Cross-user reads still look like "not found"; DELETE is handled
# separately below, because telling the caller "deleted" about a library
# we did not touch is what silently orphans libraries, and an orphan holds
# its globally-unique name forever.
foreign = (
lib is not None
and lib.owner_username
and lib.owner_username != request.user.username
)
if request.method == "GET": if request.method == "GET":
if lib is None: if lib is None or foreign:
return Response( return Response(
{"detail": "Workspace not found."}, {"detail": "Workspace not found."},
status=status.HTTP_404_NOT_FOUND, status=status.HTTP_404_NOT_FOUND,
) )
return Response(WorkspaceStatusSerializer(_serialize_workspace(lib)).data) return Response(WorkspaceStatusSerializer(_serialize_workspace(lib)).data)
# DELETE — idempotent: a missing (or unowned) workspace returns 204. # DELETE — idempotent only where nothing exists to delete.
if lib is None: if lib is None:
return Response(status=status.HTTP_204_NO_CONTENT) return Response(status=status.HTTP_204_NO_CONTENT)
if foreign:
# Still opaque about ownership, but never a false success: the
# caller must not record this workspace as cleaned up.
logger.warning(
"workspace_delete owner_conflict workspace_id=%s library_uid=%s "
"caller=%s",
workspace_id, lib.uid, request.user.username,
)
return Response(
{
"detail": "Workspace id is already in use.",
"code": "owner_conflict",
},
status=status.HTTP_409_CONFLICT,
)
# Delete the Library and everything reachable + unique to it, plus # Delete the Library and everything reachable + unique to it, plus
# orphan-Concept GC. Shared with the admin/HTML delete path. # orphan-Concept GC. Shared with the admin/HTML delete path.
result = delete_library_cascade(lib) result = delete_library_cascade(lib)

View File

@@ -241,6 +241,38 @@ LIBRARY_TYPE_DEFAULTS = {
"4) The commercial purpose — positioning, pricing, capability demonstration." "4) The commercial purpose — positioning, pricing, capability demonstration."
), ),
}, },
"email": {
"chunking_config": {
"strategy": "entry_level",
"chunk_size": 512,
"chunk_overlap": 32,
"respect_boundaries": ["message", "quote", "paragraph"],
},
"embedding_instruction": (
"Represent this email message for retrieval. "
"Focus on the sender, recipients, subject, dates, requests and "
"commitments made, and the people, organizations, and events discussed."
),
"reranker_instruction": (
"Re-rank email messages based on relevance to the query. "
"Prioritize messages matching the correspondents, subject matter, "
"time period, and any specific commitments or requests mentioned."
),
"llm_context_prompt": (
"The following excerpts are from personal email correspondence. "
"This is private content — answer with discretion. Attribute "
"statements to their senders, note dates, and distinguish what was "
"asked from what was agreed. Quoted text below a reply is earlier "
"context, not the sender's own words."
),
"vision_prompt": (
"Analyze this image from an email message. Identify:\n"
"1) Image type (photograph, screenshot, scanned document, chart, signature graphic).\n"
"2) What it depicts — people, places, documents, data.\n"
"3) Any visible text, dates, or figures.\n"
"4) Its role in the message — attachment content, inline illustration, or boilerplate."
),
},
"finance": { "finance": {
"chunking_config": { "chunking_config": {
"strategy": "section_aware", "strategy": "section_aware",
@@ -282,7 +314,7 @@ def get_library_type_config(library_type):
Args: Args:
library_type: One of 'fiction', 'nonfiction', 'technical', 'music', library_type: One of 'fiction', 'nonfiction', 'technical', 'music',
'film', 'art', 'journal', 'business', 'finance' 'film', 'art', 'journal', 'email', 'business', 'finance'
Returns: Returns:
dict with keys: chunking_config, embedding_instruction, dict with keys: chunking_config, embedding_instruction,

View File

@@ -0,0 +1,111 @@
"""One-off backfill of ``Library.managed_by`` for pre-existing libraries.
``managed_by`` is stamped from the creating API token's name, so
libraries created before the property existed have it null. This
command labels them:
* Workspace-scoped libraries get :func:`infer_legacy_manager`'s answer —
``Kairos`` for ``kairos-mail-*`` workspace ids, ``Daedalus`` otherwise.
* Global libraries that Spelunker ingested into (any ``IngestJob`` with
``source="spelunker"``) get the Spelunker label.
* Everything else stays null (hand-made in the web UI).
Label flags let the operator match the *actual* production token names
so backfilled rows render identically to newly stamped ones.
Idempotent: only null ``managed_by`` is ever written, and the lazy fill
in ``workspace_create`` is also null-only, so re-runs and later API
traffic never overwrite these labels.
"""
from __future__ import annotations
import logging
from django.core.management.base import BaseCommand, CommandError
from library.models import IngestJob
logger = logging.getLogger(__name__)
class Command(BaseCommand):
help = (
"Backfill Library.managed_by for libraries created before the "
"property existed. Workspace libraries are labelled by inference "
"(kairos-mail-* → Kairos, else Daedalus); global libraries with "
"Spelunker ingest jobs get the Spelunker label."
)
def add_arguments(self, parser):
parser.add_argument(
"--daedalus-label", default="Daedalus",
help="Label for non-Kairos workspace libraries (default: Daedalus).",
)
parser.add_argument(
"--kairos-label", default="Kairos",
help="Label for kairos-mail-* workspace libraries (default: Kairos).",
)
parser.add_argument(
"--spelunker-label", default="Spelunker",
help="Label for global libraries with Spelunker ingest jobs "
"(default: Spelunker).",
)
parser.add_argument(
"--dry-run", action="store_true",
help="Report what would be labelled, don't persist.",
)
def handle(self, *args, **options):
try:
from library.models import Library, infer_legacy_manager
except Exception as exc: # pragma: no cover
raise CommandError(
f"Could not import library.models.Library (Neo4j unreachable?): {exc}"
) from exc
overrides = {
"Daedalus": options["daedalus_label"],
"Kairos": options["kairos_label"],
}
spelunker_uids = set(
IngestJob.objects
.filter(source="spelunker")
.values_list("library_uid", flat=True)
.distinct()
)
candidates = list(Library.nodes.filter(managed_by__isnull=True))
to_label = []
for lib in candidates:
label = infer_legacy_manager(lib.workspace_id)
if label:
label = overrides[label]
elif lib.uid in spelunker_uids:
label = options["spelunker_label"]
if label:
to_label.append((lib, label))
self.stdout.write(f"Libraries with null managed_by: {len(candidates)}")
self.stdout.write(
self.style.SUCCESS(f"Will label: {len(to_label)} "
f"(unmanaged, left null: {len(candidates) - len(to_label)})")
)
for lib, label in to_label:
self.stdout.write(f" {lib.uid} {lib.name!r}{label}")
if options["dry_run"]:
self.stdout.write(self.style.WARNING("--dry-run: nothing written."))
return
for lib, label in to_label:
lib.managed_by = label
lib.save()
logger.info(
"backfill_managed_by uid=%s name=%s label=%s",
lib.uid, lib.name, label,
)
self.stdout.write(self.style.SUCCESS("Done."))

View File

@@ -51,6 +51,35 @@ class NearbyImageRel(StructuredRel):
# --- Node models --- # --- Node models ---
def infer_legacy_manager(workspace_id):
"""Manager label for pre-``managed_by`` workspace libraries, else None.
Kairos mail workspaces are recognisable by their deterministic
``kairos-mail-`` id prefix; every other workspace id is a Daedalus
workspace UUID.
"""
if not workspace_id:
return None
return "Kairos" if workspace_id.startswith("kairos-mail-") else "Daedalus"
def find_library_by_name_ci(name):
"""Case-insensitively find a Library by name, or None.
Parameterised Cypher rather than neomodel's ``iexact``, which embeds
the value in a regex and so breaks on names containing regex
metacharacters (e.g. "C++ Notes").
"""
from neomodel import db
rows, _ = db.cypher_query(
"MATCH (l:Library) WHERE toLower(l.name) = toLower($name) "
"RETURN l LIMIT 1",
{"name": name},
)
return Library.inflate(rows[0][0]) if rows else None
class Library(StructuredNode): class Library(StructuredNode):
""" """
Top-level container representing a content library. Top-level container representing a content library.
@@ -63,6 +92,11 @@ class Library(StructuredNode):
across the whole instance) or *workspace-scoped* (workspace_id set — across the whole instance) or *workspace-scoped* (workspace_id set —
visible only to agents inside that Daedalus workspace). Scoping is visible only to agents inside that Daedalus workspace). Scoping is
enforced structurally by every search query. enforced structurally by every search query.
Independently of scoping, a library may be *app-managed*
(``managed_by`` set — created through the API by an external app such
as Daedalus, Kairos, or Spelunker, which owns its content lifecycle)
or unmanaged (created by hand in the web UI).
""" """
uid = UniqueIdProperty() uid = UniqueIdProperty()
@@ -77,6 +111,7 @@ class Library(StructuredNode):
"film": "Film", "film": "Film",
"art": "Art", "art": "Art",
"journal": "Journal", "journal": "Journal",
"email": "Email",
"business": "Business", "business": "Business",
"finance": "Finance", "finance": "Finance",
}, },
@@ -92,6 +127,12 @@ class Library(StructuredNode):
# this user. Null for global libraries. # this user. Null for global libraries.
owner_username = StringProperty(required=False, index=True) owner_username = StringProperty(required=False, index=True)
# Name of the API token that created this library ("Daedalus",
# "Kairos", "Spelunker", ...). Null for libraries created in the
# web UI. Stamped at create time only — token rotation or edits by
# another token never change it.
managed_by = StringProperty(required=False, index=True)
# Content-type configuration # Content-type configuration
chunking_config = JSONProperty(default={}) chunking_config = JSONProperty(default={})
embedding_instruction = StringProperty(default="") embedding_instruction = StringProperty(default="")
@@ -103,6 +144,16 @@ class Library(StructuredNode):
# Relationships # Relationships
collections = RelationshipTo("Collection", "CONTAINS") collections = RelationshipTo("Collection", "CONTAINS")
@property
def managed_by_display(self):
"""Managing-app label for display; empty string when unmanaged.
Falls back to inference for workspace libraries created before
``managed_by`` existed, so rendering is identical before and
after the backfill command runs.
"""
return self.managed_by or infer_legacy_manager(self.workspace_id) or ""
def __str__(self): def __str__(self):
return f"{self.name} ({self.library_type})" return f"{self.name} ({self.library_type})"

View File

@@ -106,3 +106,99 @@ def delete_library_cascade(lib) -> dict:
"item_s3_keys": item_s3_keys, "item_s3_keys": item_s3_keys,
"orphans_deleted": orphans_deleted, "orphans_deleted": orphans_deleted,
} }
def delete_collection_cascade(col) -> dict:
"""Delete ``col`` and all content reachable and unique to it.
Removes the Collection's Items with their Chunks, Images, and
ImageEmbeddings, then the Collection itself. No orphan-Concept GC —
that invariant belongs to library-level deletes only (see
:func:`delete_library_cascade`); Concepts orphaned here are collected
on the next library delete.
:param col: A ``library.models.Collection`` node instance.
:returns: Dict with ``collection_uid``, ``name``, ``item_count``, and
``item_s3_keys`` (list of ``(uid, s3_key)`` for async S3 cleanup).
"""
collection_uid = col.uid
collection_name = col.name
s3_rows, _ = db.cypher_query(
"MATCH (col:Collection {uid: $uid})-[:CONTAINS]->(i:Item) "
"RETURN i.uid, i.s3_key",
{"uid": collection_uid},
)
item_s3_keys = [(r[0], r[1]) for r in s3_rows if r[1]]
db.cypher_query(
"""
MATCH (col:Collection {uid: $uid})-[:CONTAINS]->(i:Item)
-[:HAS_CHUNK]->(c:Chunk)
DETACH DELETE c
""",
{"uid": collection_uid},
)
db.cypher_query(
"""
MATCH (col:Collection {uid: $uid})-[:CONTAINS]->(i:Item)
-[:HAS_IMAGE]->(img:Image)
OPTIONAL MATCH (img)-[:HAS_EMBEDDING]->(emb:ImageEmbedding)
DETACH DELETE img, emb
""",
{"uid": collection_uid},
)
db.cypher_query(
"""
MATCH (col:Collection {uid: $uid})-[:CONTAINS]->(i:Item)
DETACH DELETE i
""",
{"uid": collection_uid},
)
db.cypher_query(
"MATCH (col:Collection {uid: $uid}) DETACH DELETE col",
{"uid": collection_uid},
)
logger.info(
"Collection cascade-deleted collection_uid=%s name=%s items=%d",
collection_uid, collection_name, len(item_s3_keys),
)
return {
"collection_uid": collection_uid,
"name": collection_name,
"item_count": len(item_s3_keys),
"item_s3_keys": item_s3_keys,
}
def delete_item_cascade(item_uid: str) -> dict:
"""Delete the Item ``item_uid`` with its Chunks, Images, and embeddings.
Keyed on the uid (not a node instance) so the ingest supersede path in
``library.tasks`` can share it. No orphan-Concept GC — see
:func:`delete_collection_cascade`.
:returns: Dict with ``item_uid`` and ``s3_key`` (empty string when the
item had no stored file) for async S3 cleanup.
"""
s3_rows, _ = db.cypher_query(
"MATCH (i:Item {uid: $uid}) RETURN i.s3_key",
{"uid": item_uid},
)
s3_key = (s3_rows[0][0] or "") if s3_rows else ""
db.cypher_query(
"""
MATCH (i:Item {uid: $uid})
OPTIONAL MATCH (i)-[:HAS_CHUNK]->(c:Chunk)
OPTIONAL MATCH (i)-[:HAS_IMAGE]->(img:Image)
OPTIONAL MATCH (img)-[:HAS_EMBEDDING]->(emb:ImageEmbedding)
DETACH DELETE c, img, emb, i
""",
{"uid": item_uid},
)
logger.info("Item cascade-deleted item_uid=%s", item_uid)
return {"item_uid": item_uid, "s3_key": s3_key}

View File

@@ -22,6 +22,17 @@ from .text_utils import remove_excessive_whitespace, sanitize_text
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
class UnsupportedFileTypeError(ValueError):
"""Raised when a file's type cannot be parsed.
Deterministic and input-driven — re-parsing identical bytes can never
succeed — so ingest must treat it as a terminal failure and never retry.
Subclasses ``ValueError`` so existing ``except ValueError`` callers still
catch it.
"""
# File extensions supported by PyMuPDF # File extensions supported by PyMuPDF
PYMUPDF_EXTENSIONS = { PYMUPDF_EXTENSIONS = {
"pdf", "epub", "xps", "mobi", "fb2", "cbz", "svg", "pdf", "epub", "xps", "mobi", "fb2", "cbz", "svg",
@@ -31,8 +42,11 @@ PYMUPDF_EXTENSIONS = {
# Plain text extensions — read directly, no PyMuPDF needed # Plain text extensions — read directly, no PyMuPDF needed
PLAINTEXT_EXTENSIONS = {"txt", "md", "csv", "tsv", "log", "json", "yaml", "yml", "xml"} PLAINTEXT_EXTENSIONS = {"txt", "md", "csv", "tsv", "log", "json", "yaml", "yml", "xml"}
# Image extensions — store as Image nodes directly # Image extensions — store as Image nodes directly.
IMAGE_EXTENSIONS = {"jpg", "jpeg", "png", "gif", "bmp", "tiff", "tif", "webp", "svg"} # SVG is deliberately absent: it is vector XML that Pillow cannot decode and
# that the vision stage cannot send as a data URI, so it gets rasterized by
# _parse_svg_file instead.
IMAGE_EXTENSIONS = {"jpg", "jpeg", "png", "gif", "bmp", "tiff", "tif", "webp"}
# Minimum image dimensions to extract (skip tiny icons/bullets) # Minimum image dimensions to extract (skip tiny icons/bullets)
MIN_IMAGE_WIDTH = 50 MIN_IMAGE_WIDTH = 50
@@ -85,7 +99,7 @@ class DocumentParser:
:param file_path: Path to the document file. :param file_path: Path to the document file.
:param file_type: File extension (without dot), e.g. 'pdf', 'epub'. :param file_type: File extension (without dot), e.g. 'pdf', 'epub'.
:returns: ParseResult with text blocks, images, and metadata. :returns: ParseResult with text blocks, images, and metadata.
:raises ValueError: If the file type is not supported. :raises UnsupportedFileTypeError: If the file type is not supported.
""" """
file_type = file_type.lower().lstrip(".") file_type = file_type.lower().lstrip(".")
@@ -98,6 +112,12 @@ class DocumentParser:
if file_type in PLAINTEXT_EXTENSIONS: if file_type in PLAINTEXT_EXTENSIONS:
return self._parse_plaintext(file_path, file_type) return self._parse_plaintext(file_path, file_type)
# Checked before PYMUPDF_EXTENSIONS: PyMuPDF can open an SVG, but
# rendering a multi-page Write note as one document yields a blank or
# illegible image (see svg_raster), so it needs page-aware handling.
if file_type == "svg":
return self._parse_svg_file(file_path, file_type)
if file_type in IMAGE_EXTENSIONS: if file_type in IMAGE_EXTENSIONS:
return self._parse_image_file(file_path, file_type) return self._parse_image_file(file_path, file_type)
@@ -108,7 +128,7 @@ class DocumentParser:
if file_type in ("html", "htm"): if file_type in ("html", "htm"):
return self._parse_with_pymupdf(file_path, file_type) return self._parse_with_pymupdf(file_path, file_type)
raise ValueError( raise UnsupportedFileTypeError(
f"Unsupported file type '{file_type}'. " f"Unsupported file type '{file_type}'. "
f"Supported: {sorted(PYMUPDF_EXTENSIONS | PLAINTEXT_EXTENSIONS | IMAGE_EXTENSIONS)}" f"Supported: {sorted(PYMUPDF_EXTENSIONS | PLAINTEXT_EXTENSIONS | IMAGE_EXTENSIONS)}"
) )
@@ -309,6 +329,61 @@ class DocumentParser:
file_type=file_type, file_type=file_type,
) )
def _parse_svg_file(self, file_path: str, file_type: str) -> ParseResult:
"""
Rasterize an SVG into one ExtractedImage per page.
SVG is vector XML: Pillow cannot decode it and the vision stage cannot
put it in a data URI, so it is rendered to PNG here. Multi-page Write
notes become one image per page so each page reaches vision/OCR at a
legible size.
:param file_path: Path to the SVG file.
:param file_type: Normalized file extension ("svg").
:returns: ParseResult with one image per rendered page.
"""
with DOCUMENT_PARSE_DURATION.labels(file_type=file_type).time():
try:
from library.services.svg_raster import render_svg_pages
with open(file_path, "rb") as f:
data = f.read()
pages = render_svg_pages(data)
except Exception as exc:
DOCUMENTS_PARSED_TOTAL.labels(file_type=file_type, status="error").inc()
logger.error("Failed to rasterize SVG file_type=%s: %s", file_type, exc)
raise
images = [
ExtractedImage(
data=png,
ext="png",
width=width,
height=height,
source_page=index,
source_index=0,
)
for index, (png, width, height) in enumerate(pages)
]
DOCUMENTS_PARSED_TOTAL.labels(file_type=file_type, status="success").inc()
IMAGES_EXTRACTED_TOTAL.labels(file_type=file_type).inc(len(images))
logger.info(
"Parsed SVG file_type=%s pages=%d bytes=%d",
file_type,
len(images),
len(data),
)
return ParseResult(
text_blocks=[],
images=images,
metadata={"page_count": len(images)},
file_type=file_type,
)
def _parse_image_file(self, file_path: str, file_type: str) -> ParseResult: def _parse_image_file(self, file_path: str, file_type: str) -> ParseResult:
""" """
Handle a standalone image file — store as a single ExtractedImage. Handle a standalone image file — store as a single ExtractedImage.

View File

@@ -0,0 +1,263 @@
"""SVG rasterization for the ingest pipeline.
The vision stage sends each extracted image to a vision LLM as a ``data:`` URI,
which cannot carry ``image/svg+xml`` — so an SVG has to become raster before it
can be described, OCR'd, or embedded. PyMuPDF (already a dependency for PDF
parsing) renders SVG natively, so this needs no cairo/rsvg system libraries.
The non-obvious part is page splitting. Handwritten notes from the Write
(Stylus Labs) app are a single SVG document holding one or more
``<svg class="write-page">`` children stacked vertically via x/y offsets. Two
root formats exist in the wild and *both* must be split per page:
- Older files carry no width/height on the root ``<svg>``. Rendering the
document as-is makes PyMuPDF fall back to US-Letter and emit the top-left
corner only — ruled lines and no handwriting, in a perfectly valid PNG.
- Newer files (Write commit eeab021) do carry root width/height, spanning the
full stacked extent. Rendering those as-is yields one very tall strip; capped
to a sane longest edge, a 5-page note squashes to ~209px wide, well past
illegible.
Write never rewrites existing files, so the old format is permanent rather than
a migration window. Page geometry lives on the ``write-page`` element and is
identical across both formats, so splitting ignores the root dimensions
entirely and needs no format detection.
.. note::
Daedalus carries a twin of this module at
``backend/daedalus/extraction/svg.py``, which returns base64 for direct
chat attachment. The two are deliberately duplicated rather than shared —
the repos ship no common package — so fixes belong in both.
"""
from __future__ import annotations
import copy
import logging
import re
from lxml import etree
logger = logging.getLogger(__name__)
SVG_NS = "http://www.w3.org/2000/svg"
XLINK_NS = "http://www.w3.org/1999/xlink"
_SVG = f"{{{SVG_NS}}}"
_PAGE_CLASS = "write-page"
# Write documents render on a grey backdrop; pages themselves are transparent,
# so each rendered page needs an explicit white underlay or strokes land on
# black in the flattened PNG.
_PAGE_BACKGROUND = "#ffffff"
#: Longest edge of a rendered page, in pixels. Enough for a vision model to
#: read handwriting without spending tokens on unusable resolution.
DEFAULT_TARGET_PX = 1568
#: Cap on pages rendered from one document.
DEFAULT_MAX_PAGES = 20
# Unquoted attribute value in the root tag, e.g. Write's `width=auto`. Matches
# only bare alphabetic values so numeric or already-quoted attributes are left
# alone.
_UNQUOTED_ATTR = re.compile(rb"(\s[-\w:]+)=([A-Za-z][-\w]*)(?=[\s>])")
_ROOT_TAG = re.compile(rb"<svg[^>]*>")
class SvgRenderError(Exception):
"""Raised when an SVG cannot be parsed or contains no renderable page."""
def _repair_root_tag(data: bytes) -> bytes:
"""Quote unquoted attribute values in the root ``<svg>`` tag.
Some Write files emit ``width=auto height=auto``, which is not valid XML.
A malformation *inside* the root tag defeats ``recover=True`` differently
from one in the body: rather than dropping a subtree, libxml2 abandons the
whole document and yields a bare root, so every page becomes invisible.
Quoting the values first recovers the full tree.
"""
match = _ROOT_TAG.search(data)
if not match:
return data
repaired = _UNQUOTED_ATTR.sub(rb'\1="\2"', match.group(0))
if repaired == match.group(0):
return data
return data[: match.start()] + repaired + data[match.end() :]
def _parser() -> etree.XMLParser:
"""Build the hardened parser used for all untrusted SVG input.
``resolve_entities=False`` blocks XXE — ingest content is untrusted.
``huge_tree`` is required because handwriting path data runs to megabytes.
``recover`` salvages the handful of Write files that emit unescaped
attribute content and are not well-formed XML.
"""
return etree.XMLParser(
huge_tree=True,
resolve_entities=False,
no_network=True,
recover=True,
)
def _dimension(element: etree._Element, name: str) -> float | None:
"""Read a CSS-pixel dimension attribute, tolerating a ``px`` suffix."""
raw = element.get(name)
if not raw:
return None
try:
return float(raw.strip().removesuffix("px"))
except ValueError:
return None
def _viewbox_size(element: etree._Element) -> tuple[float, float] | None:
"""Derive width/height from a viewBox extent."""
raw = element.get("viewBox")
if not raw:
return None
parts = raw.replace(",", " ").split()
if len(parts) != 4:
return None
try:
width, height = float(parts[2]), float(parts[3])
except ValueError:
return None
return (width, height) if width > 0 and height > 0 else None
def _element_size(element: etree._Element) -> tuple[float, float] | None:
"""Resolve an element's rendered size from width/height, else viewBox."""
width = _dimension(element, "width")
height = _dimension(element, "height")
if width and height:
return width, height
return _viewbox_size(element)
def _standalone_page(
page: etree._Element,
defs: etree._Element | None,
width: float,
height: float,
) -> bytes:
"""Wrap one ``write-page`` element as its own renderable SVG document.
The page is repositioned to the origin (its x/y place it within the stacked
parent) and given an explicit viewBox so the renderer has an unambiguous
size. ``defs`` is copied in because pen and ruling definitions live on the
root and are referenced by page content.
"""
root = etree.Element(f"{_SVG}svg", nsmap={None: SVG_NS, "xlink": XLINK_NS})
root.set("width", str(width))
root.set("height", str(height))
root.set("viewBox", f"0 0 {width} {height}")
background = etree.SubElement(root, f"{_SVG}rect")
background.set("width", "100%")
background.set("height", "100%")
background.set("fill", _PAGE_BACKGROUND)
if defs is not None:
root.append(copy.deepcopy(defs))
element = copy.deepcopy(page)
for positional in ("x", "y"):
element.attrib.pop(positional, None)
element.set("viewBox", f"0 0 {width} {height}")
root.append(element)
return etree.tostring(root)
def split_svg_pages(data: bytes) -> list[bytes]:
"""Split an SVG into one standalone document per renderable page.
Write multi-page notes yield one document per ``write-page`` child. Any
other SVG (a diagram, a logo) yields a single document — the input itself,
which the renderer sizes from its own width/height or viewBox.
:param data: Raw SVG bytes.
:returns: One or more standalone SVG documents, in page order.
:raises SvgRenderError: If the SVG cannot be parsed or has no usable size.
"""
try:
root = etree.fromstring(_repair_root_tag(data), _parser())
except etree.XMLSyntaxError as exc:
raise SvgRenderError(f"Could not parse SVG: {exc}") from exc
if root is None:
raise SvgRenderError("Could not parse SVG: no root element.")
defs = root.find(f"{_SVG}defs")
pages: list[bytes] = []
for element in root.iter(f"{_SVG}svg"):
if element is root:
continue
if _PAGE_CLASS not in (element.get("class") or "").split():
continue
size = _element_size(element)
if not size:
continue
pages.append(_standalone_page(element, defs, *size))
if pages:
return pages
# Generic SVG: render as-is. It must still be sizeable, or the renderer
# would silently substitute a default page box.
if not _element_size(root):
raise SvgRenderError(
"SVG has no width/height or viewBox, so its size is undefined."
)
return [data]
def render_svg_pages(
data: bytes,
max_pages: int = DEFAULT_MAX_PAGES,
target_px: int = DEFAULT_TARGET_PX,
) -> list[tuple[bytes, int, int]]:
"""Rasterize an SVG to one PNG per page.
:param data: Raw SVG bytes.
:param max_pages: Cap on pages rendered.
:param target_px: Longest edge of each rendered page, in pixels.
:returns: One ``(png_bytes, width, height)`` tuple per page, in page order.
:raises SvgRenderError: If nothing could be rendered.
"""
import fitz
pages = split_svg_pages(data)
rendered: list[tuple[bytes, int, int]] = []
for index, page_svg in enumerate(pages[:max_pages]):
try:
with fitz.open(stream=page_svg, filetype="svg") as document:
page = document[0]
longest = max(page.rect.width, page.rect.height)
zoom = target_px / longest if longest else 1.0
pixmap = page.get_pixmap(matrix=fitz.Matrix(zoom, zoom))
rendered.append(
(pixmap.tobytes("png"), pixmap.width, pixmap.height)
)
except Exception as exc:
raise SvgRenderError(
f"Could not render SVG page {index + 1}: {exc}"
) from exc
if not rendered:
raise SvgRenderError("No renderable pages in SVG.")
logger.info(
"Rasterized SVG pages_rendered=%d total_pages=%d target_px=%d",
len(rendered),
len(pages),
target_px,
)
return rendered

View File

@@ -347,6 +347,7 @@ def ingest_from_daedalus(self, job_id: str):
from datetime import datetime, timezone from datetime import datetime, timezone
from library.models import IngestJob, Item, Library from library.models import IngestJob, Item, Library
from library.services.parsers import UnsupportedFileTypeError
from library.services.source_s3 import ( from library.services.source_s3 import (
copy_into_mnemosyne, copy_into_mnemosyne,
fetch_from_source, fetch_from_source,
@@ -465,6 +466,24 @@ def ingest_from_daedalus(self, job_id: str):
**result, **result,
} }
except UnsupportedFileTypeError as exc:
# Deterministic, input-driven — re-parsing identical bytes can never
# succeed. Terminal client-data failure, never retried, and logged at
# WARNING (not ERROR) because an unparseable input is not a server fault.
logger.warning(
"Task ingest_from_daedalus rejected job_id=%s: %s", job_id, exc,
)
job.status = "failed"
job.error = str(exc)
job.completed_at = datetime.now(timezone.utc)
job.save(update_fields=["status", "error", "completed_at"])
return {
"success": False,
"job_id": job_id,
"error": str(exc),
"reason": "unsupported_file_type",
}
except Exception as exc: except Exception as exc:
logger.error( logger.error(
"Task ingest_from_daedalus failed job_id=%s: %s", "Task ingest_from_daedalus failed job_id=%s: %s",
@@ -484,16 +503,9 @@ def ingest_from_daedalus(self, job_id: str):
def _delete_item_and_chunks(item_uid: str): def _delete_item_and_chunks(item_uid: str):
"""Delete an Item, its chunks, and its images. Concept GC is workspace-delete only.""" """Delete an Item, its chunks, and its images. Concept GC is workspace-delete only."""
db.cypher_query( from library.services.library_delete import delete_item_cascade
"""
MATCH (i:Item {uid: $uid}) delete_item_cascade(item_uid)
OPTIONAL MATCH (i)-[:HAS_CHUNK]->(c:Chunk)
OPTIONAL MATCH (i)-[:HAS_IMAGE]->(img:Image)
OPTIONAL MATCH (img)-[:HAS_EMBEDDING]->(emb:ImageEmbedding)
DETACH DELETE c, img, emb, i
""",
{"uid": item_uid},
)
def _resolve_or_create_default_collection(lib, collection_uid: str = ""): def _resolve_or_create_default_collection(lib, collection_uid: str = ""):

View File

@@ -12,15 +12,22 @@
<div class="alert alert-warning mb-6"> <div class="alert alert-warning mb-6">
<span>Are you sure you want to delete <strong>{{ library.name }}</strong>? This action cannot be undone.</span> <span>Are you sure you want to delete <strong>{{ library.name }}</strong>? This action cannot be undone.</span>
</div> </div>
{% if library.workspace_id %} {% if library.managed_by_display %}
<div class="alert alert-error mb-6"> <div class="alert alert-error mb-6">
<span> <span>
<strong>This Library is managed by Daedalus</strong> <strong>This Library is managed by {{ library.managed_by_display }}</strong>{% if library.workspace_id %}
(workspace <code>{{ library.workspace_id }}</code>). (workspace <code>{{ library.workspace_id }}</code>){% endif %}.
{% if library.workspace_id %}
Deleting it here removes its embedded content from Mnemosyne, but the Deleting it here removes its embedded content from Mnemosyne, but the
source files still live in Daedalus — it will be <strong>recreated and source files still live in {{ library.managed_by_display }} — it will
re-embedded on the next Daedalus sync</strong>. Use this to clear an be <strong>recreated and re-embedded on the next sync</strong>. Use
orphaned Library that is blocking workspace re-registration. this to clear an orphaned Library that is blocking workspace
re-registration.
{% else %}
Deleting it here removes its embedded content from Mnemosyne;
{{ library.managed_by_display }} may recreate and re-embed it on its
next sync.
{% endif %}
</span> </span>
</div> </div>
{% endif %} {% endif %}

View File

@@ -12,10 +12,10 @@
<h1 class="text-3xl font-bold">{{ library.name }}</h1> <h1 class="text-3xl font-bold">{{ library.name }}</h1>
<div class="flex flex-wrap gap-2 mt-2"> <div class="flex flex-wrap gap-2 mt-2">
<div class="badge badge-primary">{{ library.library_type }}</div> <div class="badge badge-primary">{{ library.library_type }}</div>
{% if library.workspace_id %} {% if library.managed_by_display %}
<div class="badge badge-warning gap-1" <div class="badge badge-warning gap-1"
title="Workspace {{ library.workspace_id }}"> {% if library.workspace_id %}title="Workspace {{ library.workspace_id }}"{% endif %}>
Daedalus workspace Managed by {{ library.managed_by_display }}
</div> </div>
{% endif %} {% endif %}
</div> </div>
@@ -29,18 +29,25 @@
</div> </div>
</div> </div>
{% if library.workspace_id %} {% if library.managed_by_display %}
<div class="alert alert-warning mb-6"> <div class="alert alert-warning mb-6">
<div> <div>
<div class="font-semibold">Managed by Daedalus</div> <div class="font-semibold">Managed by {{ library.managed_by_display }}</div>
<div class="text-sm opacity-80"> <div class="text-sm opacity-80">
This library was created for Daedalus workspace {% if library.workspace_id %}
This library was created for workspace
<code class="font-mono">{{ library.workspace_id }}</code>. <code class="font-mono">{{ library.workspace_id }}</code>.
Normally you manage it from Daedalus. Deleting it here removes its Normally you manage it from {{ library.managed_by_display }}.
embedded content from Mnemosyne, but the source files still live in Deleting it here removes its embedded content from Mnemosyne, but
Daedalus — it will be recreated and re-embedded on the next sync. the source files still live in {{ library.managed_by_display }} —
it will be recreated and re-embedded on the next sync.
Use Delete to clear an orphaned library that is blocking workspace Use Delete to clear an orphaned library that is blocking workspace
re-registration. re-registration.
{% else %}
Content in this library is pushed by
{{ library.managed_by_display }}. Edits made here may be
overwritten or re-created on its next sync.
{% endif %}
</div> </div>
</div> </div>
</div> </div>

View File

@@ -19,9 +19,9 @@
<div class="form-control"> <div class="form-control">
<label class="label"><span class="label-text">Scope</span></label> <label class="label"><span class="label-text">Scope</span></label>
<select name="scope" class="select select-bordered select-sm"> <select name="scope" class="select select-bordered select-sm">
<option value="all" {% if scope == "all" %}selected{% endif %}>All libraries</option> <option value="all" {% if scope == "all" %}selected{% endif %}>All libraries</option>
<option value="global" {% if scope == "global" %}selected{% endif %}>Global only</option> <option value="unmanaged" {% if scope == "unmanaged" %}selected{% endif %}>Unmanaged only</option>
<option value="daedalus" {% if scope == "daedalus" %}selected{% endif %}>Daedalus workspaces only</option> <option value="managed" {% if scope == "managed" %}selected{% endif %}>App-managed only</option>
</select> </select>
</div> </div>
<button type="submit" class="btn btn-sm btn-outline">Filter</button> <button type="submit" class="btn btn-sm btn-outline">Filter</button>
@@ -45,9 +45,9 @@
</h2> </h2>
<div class="flex flex-wrap gap-1"> <div class="flex flex-wrap gap-1">
<div class="badge badge-outline">{{ lib.library_type }}</div> <div class="badge badge-outline">{{ lib.library_type }}</div>
{% if lib.workspace_id %} {% if lib.managed_by_display %}
<div class="badge badge-warning gap-1" title="Managed by Daedalus workspace {{ lib.workspace_id }} — do not delete from Mnemosyne."> <div class="badge badge-warning gap-1" title="Managed by {{ lib.managed_by_display }}{% if lib.workspace_id %} (workspace {{ lib.workspace_id }}){% endif %} — do not delete from Mnemosyne.">
Daedalus workspace Managed by {{ lib.managed_by_display }}
</div> </div>
{% endif %} {% endif %}
</div> </div>

View File

@@ -21,6 +21,7 @@ class LibraryTypeDefaultsTests(TestCase):
"film", "film",
"art", "art",
"journal", "journal",
"email",
"business", "business",
"finance", "finance",
} }

View File

@@ -0,0 +1,148 @@
"""Tests for the plain library REST endpoints beyond create.
Currently covers the DELETE cascades: library, collection, and item DELETE
endpoints must go through the shared ``library_delete`` service functions —
bare ``.delete()`` calls leak child nodes (Collections, Items, Chunks,
Images) and, for libraries, skip orphan-Concept GC. Neo4j is stubbed via
``sys.modules``, same style as ``test_managed_by.py``.
"""
from __future__ import annotations
from types import SimpleNamespace
from unittest.mock import MagicMock, patch
from django.contrib.auth import get_user_model
from django.test import TestCase
from rest_framework.test import APIClient
User = get_user_model()
def _fake_node_cls(instance):
"""A neomodel-class stand-in whose ``nodes.get`` returns ``instance``.
``instance=None`` makes ``nodes.get`` raise the class's DoesNotExist.
"""
fake_nodes = MagicMock()
class DoesNotExist(Exception):
pass
if instance is None:
fake_nodes.get.side_effect = DoesNotExist()
else:
fake_nodes.get.return_value = instance
return SimpleNamespace(nodes=fake_nodes, DoesNotExist=DoesNotExist)
class LibraryApiDeleteTests(TestCase):
"""DELETE on the plain library endpoint cascades."""
def setUp(self):
self.user = User.objects.create_user(username="op", password="pw")
self.client = APIClient()
self.client.force_authenticate(user=self.user)
def _fake_models_module(self, lib):
return SimpleNamespace(Library=_fake_node_cls(lib))
def test_delete_uses_shared_cascade(self):
lib = SimpleNamespace(uid="lib-1", name="Docs")
cascade_result = {
"library_uid": "lib-1",
"name": "Docs",
"item_count": 3,
"item_s3_keys": [],
"orphans_deleted": 1,
}
with patch.dict(
"sys.modules", {"library.models": self._fake_models_module(lib)}
), patch(
"library.api.views.delete_library_cascade",
return_value=cascade_result,
) as mock_cascade:
response = self.client.delete("/library/api/libraries/lib-1/")
self.assertEqual(response.status_code, 204)
mock_cascade.assert_called_once_with(lib)
def test_delete_missing_library_returns_404(self):
with patch.dict(
"sys.modules", {"library.models": self._fake_models_module(None)}
), patch("library.api.views.delete_library_cascade") as mock_cascade:
response = self.client.delete("/library/api/libraries/nope/")
self.assertEqual(response.status_code, 404)
mock_cascade.assert_not_called()
class CollectionApiDeleteTests(TestCase):
"""DELETE on the collection endpoint cascades Items/Chunks/Images."""
def setUp(self):
self.user = User.objects.create_user(username="op", password="pw")
self.client = APIClient()
self.client.force_authenticate(user=self.user)
def test_delete_uses_shared_cascade(self):
col = SimpleNamespace(uid="col-1", name="Default")
with patch.dict(
"sys.modules",
{"library.models": SimpleNamespace(Collection=_fake_node_cls(col))},
), patch(
"library.api.views.delete_collection_cascade",
return_value={
"collection_uid": "col-1",
"name": "Default",
"item_count": 2,
"item_s3_keys": [],
},
) as mock_cascade:
response = self.client.delete("/library/api/collections/col-1/")
self.assertEqual(response.status_code, 204)
mock_cascade.assert_called_once_with(col)
def test_delete_missing_collection_returns_404(self):
with patch.dict(
"sys.modules",
{"library.models": SimpleNamespace(Collection=_fake_node_cls(None))},
), patch("library.api.views.delete_collection_cascade") as mock_cascade:
response = self.client.delete("/library/api/collections/nope/")
self.assertEqual(response.status_code, 404)
mock_cascade.assert_not_called()
class ItemApiDeleteTests(TestCase):
"""DELETE on the item endpoint cascades Chunks/Images/embeddings."""
def setUp(self):
self.user = User.objects.create_user(username="op", password="pw")
self.client = APIClient()
self.client.force_authenticate(user=self.user)
def test_delete_uses_shared_cascade(self):
item = SimpleNamespace(uid="item-1", title="Doc")
with patch.dict(
"sys.modules",
{"library.models": SimpleNamespace(Item=_fake_node_cls(item))},
), patch(
"library.api.views.delete_item_cascade",
return_value={"item_uid": "item-1", "s3_key": ""},
) as mock_cascade:
response = self.client.delete("/library/api/items/item-1/")
self.assertEqual(response.status_code, 204)
mock_cascade.assert_called_once_with("item-1")
def test_delete_missing_item_returns_404(self):
with patch.dict(
"sys.modules",
{"library.models": SimpleNamespace(Item=_fake_node_cls(None))},
), patch("library.api.views.delete_item_cascade") as mock_cascade:
response = self.client.delete("/library/api/items/nope/")
self.assertEqual(response.status_code, 404)
mock_cascade.assert_not_called()

View File

@@ -0,0 +1,207 @@
"""Tests for the per-app ``managed_by`` concept.
Covers the pure helpers (``infer_legacy_manager``,
``Library.managed_by_display``), the token-derived stamping and
duplicate-name rejection on the plain create endpoint (Neo4j stubbed
via ``sys.modules``, same style as ``test_views.py``), and the
``WorkspaceStatusSerializer`` surface. Cypher-touching paths are
covered by the manual end-to-end plan, not these unit tests.
"""
from __future__ import annotations
from types import SimpleNamespace
from unittest.mock import patch
from django.contrib.auth import get_user_model
from django.test import TestCase
from rest_framework.test import APIClient
from library.api.serializers import WorkspaceStatusSerializer
from library.models import Library, infer_legacy_manager
from mcp_server.models import UserToken
User = get_user_model()
class InferLegacyManagerTests(TestCase):
"""Truth table for the pre-``managed_by`` inference."""
def test_null_workspace_is_unmanaged(self):
self.assertIsNone(infer_legacy_manager(None))
self.assertIsNone(infer_legacy_manager(""))
def test_kairos_mail_prefix_is_kairos(self):
self.assertEqual(
infer_legacy_manager("kairos-mail-abc123-7"), "Kairos"
)
def test_other_workspace_is_daedalus(self):
self.assertEqual(
infer_legacy_manager("2f9c4a1e-uuid-ish"), "Daedalus"
)
class ManagedByDisplayTests(TestCase):
"""``managed_by_display`` on in-memory (unsaved) Library nodes."""
def test_stamped_label_wins(self):
lib = Library(name="x", managed_by="Spelunker")
self.assertEqual(lib.managed_by_display, "Spelunker")
def test_stamped_label_wins_over_inference(self):
lib = Library(
name="x", managed_by="My Token", workspace_id="kairos-mail-a-1"
)
self.assertEqual(lib.managed_by_display, "My Token")
def test_legacy_workspace_falls_back_to_inference(self):
lib = Library(name="x", workspace_id="ws-uuid")
self.assertEqual(lib.managed_by_display, "Daedalus")
def test_unmanaged_is_empty_string(self):
lib = Library(name="x")
self.assertEqual(lib.managed_by_display, "")
class _FakeLibrary:
"""Stand-in for the neomodel Library on the plain create endpoint."""
class DoesNotExist(Exception):
pass
existing = None # what find_library_by_name_ci returns
instances = [] # constructor kwargs, in order
save_raises = None # exception instance save() should raise, if any
def __init__(self, **kwargs):
type(self).instances.append(kwargs)
self.__dict__.update(kwargs)
self.uid = "lib-new"
self.workspace_id = None
self.created_at = None
def save(self):
if type(self).save_raises is not None:
raise type(self).save_raises
return self
def _fake_find_ci(name):
_FakeLibrary.ci_queries.append(name)
return _FakeLibrary.existing
def _fake_models_module():
return SimpleNamespace(
Library=_FakeLibrary, find_library_by_name_ci=_fake_find_ci
)
class LibraryCreateStampingTests(TestCase):
"""POST /library/api/libraries/ stamps ``managed_by`` and rejects dupes."""
def setUp(self):
self.user = User.objects.create_user(username="op", password="pw")
self.client = APIClient()
_FakeLibrary.existing = None
_FakeLibrary.instances = []
_FakeLibrary.save_raises = None
_FakeLibrary.ci_queries = []
def _post(self, token=None, name="Docs"):
self.client.force_authenticate(user=self.user, token=token)
with patch.dict("sys.modules", {"library.models": _fake_models_module()}):
return self.client.post(
"/library/api/libraries/",
{"name": name, "library_type": "technical"},
format="json",
)
def test_token_create_stamps_token_name(self):
response = self._post(token=UserToken(name="Spelunker"))
self.assertEqual(response.status_code, 201)
self.assertEqual(_FakeLibrary.instances[0]["managed_by"], "Spelunker")
self.assertEqual(response.json()["managed_by"], "Spelunker")
def test_session_create_leaves_managed_by_null(self):
response = self._post(token=None)
self.assertEqual(response.status_code, 201)
self.assertIsNone(_FakeLibrary.instances[0]["managed_by"])
def test_duplicate_name_returns_409_name_conflict(self):
_FakeLibrary.existing = SimpleNamespace(
uid="lib-old", name="Docs", managed_by_display="Daedalus"
)
response = self._post(token=UserToken(name="Spelunker"))
self.assertEqual(response.status_code, 409)
body = response.json()
self.assertEqual(body["code"], "name_conflict")
self.assertIn("Docs", body["detail"])
self.assertEqual(body["uid"], "lib-old")
self.assertEqual(body["managed_by"], "Daedalus")
self.assertEqual(_FakeLibrary.instances, [])
def test_duplicate_check_is_case_insensitive(self):
"""A case-variant name 409s and reports the existing spelling."""
_FakeLibrary.existing = SimpleNamespace(
uid="lib-old", name="Amazon Connect", managed_by_display="Spelunker"
)
response = self._post(token=None, name="amazon connect")
self.assertEqual(response.status_code, 409)
# The lookup received the posted name (case-folding happens in
# Cypher), and the response names the existing spelling.
self.assertEqual(_FakeLibrary.ci_queries, ["amazon connect"])
self.assertIn("Amazon Connect", response.json()["detail"])
self.assertEqual(_FakeLibrary.instances, [])
def test_duplicate_of_unmanaged_reports_null_manager(self):
_FakeLibrary.existing = SimpleNamespace(
uid="lib-old", name="Docs", managed_by_display=""
)
response = self._post(token=None)
self.assertEqual(response.status_code, 409)
self.assertIsNone(response.json()["managed_by"])
def test_save_race_returns_409_not_500(self):
"""An exact-name twin landing between pre-check and save 409s."""
from neomodel.exceptions import UniqueProperty
_FakeLibrary.save_raises = UniqueProperty("name")
response = self._post(token=None)
self.assertEqual(response.status_code, 409)
self.assertEqual(response.json()["code"], "name_conflict")
class WorkspaceStatusSerializerManagedByTests(TestCase):
"""The workspace status payload carries ``managed_by`` (nullable)."""
BASE = {
"workspace_id": "ws_a",
"library_uid": "lib_1",
"name": "W",
"library_type": "technical",
"description": "",
"item_count": 0,
"chunk_count": 0,
"created_at": "2026-01-01T00:00:00Z",
}
def test_managed_by_value_round_trips(self):
s = WorkspaceStatusSerializer(data={**self.BASE, "managed_by": "Daedalus"})
self.assertTrue(s.is_valid(), s.errors)
self.assertEqual(s.validated_data["managed_by"], "Daedalus")
def test_managed_by_null_accepted(self):
s = WorkspaceStatusSerializer(data={**self.BASE, "managed_by": None})
self.assertTrue(s.is_valid(), s.errors)
def test_managed_by_absent_accepted(self):
s = WorkspaceStatusSerializer(data=self.BASE)
self.assertTrue(s.is_valid(), s.errors)

View File

@@ -0,0 +1,195 @@
"""
Tests for SVG rasterization in the ingest pipeline.
SVG is vector XML: Pillow cannot decode it and the vision stage cannot send it
as a data URI, so it is rendered to PNG at parse time. The cases that matter
are the ones that fail *silently* — Write notes stack their pages inside one
document, and rendering that document whole yields a plausible-looking PNG that
is either blank or an illegible tall strip.
"""
import io
import os
import tempfile
from django.test import TestCase
from PIL import Image
from library.services.parsers import (
IMAGE_EXTENSIONS,
DocumentParser,
)
from library.services.svg_raster import (
SvgRenderError,
render_svg_pages,
split_svg_pages,
)
# Geometry copied from a real Write note; the page element is identical across
# both root formats, which is what makes splitting format-agnostic.
PAGE_WIDTH = 1094
PAGE_HEIGHT = 1654
PAGE_PITCH = 1674 # page height + inter-page gap
def write_document(pages: int, root_size: bool = False) -> bytes:
"""
Build a Write-style document.
:param pages: Number of stacked pages.
:param root_size: Emit width/height on the root <svg>. False models
pre-eeab021 files (the permanent majority), True models newer saves.
"""
root_attrs = ""
if root_size:
total = 10 + pages * PAGE_PITCH
root_attrs = f' width="{PAGE_WIDTH + 20}" height="{total}"'
body = "".join(
f'<svg class="write-page" x="10" y="{10 + i * PAGE_PITCH}" '
f'width="{PAGE_WIDTH}px" height="{PAGE_HEIGHT}px" '
f'xmlns="http://www.w3.org/2000/svg">'
f'<path d="M 100 {100 + i * 40} L {600 + i * 120} {700 + i * 40}" '
f'stroke="#000000" stroke-width="12" fill="none"/></svg>'
for i in range(pages)
)
return (
f'<svg id="write-document"{root_attrs} '
f'xmlns="http://www.w3.org/2000/svg" '
f'xmlns:xlink="http://www.w3.org/1999/xlink">'
f'<rect id="write-doc-background" width="100%" height="100%" fill="#808080"/>'
f'<defs id="write-defs"><style>.write-flat-pen{{fill:none}}</style></defs>'
f"{body}</svg>"
).encode()
GENERIC_SVG = (
b'<svg xmlns="http://www.w3.org/2000/svg" width="400" height="200" '
b'viewBox="0 0 400 200">'
b'<path d="M 20 20 L 380 180" stroke="#000" stroke-width="10"/></svg>'
)
def ink_fraction(png: bytes) -> float:
"""Fraction of non-white pixels — the blank-render canary."""
grey = Image.open(io.BytesIO(png)).convert("L")
histogram = grey.histogram()
return sum(histogram[:200]) / sum(histogram)
class SvgSplitTests(TestCase):
"""Splitting a Write document into per-page SVGs."""
def test_splits_one_document_per_page(self):
for pages in (1, 2, 5):
for root_size in (False, True):
with self.subTest(pages=pages, root_size=root_size):
document = write_document(pages, root_size=root_size)
self.assertEqual(len(split_svg_pages(document)), pages)
def test_both_root_formats_split_identically(self):
"""Root width/height (Write eeab021) must not change the outcome.
Existing notes are never rewritten, so both formats persist
indefinitely and have to render the same.
"""
without = split_svg_pages(write_document(3, root_size=False))
with_size = split_svg_pages(write_document(3, root_size=True))
self.assertEqual(len(without), len(with_size))
self.assertEqual(len(without), 3)
def test_generic_svg_is_a_single_page(self):
self.assertEqual(len(split_svg_pages(GENERIC_SVG)), 1)
def test_unsized_svg_is_rejected_rather_than_guessed(self):
unsized = b'<svg xmlns="http://www.w3.org/2000/svg"><rect/></svg>'
with self.assertRaises(SvgRenderError):
split_svg_pages(unsized)
def test_script_inside_defs_does_not_desync_the_split(self):
"""Regression: a non-greedy <defs>...</defs> regex mis-parses these."""
document = write_document(2).replace(
b'<defs id="write-defs">',
b'<defs id="write-defs"><script><float value="770" /></script>',
)
self.assertEqual(len(split_svg_pages(document)), 2)
def test_unquoted_root_attributes_are_repaired(self):
"""Write can emit ``width=auto``, which is not valid XML.
A malformation inside the root tag defeats recovery differently from
one in the body: libxml2 abandons the whole document and returns a
bare root, so every page silently disappears.
"""
document = write_document(2).replace(
b'<svg id="write-document"',
b'<svg width=auto height=auto id="write-document"',
)
self.assertEqual(len(split_svg_pages(document)), 2)
class SvgRenderTests(TestCase):
"""Rasterizing pages to PNG."""
def test_renders_one_png_per_page(self):
pages = render_svg_pages(write_document(3))
self.assertEqual(len(pages), 3)
for png, width, height in pages:
self.assertEqual(Image.open(io.BytesIO(png)).format, "PNG")
self.assertEqual(max(width, height), 1568)
def test_rendered_pages_are_not_blank(self):
"""The pre-patch trap: a whole-document render emits ruling, no ink."""
for root_size in (False, True):
with self.subTest(root_size=root_size):
for png, _, _ in render_svg_pages(
write_document(3, root_size=root_size)
):
self.assertGreater(ink_fraction(png), 0.001)
def test_pages_are_page_shaped_not_a_stacked_strip(self):
"""The post-patch trap: root dimensions span every stacked page.
Rendering that whole gives one tall strip which, once capped, squashes
a 5-page note to ~209px wide.
"""
expected = PAGE_HEIGHT / PAGE_WIDTH
for root_size in (False, True):
with self.subTest(root_size=root_size):
for _, width, height in render_svg_pages(
write_document(5, root_size=root_size)
):
self.assertAlmostEqual(height / width, expected, delta=0.05)
def test_page_cap_is_respected(self):
self.assertEqual(len(render_svg_pages(write_document(8), max_pages=5)), 5)
class SvgParserIntegrationTests(TestCase):
"""The parser dispatch — SVG must not reach the Pillow image path."""
def setUp(self):
self.parser = DocumentParser()
def test_svg_is_not_in_image_extensions(self):
# It was, and PIL.Image.open cannot decode SVG, so ingest always failed.
self.assertNotIn("svg", IMAGE_EXTENSIONS)
def test_parse_multipage_svg_yields_one_image_per_page(self):
with tempfile.NamedTemporaryFile(suffix=".svg", delete=False) as f:
f.write(write_document(3))
f.flush()
path = f.name
try:
result = self.parser.parse(path, "svg")
finally:
os.unlink(path)
self.assertEqual(len(result.images), 3)
self.assertEqual(result.metadata["page_count"], 3)
self.assertEqual(result.text_blocks, [])
for index, image in enumerate(result.images):
# PNG, not svg — the vision stage needs raster for its data URI.
self.assertEqual(image.ext, "png")
self.assertEqual(image.source_page, index)
self.assertGreater(ink_fraction(image.data), 0.001)

View File

@@ -67,6 +67,114 @@ class ReembedItemTaskTests(TestCase):
mock_pipeline.reprocess_item.assert_called_once() mock_pipeline.reprocess_item.assert_called_once()
@override_settings(CELERY_TASK_ALWAYS_EAGER=True)
class IngestFromDaedalusFailureClassificationTests(TestCase):
"""A deterministic parse failure is terminal; a transient error retries.
Neo4j and S3 are mocked at the task's boundaries so the test exercises the
exception-classification branch (parsers.UnsupportedFileTypeError vs. any
other Exception) without a live graph or object store.
"""
def setUp(self):
from library.tasks import ingest_from_daedalus
# Calling .run() bypasses Celery's request setup, but the task
# persists self.request.id into the NOT NULL celery_task_id column
# and reads self.request.retries — push a real request context.
ingest_from_daedalus.push_request(id="test-task-id", retries=0)
self.addCleanup(ingest_from_daedalus.pop_request)
def _make_job(self):
from library.models import IngestJob
return IngestJob.objects.create(
id="job_test_unsupported",
library_uid="lib-uid-123",
source="daedalus",
s3_key="incoming/bad.drawio",
file_type="vnd.jgraph.mxfile",
title="bad.drawio",
content_hash="abc123",
)
def _patched_boundaries(self, pipeline_side_effect):
"""Patch every boundary the task hits before the pipeline runs.
Returns a context-manager list; the pipeline's ``process_item`` is set
to raise ``pipeline_side_effect``.
"""
from library.services.parsers import UnsupportedFileTypeError # noqa: F401
patchers = [
patch("library.tasks.db"),
patch("library.models.Library"),
patch("library.models.Item"),
patch("library.services.source_s3.fetch_from_source", return_value=b"data"),
patch("library.services.source_s3.copy_into_mnemosyne"),
patch("library.tasks._resolve_or_create_default_collection"),
patch("library.services.pipeline.EmbeddingPipeline"),
]
mocks = [p.start() for p in patchers]
self.addCleanup(lambda: [p.stop() for p in patchers])
# db.cypher_query returns (rows, meta); no prior item to supersede.
mocks[0].cypher_query.return_value = ([], None)
# Library.nodes.get returns a stand-in library node.
mocks[1].nodes.get.return_value = MagicMock(uid="lib-uid-123")
# Item() instances carry a uid used for the S3 key.
item_instance = MagicMock(uid="item-uid-abc")
mocks[2].return_value = item_instance
# The pipeline raises the classification-relevant error.
pipeline_instance = MagicMock()
pipeline_instance.process_item.side_effect = pipeline_side_effect
mocks[6].return_value = pipeline_instance
return pipeline_instance
def test_unsupported_file_type_is_terminal_and_not_retried(self):
from library.models import IngestJob
from library.services.parsers import UnsupportedFileTypeError
from library.tasks import ingest_from_daedalus
job = self._make_job()
self._patched_boundaries(
UnsupportedFileTypeError("Unsupported file type 'vnd.jgraph.mxfile'.")
)
with patch.object(ingest_from_daedalus, "retry") as mock_retry:
result = ingest_from_daedalus.run(job.id)
# Never retried.
mock_retry.assert_not_called()
# Terminal failure with a machine-readable reason.
self.assertFalse(result["success"])
self.assertEqual(result["reason"], "unsupported_file_type")
job.refresh_from_db()
self.assertEqual(job.status, "failed")
self.assertEqual(job.retry_count, 0)
self.assertIsNotNone(job.completed_at)
self.assertIn("Unsupported file type", job.error)
def test_transient_error_takes_the_retry_path(self):
from library.tasks import ingest_from_daedalus
job = self._make_job()
self._patched_boundaries(ConnectionError("S3 hiccup"))
# self.retry raises Retry in real Celery; simulate that so the task
# doesn't fall through to the terminal branch.
from celery.exceptions import Retry
with patch.object(ingest_from_daedalus, "retry", side_effect=Retry()) as mock_retry:
with self.assertRaises(Retry):
ingest_from_daedalus.run(job.id)
mock_retry.assert_called_once()
job.refresh_from_db()
self.assertEqual(job.retry_count, 1)
class ResolveUserTests(TestCase): class ResolveUserTests(TestCase):
"""Tests for the _resolve_user helper.""" """Tests for the _resolve_user helper."""

View File

@@ -1,13 +1,13 @@
"""Tests for the library CRUD HTML views. """Tests for the library CRUD HTML views.
Currently covers ``library_list``'s Daedalus-workspace scope filter. The Currently covers ``library_list``'s app-managed scope filter. The view
view loads every ``Library`` node from Neo4j and narrows it by a ``scope`` loads every ``Library`` node from Neo4j and narrows it in Python by a
GET param (``all`` / ``global`` / ``daedalus``). These tests stub out ``scope`` GET param (``all`` / ``unmanaged`` / ``managed``, with legacy
Neo4j entirely — patching ``neo4j_available`` and injecting a fake ``global`` / ``daedalus`` aliases). These tests stub out Neo4j entirely —
``Library`` class via ``sys.modules`` — so they assert on the queryset patching ``neo4j_available`` and injecting a fake ``Library`` class via
``.filter(...)`` call the view makes and the context it renders, not on ``sys.modules`` — so they assert on the filtering the view does and the
real graph behaviour. Mirrors the mocking style in context it renders, not on real graph behaviour. Mirrors the mocking
``test_search_views_admin_scope.py``. style in ``test_search_views_admin_scope.py``.
""" """
from __future__ import annotations from __future__ import annotations
@@ -22,6 +22,17 @@ from django.urls import reverse
User = get_user_model() User = get_user_model()
def _lib(name, managed_by_display):
return SimpleNamespace(
uid=f"uid-{name}",
name=name,
library_type="technical",
description="",
workspace_id=None,
managed_by_display=managed_by_display,
)
class LibraryListScopeFilterTests(TestCase): class LibraryListScopeFilterTests(TestCase):
"""Cover the ``scope`` filter branches of ``library_list``.""" """Cover the ``scope`` filter branches of ``library_list``."""
@@ -31,57 +42,64 @@ class LibraryListScopeFilterTests(TestCase):
) )
self.client.force_login(self.user) self.client.force_login(self.user)
self.url = reverse("library:library-list") self.url = reverse("library:library-list")
self.managed = _lib("Docs", "Spelunker")
self.unmanaged = _lib("Notes", "")
def _fake_library_cls(self): def _fake_library_cls(self):
"""Return (Library stub, nodes mock) where ``nodes`` chains fluently. """Return a Library stub whose ``nodes.order_by`` yields two libraries."""
``Library.nodes`` → ``.filter(...)`` → ``.order_by(...)`` all return
the same MagicMock so the view's queryset building works regardless
of which branch it takes, and ``.filter`` records its kwargs.
"""
fake_nodes = MagicMock() fake_nodes = MagicMock()
fake_nodes.filter.return_value = fake_nodes fake_nodes.order_by.return_value = [self.managed, self.unmanaged]
fake_nodes.order_by.return_value = [] return SimpleNamespace(nodes=fake_nodes)
return SimpleNamespace(nodes=fake_nodes), fake_nodes
def _get(self, fake_library_cls, **params): def _get(self, **params):
with patch("library.views.neo4j_available", return_value=True), \ with patch("library.views.neo4j_available", return_value=True), \
patch.dict( patch.dict(
"sys.modules", "sys.modules",
{"library.models": SimpleNamespace(Library=fake_library_cls)}, {"library.models": SimpleNamespace(Library=self._fake_library_cls())},
): ):
return self.client.get(self.url, params) return self.client.get(self.url, params)
def test_default_scope_is_all_and_does_not_filter(self): def test_default_scope_is_all_and_returns_everything(self):
fake_cls, fake_nodes = self._fake_library_cls() response = self._get()
response = self._get(fake_cls)
self.assertEqual(response.status_code, 200) self.assertEqual(response.status_code, 200)
self.assertEqual(response.context["scope"], "all") self.assertEqual(response.context["scope"], "all")
fake_nodes.filter.assert_not_called() self.assertEqual(
fake_nodes.order_by.assert_called_once_with("name") list(response.context["libraries"]), [self.managed, self.unmanaged]
)
def test_global_scope_filters_workspace_isnull_true(self): def test_managed_scope_keeps_only_managed(self):
fake_cls, fake_nodes = self._fake_library_cls() response = self._get(scope="managed")
response = self._get(fake_cls, scope="global")
self.assertEqual(response.context["scope"], "global") self.assertEqual(response.context["scope"], "managed")
fake_nodes.filter.assert_called_once_with(workspace_id__isnull=True) self.assertEqual(list(response.context["libraries"]), [self.managed])
def test_daedalus_scope_filters_workspace_isnull_false(self): def test_unmanaged_scope_keeps_only_unmanaged(self):
fake_cls, fake_nodes = self._fake_library_cls() response = self._get(scope="unmanaged")
response = self._get(fake_cls, scope="daedalus")
self.assertEqual(response.context["scope"], "daedalus") self.assertEqual(response.context["scope"], "unmanaged")
fake_nodes.filter.assert_called_once_with(workspace_id__isnull=False) self.assertEqual(list(response.context["libraries"]), [self.unmanaged])
def test_legacy_daedalus_scope_aliases_to_managed(self):
response = self._get(scope="daedalus")
self.assertEqual(response.context["scope"], "managed")
self.assertEqual(list(response.context["libraries"]), [self.managed])
def test_legacy_global_scope_aliases_to_unmanaged(self):
response = self._get(scope="global")
self.assertEqual(response.context["scope"], "unmanaged")
self.assertEqual(list(response.context["libraries"]), [self.unmanaged])
def test_unknown_scope_does_not_filter(self): def test_unknown_scope_does_not_filter(self):
"""An unexpected scope value degrades to the unfiltered list.""" """An unexpected scope value degrades to the unfiltered list."""
fake_cls, fake_nodes = self._fake_library_cls() response = self._get(scope="bogus")
response = self._get(fake_cls, scope="bogus")
self.assertEqual(response.context["scope"], "bogus") self.assertEqual(response.context["scope"], "bogus")
fake_nodes.filter.assert_not_called() self.assertEqual(
list(response.context["libraries"]), [self.managed, self.unmanaged]
)
def test_neo4j_unavailable_sets_error_and_empty_list(self): def test_neo4j_unavailable_sets_error_and_empty_list(self):
with patch("library.views.neo4j_available", return_value=False): with patch("library.views.neo4j_available", return_value=False):

View File

@@ -7,6 +7,9 @@ search scoping) require Neo4j and are validated by the manual end-to-end
test plan, not these unit tests. test plan, not these unit tests.
""" """
from unittest.mock import Mock, patch
from django.contrib.auth.models import User
from django.test import TestCase from django.test import TestCase
from rest_framework.test import APIClient from rest_framework.test import APIClient
@@ -208,3 +211,113 @@ class WorkspaceEndpointAuthTests(TestCase):
def test_workspace_delete_requires_auth(self): def test_workspace_delete_requires_auth(self):
response = self.client.delete("/library/api/workspaces/ws_a/") response = self.client.delete("/library/api/workspaces/ws_a/")
self.assertIn(response.status_code, [401, 403]) self.assertIn(response.status_code, [401, 403])
class WorkspaceDeleteOwnershipTests(TestCase):
"""DELETE must never report success for a library it did not delete.
A false 204 is how libraries get orphaned: Daedalus records the
workspace as cleaned up and drops its own row, while the Library node
survives holding its globally-unique name forever — which then blocks
ever recreating a workspace under that name.
"""
def setUp(self):
self.client = APIClient()
self.user = User.objects.create_user(
username="owner", password="pw" # noqa: S106 — test credential
)
self.client.force_authenticate(user=self.user)
def _library(self, owner_username):
lib = Mock()
lib.uid = "lib_1"
lib.owner_username = owner_username
return lib
def test_absent_library_still_returns_204(self):
"""Genuine idempotency is preserved — nothing exists, nothing to do."""
from library.models import Library
with patch(
"neomodel.sync_.match.NodeSet.get",
side_effect=Library.DoesNotExist("nope"),
):
response = self.client.delete("/library/api/workspaces/ws_gone/")
self.assertEqual(response.status_code, 204)
def test_unowned_library_returns_409_and_is_not_deleted(self):
from library.models import Library
lib = self._library("someone_else")
with (
patch("neomodel.sync_.match.NodeSet.get", return_value=lib),
patch(
"library.api.workspaces.delete_library_cascade"
) as cascade,
):
response = self.client.delete("/library/api/workspaces/ws_a/")
self.assertEqual(response.status_code, 409)
self.assertEqual(response.json()["code"], "owner_conflict")
cascade.assert_not_called()
def test_owned_library_is_deleted(self):
from library.models import Library
lib = self._library("owner")
with (
patch("neomodel.sync_.match.NodeSet.get", return_value=lib),
patch(
"library.api.workspaces.delete_library_cascade",
return_value={
"library_uid": "lib_1",
"name": "Assistant",
"item_count": 0,
"orphans_deleted": 0,
},
) as cascade,
):
response = self.client.delete("/library/api/workspaces/ws_a/")
self.assertEqual(response.status_code, 204)
cascade.assert_called_once_with(lib)
def test_unclaimed_library_is_deleted(self):
"""A NULL owner is unclaimed, not foreign.
Workspace libraries predating owner stamping (and any created via the
plain /libraries/ endpoint, which never sets an owner) have
owner_username=None. Refusing those would make them permanently
undeletable through this endpoint — orphans by construction.
"""
from library.models import Library
lib = self._library(None)
with (
patch("neomodel.sync_.match.NodeSet.get", return_value=lib),
patch(
"library.api.workspaces.delete_library_cascade",
return_value={
"library_uid": "lib_1",
"name": "Assistant",
"item_count": 0,
"orphans_deleted": 0,
},
) as cascade,
):
response = self.client.delete("/library/api/workspaces/ws_a/")
self.assertEqual(response.status_code, 204)
cascade.assert_called_once_with(lib)
def test_unowned_library_get_still_looks_absent(self):
"""Ownership must stay opaque on reads — 404, not 409."""
from library.models import Library
lib = self._library("someone_else")
with patch("neomodel.sync_.match.NodeSet.get", return_value=lib):
response = self.client.get("/library/api/workspaces/ws_a/")
self.assertEqual(response.status_code, 404)

View File

@@ -31,20 +31,23 @@ logger = logging.getLogger(__name__)
@login_required @login_required
def library_list(request): def library_list(request):
"""List libraries, optionally filtered by Daedalus-workspace scope.""" """List libraries, optionally filtered by app-managed scope."""
scope = request.GET.get("scope", "all") scope = request.GET.get("scope", "all")
# Legacy bookmark values from before managed_by existed.
scope = {"daedalus": "managed", "global": "unmanaged"}.get(scope, scope)
libraries = [] libraries = []
error = None error = None
if neo4j_available(): if neo4j_available():
try: try:
from .models import Library from .models import Library
qs = Library.nodes libraries = list(Library.nodes.order_by("name"))
if scope == "daedalus": # managed_by_display covers legacy workspace libraries that
qs = qs.filter(workspace_id__isnull=False) # predate the managed_by property, so filter in Python.
elif scope == "global": if scope == "managed":
qs = qs.filter(workspace_id__isnull=True) libraries = [l for l in libraries if l.managed_by_display]
libraries = qs.order_by("name") elif scope == "unmanaged":
libraries = [l for l in libraries if not l.managed_by_display]
except Exception as e: except Exception as e:
error = f"Could not connect to Neo4j: {e}" error = f"Could not connect to Neo4j: {e}"
logger.error(error) logger.error(error)
@@ -457,7 +460,11 @@ def collection_delete(request, uid):
if request.method == "POST": if request.method == "POST":
name = col.name name = col.name
col.delete() # Shared cascade so Items/Chunks/Images go too — a bare
# col.delete() would leak them.
from .services.library_delete import delete_collection_cascade
delete_collection_cascade(col)
messages.success(request, f'Collection "{name}" deleted.') messages.success(request, f'Collection "{name}" deleted.')
return redirect("library:library-list") return redirect("library:library-list")
return render( return render(
@@ -645,7 +652,11 @@ def item_delete(request, uid):
if request.method == "POST": if request.method == "POST":
title = item.title title = item.title
item.delete() # Shared cascade so Chunks/Images/embeddings go too — a bare
# item.delete() would leak them.
from .services.library_delete import delete_item_cascade
delete_item_cascade(item.uid)
messages.success(request, f'Item "{title}" deleted.') messages.success(request, f'Item "{title}" deleted.')
return redirect("library:library-list") return redirect("library:library-list")
return render(request, "library/item_confirm_delete.html", {"item": item}) return render(request, "library/item_confirm_delete.html", {"item": item})

View File

@@ -26,6 +26,20 @@ from rest_framework import authentication, exceptions
from .auth import MCPAuthError, resolve_mcp_user from .auth import MCPAuthError, resolve_mcp_user
def request_token_label(request):
"""Name of the ``UserToken`` authenticating this request, or None.
Session-authenticated requests (``request.auth`` is None) and blank
token names return None — the caller treats both as "no managing app".
"""
from .models import UserToken
token = getattr(request, "auth", None)
if isinstance(token, UserToken):
return token.name.strip() or None
return None
class UserTokenAuthentication(authentication.BaseAuthentication): class UserTokenAuthentication(authentication.BaseAuthentication):
"""Authenticate DRF requests with a ``UserToken`` bearer.""" """Authenticate DRF requests with a ``UserToken`` bearer."""

View File

@@ -57,6 +57,12 @@ class UserTokenCreateForm(forms.Form):
"class": "input input-bordered w-full", "class": "input input-bordered w-full",
"placeholder": "e.g. Claude Desktop, CI script", "placeholder": "e.g. Claude Desktop, CI script",
}), }),
help_text=(
"A friendly label so you can identify this token later. It also "
"labels any library the token creates (shown as “Managed by "
"<name>”) — for an app integration, use the app's name, e.g. "
"Daedalus, Kairos, Spelunker."
),
) )
expires_at = forms.DateTimeField( expires_at = forms.DateTimeField(
required=False, required=False,

View File

@@ -21,7 +21,7 @@
</label> </label>
{{ form.name }} {{ form.name }}
<label class="label"> <label class="label">
<span class="label-text-alt opacity-60">A friendly label so you can identify this token later (e.g. “Claude Desktop”).</span> <span class="label-text-alt opacity-60">{{ form.name.help_text }}</span>
</label> </label>
</div> </div>
<div class="form-control mt-4"> <div class="form-control mt-4">

View File

@@ -105,6 +105,44 @@ class UserTokenAuthenticationTest(TestCase):
resp = self._get(f"Bearer {self.plaintext} extra") resp = self._get(f"Bearer {self.plaintext} extra")
self.assertEqual(resp.status_code, status.HTTP_401_UNAUTHORIZED) self.assertEqual(resp.status_code, status.HTTP_401_UNAUTHORIZED)
def test_request_token_label_reads_token_name(self):
from types import SimpleNamespace
from mcp_server.drf_auth import request_token_label
token = UserToken(name=" Spelunker ")
self.assertEqual(
request_token_label(SimpleNamespace(auth=token)), "Spelunker"
)
def test_request_token_label_none_for_session(self):
from types import SimpleNamespace
from mcp_server.drf_auth import request_token_label
self.assertIsNone(request_token_label(SimpleNamespace(auth=None)))
# A request object with no auth attribute at all (plain Django).
self.assertIsNone(request_token_label(SimpleNamespace()))
def test_request_token_label_none_for_blank_name(self):
from types import SimpleNamespace
from mcp_server.drf_auth import request_token_label
self.assertIsNone(
request_token_label(SimpleNamespace(auth=UserToken(name=" ")))
)
def test_request_token_label_none_for_foreign_auth_object(self):
from types import SimpleNamespace
from mcp_server.drf_auth import request_token_label
# e.g. a JWT dict from another auth class — not a UserToken.
self.assertIsNone(
request_token_label(SimpleNamespace(auth={"iss": "daedalus"}))
)
def test_request_auth_stashes_token(self): def test_request_auth_stashes_token(self):
# The auth class returns (user, token); DRF places the token on # The auth class returns (user, token); DRF places the token on
# request.auth. Re-use a UserToken-aware endpoint to verify. # request.auth. Re-use a UserToken-aware endpoint to verify.

View File

@@ -30,6 +30,9 @@ dependencies = [
"semantic-text-splitter>=0.20,<1.0", "semantic-text-splitter>=0.20,<1.0",
"tokenizers>=0.20,<1.0", "tokenizers>=0.20,<1.0",
"Pillow>=10.0,<12.0", "Pillow>=10.0,<12.0",
# SVG page splitting — needs recover=True for Write notes that emit
# unescaped attribute content and aren't well-formed XML
"lxml>=5.3,<7",
"requests>=2.31,<3.0", "requests>=2.31,<3.0",
# Phase 5: MCP Server # Phase 5: MCP Server
"fastmcp>=2.0,<3.0", "fastmcp>=2.0,<3.0",