Repository navigation
feat: scope reads and writes to the calling principal - #131
Merged
Merged
Conversation
…the default principal The visibility column, the access-control migration, the write-tool parameters, the batch validators and the settings accept exactly 'private' and 'public'. A private entry is readable by its owner and every grantee, so grants take effect without a separate visibility value. Requests authenticated with simple_token resolve to ACCESS_CONTROL_DEFAULT_PRINCIPAL like stdio and unauthenticated HTTP, because the bearer token proves access but names no caller. MCP_AUTH_CLIENT_ID identifies the client to FastMCP and is not an owner identity. Schema comments, docstrings, tool descriptions and server.json describe the two-value model.
app/access_scope.py defines AccessScope, SystemScope, AccessMode and build_access_predicate, which renders the READ, WRITE and OWNER row filters for SQLite and PostgreSQL. READ admits the owner, public rows and any read or write grant to the caller or one of its groups; WRITE admits the owner and write grants; OWNER admits the owner only. Groups bind as one parameter, so each predicate has a constant bind count and an empty group set needs no special case. The module imports only the standard library, so repositories can depend on it without loading settings or the MCP framework. RequestPrincipal.access_scope and resolve_access_scope turn the effective principal into the scope a request reads and writes as, leaving roles to the publish gate. The test helpers gain LOCAL_SCOPE, as_principal, insert_grant and read_grants for scoped repository and tool tests on both backends.
Add a shared case registry and seed for two-principal tests of the repository seams, with SQLite and PostgreSQL entry points. The seed stores eight entries owned by two principals that cover every visibility and grant arrangement the access model distinguishes, with tags, images, embeddings, index nodes and grants. Each backend builds an fp32 and a compressed layout through the server's own database preparation, so cases run against the storage a real deployment provisions. Scaffold tests prove the seed, the expected read, write and owner tables against the access predicate on both backends, and the registry and runner contract. A sqlite_999_variables fixture caps every SQLite connection at 999 bind variables, so tests can exercise the repositories' id chunking against the lowest common limit.
The deduplicating store and its read-only pre-check take a required keyword-only scope instead of an owner id. The dedup candidate is the latest entry of the thread and source the caller may read, and a store merges into it only when the caller owns it and the text matches, so identical text from another principal becomes that principal's own entry. Only opposite-source entries the caller may read count as a new conversational turn, so an entry hidden from the caller never changes the dedup outcome. The dedup UPDATE re-asserts ownership in its WHERE clause, and the pre-check and the store build their candidate and turn statements through the same helpers so both read the same rows. store_context and store_context_batch build the scope once per call and pass it to the pre-check and every store transaction, which records the scope principal as the grantor of author-group grants. Two-principal cases on SQLite and PostgreSQL cover candidate ownership, readable and hidden turns, the pre-check, and two principals writing into one thread.
get_by_ids, find_ids_by_prefix and check_entry_exists take a required keyword-only scope and apply the read predicate in SQL before any ordering or limit, so an entry the caller may not read is indistinguishable from a missing one. check_entry_exists also reports whether the caller may modify the entry, computed from the write predicate in the same statement. The new probe_ids reports, for every readable id among many, whether the caller may modify and owns the entry, in bounded chunks and optionally on a transaction connection. Id prefixes resolve over readable entries only, so a hidden entry never makes a prefix ambiguous and a prefix only hidden entries carry matches nothing. get_context_by_ids, navigate_context, read_context_range and both delete tools build the caller's scope once per call and pass it to prefix resolution and every read. update_context and update_context_batch resolve the caller before resolving ids, so prefix resolution, the existence probe and the version re-read all run as the caller. Two-principal cases on SQLite and PostgreSQL cover the by-id read, prefix lookup and resolution, the existence probe and the access probe.
search_contexts and grep_scan_text_contents take a required keyword-only scope, and the shared filter builder appends the read predicate after every client filter. The predicate applies in SQL before ordering, pagination and the grep keyset pages, so hidden entries never take a page slot, never count toward the scan cap and never mark a scan truncated. The predicate is never counted as a filter, so filters_applied reports the client filters alone, and its PostgreSQL placeholders follow every filter bind so a filter value cannot bind into it. search_context and grep_context build the caller's scope once per call and pass it to the repository. The filter builder drops its unused params_start parameter. Two-principal cases on SQLite and PostgreSQL cover both seams unfiltered, filtered, with a filter value naming another principal, and with hidden rows that out-rank the readable ones.
update_context_entry, patch_metadata, touch_updated_at, update_content_type, entry_exists and get_content_type take a required keyword-only scope and carry the access predicate in every statement, so an update reaches only an entry the caller may modify. Content edits require ownership or a write grant, while a visibility change requires ownership, in both the existence check and the UPDATE itself. An entry the caller may not modify reports no row exactly like a missing one, and a stale compare-and-set version on such an entry never surfaces as a version conflict. The in-transaction presence check of tags-only and images-only updates applies the write predicate and locks only the parent row on PostgreSQL. execute_update_in_transaction passes the caller's scope to every statement on the entry row and ends the update as not found when the entry's content type can no longer be read. update_context and update_context_batch refuse an entry the caller may read but not modify before any generation runs, with "Not authorized to modify context entry" errors that stay distinct from not found. The post-conflict re-read returns the whole probe, so an entry that lost write access during generation ends as not authorized instead of being retried. EntryNotAuthorizedError carries the exact denial texts for modify and delete refusals. Two-principal cases on SQLite and PostgreSQL cover the write gate, content and visibility updates, a stale version, the metadata patch, the single-column writes and the content-type read.
delete_by_ids takes a required keyword-only scope and deletes only the requested entries the caller owns, carrying the owner predicate in every chunked DELETE. get_ids_matching_batch_criteria takes a required keyword-only scope and a READ or OWNER snapshot mode, runs on PostgreSQL as well as SQLite, binds older_than_days on both backends and still matches nothing without a criterion. delete_entries_with_cleanup probes the ids inside the delete transaction, cleans and deletes only the owned ones, and refuses a delete that names a readable entry the caller does not own before anything is written. delete_context by ids refuses with "Not authorized to delete context entries", while delete_context by thread deletes only the caller's entries of the thread on both backends. delete_context_batch snapshots readable entries and refuses on a visible foreign entry when it names context_ids, and otherwise snapshots and deletes only the caller's own matching entries. An entry the caller may not read is never refused and never counted, exactly like a missing id. delete_by_thread and delete_contexts_batch are removed, so every delete on both backends goes through the id snapshot and the one chokepoint. Two-principal cases on SQLite and PostgreSQL cover the owner-only delete and both snapshot modes, including the age bound and the no-criteria guard.
Every ranked statement applies the caller's read predicate after the client filters and before the rank depth, LIMIT and OFFSET, so an entry the caller cannot read never takes a rank position or a page slot. The fp32 search, the compressed candidate selection and the full-text search on both backends carry the predicate, and on PostgreSQL the full-text predicate sits inside the limited inner subquery. Compressed search re-applies the predicate when it hydrates the ranked page, so an entry that stops being readable between the two stages is dropped. The repository searches, their per-backend executors and the raw search legs take a required keyword-only scope, and the hybrid tool passes one scope to both legs. The full-text tool docstring states that SQLite BM25 scores draw on the whole index while membership and paging depend only on readable entries. Two-principal cases on both backends cover vector search on both layouts, hydration revocation and full-text search, with filtered, adversarial and page-fill variants and a pin of the backend scoring difference.
The thread listing and every entry-derived statistics figure apply the caller's read predicate before GROUP BY, ORDER BY and LIMIT, so a thread, count or top-N item never reflects an entry the caller cannot read. Image and tag figures join their parent entry and apply the predicate to it, the embedding, full-text and index-tree node counts do the same, and the top-N lists hold the caller's own top items. The thread listing numbers its LIMIT and OFFSET placeholders after the predicate binds. The database size, embedding storage size, connection metrics and configuration blocks stay deployment-wide, and the list_threads and get_statistics docstrings say which figures are per caller. The full-text migration sizes its rebuild estimate through the system scope, since the rebuild covers every entry. The thread statistics, tag statistics, multi-entry tag and image readers and the grant listing have no caller and are removed with their tests, which now read grants through raw SQL. Two-principal cases on both backends cover the thread listing with page fill, the database statistics with top-N fill and the deployment-wide size, and the summary, embedding, full-text and node counts.
Every repository method that runs SQL against a context table must be classified, and a guarded method must apply the access predicate and take a required keyword-only scope. Each guarded method names two-principal cases that run on both backends, and every reference to a child reader or writer is pinned to a reviewed call site. Sweeps keep context-table SQL inside the repositories, migrations and CLI, keep the system scope out of request paths and keep the shared visibility out of the code. The dedup statement helpers take scope as a keyword-only argument, like every other scope carrier.
…backends A harness check runs every tool as a second principal against a server that shares the primary's database, so a private entry stays unreachable, a public entry stays read-only and a hidden id answers exactly like an absent one. The second server can start from the primary server's own environment, which pins the database, the embedding model and the compression layout to the primary's. The HTTP server helpers move into a shared module with backend-aware environments built from the MCP SDK default environment. A JWT scenario on SQLite and PostgreSQL proves the group read grant, the user read grant and the user write grant end to end, including the not-authorized update, visibility and delete refusals.
mcp-context-server-migrate --reassign-owner FROM TO rewrites the owner of every context entry owned by FROM to TO in one statement, so rows written under the default principal can be handed to an identity-provider subject after switching to JWT authentication. Both values are bound as parameters with no character-set limit, which makes the mode the route for subjects such as auth0|... that the default-principal variable cannot hold. The statement bumps updated_at on both backends and leaves the version token and every grant row unchanged. The mode joins the exclusive mode group, reads only --source-url, prints the matching row count under --dry-run, treats zero matching rows as a successful no-op, and rejects empty or identical values.
The authentication guide gains an access model covering principals, ownership, the two visibilities, grants, the not-found versus not-authorized contract and the limits of the enforcement. A new section explains how to hand existing entries to an identity-provider subject when a deployment switches to JWT, through the default-principal variable or the owner reassignment CLI. The API reference, environment variable reference, backend guide, grep and navigation guide, migration guide and README describe the scoped reads, writes, deletes and statistics, and document the visibility parameter on the four write tools. The full-text tool description states that only readable entries are matched and counted, and that SQLite BM25 scores draw on the whole index while PostgreSQL scores each entry on its own. The store, update and delete tool docstrings and parameter descriptions state the deduplication candidate rule, the update authorization errors and the delete refusal rule. CLAUDE.md records the access predicate, the scope parameter rule, the delete chokepoint, the test vehicles for access scoping and the new CLI mode.
The grant EXISTS arm of the read predicate stops SQLite from answering the OR through index lookups, so unfiltered scoped statements scanned context_entries and walked each row's text overflow pages. The access-control migration now creates three SQLite-only covering indexes, idx_context_access_thread, idx_context_access_source and idx_context_access_id, on fresh, upgraded and migrate-CLI databases alike. The SQLite get_statistics tag and image figures match child rows against the readable parent ids so they read the covering id index instead of each parent row. PostgreSQL DDL and queries are unchanged. Query-plan tests pin the covering indexes for the statistics statements and the search candidate statement.
A caller who may only read an entry gets the not-authorized error for every change, a visibility change included, and only a write grantee who is not the owner gets the owner-only visibility error. The JWT switch guidance names the run that stamps the owner of existing entries: the migrate CLI for an integer-keyed v2 database, and the first upgraded server start only for a UUIDv7 database whose entries predate the ownership columns. The hybrid_search_context description now carries the SQLite note that its fts_score draws on whole-index BM25 statistics, and a description test pins it.
No application code calls the single-image writer; store_images and replace_images_for_context are the image write paths. The two tests that exercised only this method and its access-scope registry entry go with it.
The batch update tool's write-access checks move to app/tools/batch/update_access.py, keeping the tool module within the size limit. authorize_updates holds the pre-generation probe that refuses unreadable entries as not found, readable but unmodifiable entries as not authorized, and visibility changes by anyone but the owner. reraise_disambiguated_cas_conflict holds the in-transaction re-probe that turns a compare-and-set matching zero rows into a not-found error or a version conflict. Both bodies are unchanged, and a new test module covers the moved functions directly.
On SQLite the get_statistics index_tree node count joined each node to its entry through the unique id index, and the FTS indexed count joined the FTS table to its entries, so both read every readable entry's row and its text overflow pages. A new build_readable_parent_predicate renders the READ predicate for child rows as membership in the readable entry keys, which one scan of the covering idx_context_access_id index answers. The tag, image and node counts use it on SQLite, and the FTS indexed count matches the FTS5 docsize rows against the readable rowid_int values, because a query on the external-content FTS table without MATCH reads each entry's text. PostgreSQL statements are unchanged. The plan test now runs the get_statistics tool on a database built by the startup preparation and pins every statement it executes that names context_entries to a covering access index, except the summary count, which reads the summary column. A test also pins the empty-mapping round trip of per-image metadata through store_images.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Every read, existence check, deduplication lookup, aggregate, update and delete now covers only the entries the calling principal may read or write, on SQLite and PostgreSQL. With
MCP_AUTH_PROVIDER=jwt, two users of one server no longer see or change each other's private entries. Deployments without per-user identity (stdio,none,simple_token) keep working as before: every request maps toACCESS_CONTROL_DEFAULT_PRINCIPAL, which owns every existing entry.Access model
publicentry, and every user or group holding a read or write grant.private(owner plus every grantee) andpublic(everyone). The unreleasedsharedvalue is removed from the schema, the migration, the tool parameters and the settings.simple_token: maps to the default principal, likenoneand stdio.MCP_AUTH_CLIENT_IDnames the client to FastMCP and is not an owner.What callers see
get_context_by_ids, absent from every search, count, thread list and statistic, and reported as not found by navigation, range reads and updates. Prefix ids resolve among readable entries only.Not authorized to modify context entry with ID {id}. A write grantee who tries to change visibility getsOnly the owner may change the visibility of context {id}.delete_contextwith ids anddelete_context_batchwithcontext_idsrefuse withNot authorized to delete context entries: {ids}and delete nothing when a named entry is readable but not yours. Deletes by thread or criteria remove only your own entries.store_contextmerges a retransmitted entry only into an entry you own; an identical latest entry that another principal owns never absorbs your write.list_threadsandget_statisticsreport per caller. Database size, embedding storage size, connection metrics and configuration stay deployment-wide.Other changes
mcp-context-server-migrate --source-url <url> --reassign-owner FROM TOmoves every entry owned by FROM to TO, for example after switching an existing deployment tojwt. It binds both values as parameters, so it accepts subjects such asauth0|...thatACCESS_CONTROL_DEFAULT_PRINCIPALcannot hold.--dry-runprints the count.docs/authentication.mdgains an access-model section and a guide for switching an existing deployment tojwt. The environment-variable reference, API reference, migration guide, README and tool descriptions describe the enforced behavior.Known limitation
On SQLite,
fts_score(and the full-text rank inside hybrid search) comes frombm25(), which uses table-wide statistics, so hidden entries can shift a readable entry's score. They never change which entries are returned or how many fill a page. PostgreSQL scores per document and has no such effect; the docs recommend PostgreSQL for multi-principal deployments.Testing