Skip to content

Commit 97ea5fd

Browse files
feat(mcp): capture resource discovery and reads (#928)
* feat(mcp): capture resource discovery and reads * fix(mcp): keep resource bodies out of analytics * fix(mcp): redact credentials in captured resource addresses Apply existing credential redaction to resource-read names before the primary event and exception sibling are built. Preserve the original URI and resource result or exception received by the caller. Extend the existing resource tests with successful and failing reads containing an invented token; the two new cases fail before this fix under each MCP major. Document the capture boundary and before_send. Validation: MCP v1 245 passed; MCP v2 225 passed and 13 expected skips. Ruff lint and formatting pass. Mypy baseline passes (227 source files). * fix(mcp): redact credentials embedded in captured URLs Parse captured URLs to remove userinfo and credential query values, including common signed URL fields. Apply the same sanitization to URLs inside exception messages without changing handler requests or responses. Document the limits of key-based URL redaction. Verification: reproduced the credential leak before the fix. MCP v1 suite: 260 passed; v2 suite: 240 passed, 13 skipped. Ruff check and format passed; mypy baseline passed for 227 files. Regression coverage includes encoded keys, duplicate query parameters, malformed URLs, and success/error events. * fix(mcp): bound captured URL parsing Reject captured URLs over 8,192 characters before copying or parsing them and cap parse_qsl at 128 fields. Preserve caller requests and responses. Document the limits and verify the boundary behavior in plain URLs and exception messages. Add real high-level resource-adapter coverage for early/late registration, idempotency, success/failure events, duration, and response-body exclusion. Validation: MCP v1 338 passed; MCP v2 312 passed, 17 expected skips. Ruff lint/format and mypy baseline passed. The new adapter test scores 10.0 in CodeScene; broader existing sanitizer complexity is left unchanged. * fix(mcp): widen URL redaction, capture resource listings and template listings What changed - URL sanitizer: drop the leading `\b` (it left `resource_https://user:pw@host` entirely unredacted), split trailing prose punctuation off before parsing and re-append it, match sensitive query keys per `-`/`_`/`.` segment plus a short exact list (`code` exact-only so it can't eat `country_code`), normalize `;` to `&` before splitting fields, redact `=`-shaped fragments, and sanitize one level of URL nested inside a retained query value. Only the part that actually changed is re-serialized, so an untouched `#section-2` stays byte-for-byte. - `resources/templates/list` is instrumented on every adapter and emitted as `$mcp_resources_list`; the captured `request.method` separates it from `resources/list`. - Listing events carry the listing as `$mcp_response` (names/uris/mime types are metadata, not a resource body — reads still capture no response). An empty listing is not flagged as an error: a template-only server legitimately lists no static resources. - `$identify` falls back to `params.uri` when a request has no `name`, and `resource_name` is sanitized on every event rather than only on reads — so a credential-bearing read uri is redacted there too. Why Reviewer follow-ups on posthog-python#928; the same semantics ship in posthog-js#4830 so both SDKs redact identically. How tested - `.venv/bin/pytest posthog/test/mcp -q` -> 358 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 330 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer - The spec's vector table is now the parametrized `test_sanitize_url_credentials` rows, including the `sort_key` over-redaction and the `country_code` keep. - The URL-key decision lives in `_should_redact_query_key`, kept separate to keep `_sanitize_url` shallow (CodeScene flagged complexity here before). Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): strip trailing URL punctuation in linear time What changed `_URL_TRAILING_PUNCTUATION_PATTERN` (`[.,;:!?)\]}]+$`) backtracks quadratically over an interior run of punctuation, and the URL comes from an attacker- influenceable request, so one message can carry many of them. It is now a plain character set stripped with `str.rstrip`, which is the same operation in linear time. Behavior is unchanged: `rstrip` removes exactly the trailing run the anchored pattern matched. How tested - `https://x/` + 8000 `.` + `a`: 122 ms before, 0.1 ms after - `.venv/bin/pytest posthog/test/mcp -q` -> 358 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 330 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer The existing vectors cover the behavior (prose comma, `).`, `Foo_(bar)`), so no new row was added; the JS sibling should make the same swap on its own `replace(/[.,;:!?)\]}]+$/, '')`. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): keep apostrophes inside captured URLs What changed `'` is a valid URI sub-delimiter, but the URL pattern's terminal class treated it as a terminator: `https://example.com/o'reilly?token=fakesecret` matched only up to the `o`, so the token shipped unredacted in `$mcp_resource_name` and `$mcp_parameters`, and `https://user:pa'ss@example.com/doc` kept its userinfo. The class is now `[^\s<>\"]+` — `"`, `<` and `>` cannot appear unencoded in a URI so they still terminate a match — and `'` joins the trailing-punctuation set, so a single-quoted URL in prose still has its closing quote split off and re-appended. How tested Three rows added to the parametrized `test_sanitize_url_credentials`: the two vectors above and `Read 'https://example.com/x?sig=fakesignature' first.`, which must keep both quotes and redact the signature. The test also re-runs the sanitizer over each expected value, so idempotence is covered. - `.venv/bin/pytest posthog/test/mcp -q` -> 361 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 333 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer The userinfo row is spelled `fakeuser:fake'pass` rather than `user:pa'ss` to match the fake-credential naming the rest of the table uses; it asserts the same redaction. The JS sibling gets the identical pattern change. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): close three URL-redaction leaks What changed - Depth exhaustion. A value that still carried a URL after the one-level nested pass was returned untouched, so a doubly nested gateway uri (`?url=<gateway2 whose own ?url= carries ...?token=fakesecret>`) shipped the token. The budget is one level; past it a URL-bearing value is now dropped rather than trusted. - Restored punctuation as a credential's tail. `?password=fakepass!!!` came back as `password=%5Bredacted%5D!!!`. Two rules: a string that IS a single URL (a `$mcp_resource_name`, a `params.uri`, a nested query value) has no prose, so nothing is split off it at all; and in prose, when the last field of the part the URL ends in was rewritten, the punctuation goes with it instead of being re-appended. A sentence loses its comma when it ends in a redacted credential — the accepted cost. - Pass ordering. URLs were rewritten before the PostHog-token pass, and re-serializing a query percent-encodes `/`, so `?ref=/phx_...` became `ref=%2Fphx_...` where the token pattern's `\bph` boundary no longer matched. Tokens are now redacted first, then URLs, then the entropy pass as before (that one still runs last: it works on whitespace-separated words and must see the final text). How tested Nine rows added to / adjusted in the parametrized `test_sanitize_url_credentials`, including the double-nested gateway uri, `?password=fakepass!!!`, the suffix-dropped prose rows, the two suffix-KEPT rows (credential not last; tail is a prose fragment), and the `?ref=/phx_...` row. Each expected value is re-sanitized by the test, so all of them are idempotent. - `.venv/bin/pytest posthog/test/mcp -q` -> 368 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 340 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer The token-first ordering changes one existing expectation in both low-level resource tests: `?token=phx_...` now captures as `?token=[redacted]` rather than `?token=%5Bredacted%5D`. The token pass has already redacted the value by the time the URL is parsed, so the URL rewrite finds nothing changed and returns the string as-is. Still fully redacted, only the encoding differs; the JS sibling will land on the same form. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): strip intent PII before the generic pass rewrites URLs What changed `sanitize_event` ran `redact_pii(sanitize_captured_value(intent))`. The generic pass rewrites any URL it finds and a rewritten query percent-encodes `@`, so `Open https://example.com/?email=alice@example.com&token=fakesecret` reached `redact_pii` as `email=alice%40example.com` and the email pattern no longer matched it — an address main would have redacted. The two passes are now `sanitize_captured_value(redact_pii(intent))`: PII first, while the narration is still the raw string the agent wrote, then the generic redaction. The comment above it says why the order matters. How tested The intent composition test is now parametrized, with the existing token+email row and the new URL row asserting the exact captured value: `Open https://example.com/?email=%5Bredacted%5D&token=%5Bredacted%5D` — the address and the token are gone, the host is still there. - `.venv/bin/pytest posthog/test/mcp -q` -> 369 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 341 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): keep tool names out of the entropy detector What changed Widening the `resource_name` gate in `sanitize_event` (df1c814) sent tool names through `sanitize_captured_value`, whose entropy detector reads a legitimate identifier as a credential: a call to `Get_Organization_Memberships` reported `$mcp_tool_name: "[redacted]"`, which breaks per-tool attribution. A `resource_name` is only ever an identifier or a uri, so it now runs through `_sanitize_resource_name`: PostHog-token redaction then the URL pass, and neither the entropy detector nor the base64 gate. A name with no url in it passes through untouched. How tested New parametrized `test_sanitize_event_resource_name_keeps_identifiers_and_redacts_uris` covers all three shapes: a `$mcp_tool_call` name kept verbatim, an `$identify` uri with its userinfo redacted, and a read uri with its `?token=` redacted. `test_identify_on_a_resource_read_is_named_by_the_uri` still passes. - `.venv/bin/pytest posthog/test/mcp -q` -> 372 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 344 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer No parity change for @posthog/mcp: it has no entropy pass, so its `resourceName` was never at risk. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): redact credentials in resource uris with no authority What changed An MCP resource uri need not have an authority, so `resource:guide?token=...` and `file:/guide.md?token=...` never reached the URL pass and shipped their token in `$mcp_resource_name` and `$mcp_parameters`. The `//` is now optional in `_URL_PATTERN` (`[^\s<>"]+` absorbs a `//host` when there is one), and the nested-value check in `_sanitize_url_field_value` asks the pattern instead of looking for `://`, so `?url=resource:guide?token=x` is covered too. The looser pattern over-matches prose (`Error:foo`, `at12:30`, `C:\path`); a comment says why that is harmless — a match with nothing to redact is returned byte-for-byte and never re-serialized. How tested Rows added to `test_sanitize_url_credentials` for both uri forms and for the three byte-for-byte prose cases, a `resource:` row added to the resource_name test, and a focused read test in both low-level suites asserting the token is absent from every capture. - `.venv/bin/pytest posthog/test/mcp -q` -> 379 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 351 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer Two expectations differ from the ones proposed, and both are asserted as observed: - `file:/guide.md?token=...` re-serializes as `file:///guide.md?token=...`; `urlunsplit` restores the empty authority, and JS's `new URL()` does the same. - `resource:guide?token=fakesecret` through `sanitize_captured_value` (the `$mcp_parameters` path) comes back as a bare `[redacted]`: the URL pass rewrites it to `resource:guide?token=%5Bredacted%5D`, and the entropy detector that runs after it for free-text values reads that rewritten string as a credential. `$mcp_resource_name` skips that pass and keeps the readable `resource:guide?token=%5Bredacted%5D`. The token is gone on both paths, but @posthog/mcp has no entropy pass, so its `$mcp_parameters` will keep the readable form where Python drops the value. Flagged for a parity decision. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): keep the URL length bound off data uris What changed With the authority optional, a long unspaced `data:...;base64,...` string that is not valid base64 (so the binary-data branch deliberately keeps it) matched `_URL_PATTERN`, blew the 8192 bound and came back as `[redacted]`. The bound now applies only to a match that opens with an authority — the case it exists for, capping parsing work on an attacker-shaped URL. An over-long authority-less match is sanitized normally; parse_qsl is already bounded by its field count. How tested A 10,000+ char `data:application/octet-stream;base64,AAAA%ZZ...` row added to `test_sanitize_url_bounds`, asserted unchanged both standalone and inside prose; the existing over-length `https://...` row still yields `[redacted]`. - `.venv/bin/pytest posthog/test/mcp -q` -> 380 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 352 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): start the address at the authority, not at a prose word What changed With the authority optional, `Failed URL:https://alice:hunter2@example.com/doc` matched as one URL with scheme `URL`, so urlsplit put the whole address in the path and the userinfo was never redacted. `_sanitize_url` now looks for the first authority-bearing scheme in the match: when it starts past index 0, everything before it is prose (`URL:`, `a:b:`) and is handed back verbatim with only the remainder sanitized. Matches that already start at the authority, and authority-less uris, take the path they take today. How tested Rows added to `test_sanitize_url_credentials` for the userinfo case, the `?token=` case and the byte-for-byte `Note:https://example.com/doc`. - `.venv/bin/pytest posthog/test/mcp -q` -> 384 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 356 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer `see:resource:guide?token=fakesecret` is asserted in the resource_name test rather than the URL table: the URL pass produces the expected `see:resource:guide?token=%5Bredacted%5D`, but the entropy detector that runs after it for free-text values drops that rewritten string whole, the same known behavior as the plain `resource:guide?token=...` case. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): only skip a prefix that is really prose What changed The prefix skip took any authority found past index 0 as prose, so `file:/guide?password=hunter2&url=https://example.com` treated its own outer URI as the prefix and returned the password raw. The skip now requires the prefix to be a run of colon-suffixed words (`URL:`, `a:b:`). Anything with a `?`, `/` or `=` in it means the match is an outer URI, which is parsed whole — its query pass redacts its own credentials, and a retained value carrying the inner URL goes through the nested pass. How tested Rows added to `test_sanitize_url_credentials` for the outer-URI case, the `token=<credential>+<inner url>` case and the `a:b:` prose run; the `Failed URL:` / `URL:` / `Note:` rows are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 387 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 359 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer One byte differs from the proposed row and is asserted as observed: `resource:g?token=fakesecret+https://fakeuser:fakepass@b` comes back as a bare `[redacted]`, not `resource:g?token=%5Bredacted%5D`. The URL pass does produce that value — the entropy detector that runs after it for free-text values then drops the rewritten authority-less string whole, the same behavior already noted for `resource:guide?token=...`. The credential and the inner userinfo are gone on either path. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): split a match that runs two addresses together What changed Replaces the prose-prefix rule from b50cd56, which only covered a colon-suffixed word and missed URLs joined without whitespace: `https://example.com/doc,https://user:pw@other.example.com/doc` parsed as one address with the second one — userinfo and all — buried in the first one's path. One rule covers both shapes: split the match at the first authority that starts before its first `?` or `#`, and sanitize each part on its own. An authority after `?`/`#` is a query or fragment value, so the outer URI is parsed whole and its own field pass redacts it (a sensitive key, or the nested pass). `_PROSE_PREFIX_PATTERN` is gone; `_split_at_second_address` replaces it, and each half is strictly shorter so the recursion terminates. How tested Rows added for the joined-addresses case and for two markdown links run together; the `a:b:` / `Failed URL:` / `URL:` / `Note:` rows, the `file:/guide?password=` row and both gateway rows are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 389 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 361 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer `resource:g?token=fakesecret+https://fakeuser:fakepass@b` still asserts a bare `[redacted]`: the URL pass produces `resource:g?token=%5Bredacted%5D` and the entropy detector that follows for free-text values drops the rewritten authority-less string whole, as already noted for `resource:guide?token=...`. Every other proposed row matches byte for byte. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): run the credential detectors before the URL pass What changed The entropy detector scans a 200-char window, and rewriting a URL can grow a word past it: `https://example.com/<120 chars>?ref=ghp_<36>&token=x` was `[redacted]` on main and kept its GitHub token after the URL pass moved ahead of it. Both credential passes now run first (PostHog tokens, then the detector), and the URL pass runs last — it only redacts or percent-encodes, so it never exposes anything the detectors could have matched. `_is_secret` now strips this sanitizer's own redaction markers before judging a word. A value can be sanitized twice (a response's content blocks are), and the marker's character mix alone pushed a short uri like `resource:guide?token=%5Bredacted%5D` over the entropy bar, dropping a value we had already made safe. Stripping the marker rather than skipping the word keeps a real credential written around one detectable. Together these restore the readable form for authority-less uris through `sanitize_captured_value` — `resource:guide?token=%5Bredacted%5D` rather than a bare `[redacted]` — which is byte parity with @posthog/mcp, and makes the sanitizer idempotent on every vector in the table. How tested The three `resource:`/`see:resource:` rows now assert the readable form, the Codex `ghp_` row asserts `[redacted]`, and the authority-less read case moved back into the parametrized uri tables in both low-level suites (its standalone test is gone, name and parameters agree again). - `.venv/bin/pytest posthog/test/mcp -q` -> 391 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 363 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): redact credentials in hash-routed fragments What changed `https://example.com/#/callback?token=fakesecret` parsed its whole fragment as a field list, so the only key was `/callback?token` and nothing matched. A hash-routed URL keeps its route in the fragment: everything up to and including the first `?` is now held back verbatim by `_split_fragment_route` and only the remainder is parsed as fields, with the route restored on re-serialization. The "fragment must contain `=`" gate is unchanged, so `#/callback` and `#section-2` still pass through untouched. How tested Rows added to `test_sanitize_url_credentials` for the routed callback and for the two byte-for-byte cases; the existing `#access_token=...` and `#section-2` rows are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 394 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 366 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): only treat a leading fragment segment as a route What changed `#access_token=fakesecret&next=https://other.test/?page=1` split at the `?` inside the `next` value, so everything before it — the access token included — was held back as a verbatim route. A route comes first or not at all, so the split now only happens when no `=` precedes the `?`; otherwise the fragment is already a field list and is parsed whole. How tested A row for that fragment added to `test_sanitize_url_credentials`; the three hash-route rows and the two plain-fragment rows are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 395 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 367 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): gate binary intents before redacting their PII What changed `redact_pii` ran first on the intent, so a base64 blob holding a Luhn-valid run got a `[redacted]` spliced into it, stopped matching the base64 pattern, and was captured almost whole instead of as the binary marker. `_sanitize_string` is now split into the size/base64 gate (`_is_binary_blob`) and the text passes (`_sanitize_text`), and the new `sanitize_intent` composes them in the order that holds: binary gate, then PII, then the text passes. `sanitize_event` uses it for `user_intent`; non-string intents take the same path they took before. The docstring records why each step sits where it does: the gate first because splicing a redaction into a blob stops it looking like base64, PII before the URL pass because a rewritten URL percent-encodes the `@` the email pattern needs. How tested The blob intent added as a row to the parametrized intent test; the email-inside- a-URL row and the non-string intent test are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 396 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 368 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): sanitize an address carried in a plain fragment What changed A match ends at the first `#`, so a second address inside the fragment is never split off as one — and a fragment with no `=` was skipped entirely, publishing `[a](https://public.test/#intro)[b](https://user:password@private.test/doc)` with its credentials intact. A fragment that is not a field list is now run through the URL text pass one level deep, and counts as a change when it comes back different. Field-list fragments keep today's handling, and the trailing-suffix rule is untouched: no field was rewritten, so the suffix is re-appended. How tested The credential row and its clean counterpart added to `test_sanitize_url_credentials`; `#section-2`, the hash-route rows and the `#access_token=` row are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 398 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 370 passed, 18 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): report what a failing resource read actually raised What changed mcp 2.x wraps a failing read twice — `UnexpectedResourceError`, re-raised as the shared `MCPError` — and masks the handler's message out of both, so a `TimeoutError('storage backend timed out')` reached PostHog as `$mcp_error_type: MCPError` / `$mcp_error_message: Error reading resource file:///guide.md`. `_primary_exception` already steps past consecutive tool dispatch wrappers; the wrapper table now covers the resource pair too, so the scalars land on the handler's own exception. The `$exception` sibling keeps the whole chain as before, and the wrapper is still re-raised to the caller, so dispatch semantics are unchanged. mcp 1.x's `ResourceError` is deliberately NOT in the table: it keeps the handler's message in its own text, so reporting it loses nothing. How tested New `test_failed_read_reports_the_handler_failure` in `test_resources.py`, which runs against every high-level adapter on both SDK majors: the caller still gets the SDK wrapper, the message carries `storage backend timed out` on both, and the type is `TimeoutError` on v2 / the unchanged `ResourceError` on v1. Verified it fails on v2 without the wrapper-table change. - `.venv/bin/pytest posthog/test/mcp -q` -> 400 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 371 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): sanitize route prefixes and bound both URL recursions What changed - Route prefix leak. `#https://user:password@private.test/doc?page=1` splits at a `?` that precedes any `=`, and the route before it was kept verbatim — with its credentials. The route now goes through the same text pass a plain fragment gets, and the fragment is re-serialized when either the route or the fields changed. `_split_fragment_route` returns the route, the `?` and the fields separately, so the route is sanitized as text while the fields are re-encoded. - Fragment recursion. A `#`-chained uri (`resource:x#resource:x#...`) recursed once per `#`, to RecursionError. The fragment text passes now take the same one-level budget as a nested field value: past it, text still carrying an address is replaced with the marker instead of descended into. Depth is at most two. - Address-split recursion. Splitting a match at its second address recursed once per address. `_split_addresses` now cuts the whole match into pieces in one pass and `_sanitize_single_url` (the old non-splitting body) handles each. No piece can need splitting again: every piece but the last ends before the first `?`/`#`, and in the last piece a remaining authority sits in field data. How tested Two rows for the route prefix, and the two pathological chains asserted for their exact output. Before this commit the `#`-chain raised RecursionError; the address-chain returned a bare `[redacted]` because the length bound fired ahead of the recursion (a shorter one recursed ~470 deep and worked), so Python never crashed on that one — JS, with its own stack limit, is the reason both are covered. Every existing row is unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 404 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 375 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer The length bound moved from the whole match to the individual piece, which is what lets a long run of short addresses be sanitized rather than dropped whole. Work stays linear in the input: each piece is parsed once and its field count is still bounded. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): recognize a hash route that carries its own `=` What changed `#/docs/id=1?token=fakesecret` was read as one field named `/docs/id` with a value of `1?token=fakesecret`, so the token survived. A fragment is now a route when it reads as a path (starts with `/`) and has a `?`, or when nothing field-shaped precedes that `?` — the previous rule, which still covers `#/callback?k=v` and `#https://user:pw@host/doc?page=1`. Failing both, it is a field list when it holds a `=`, and plain text otherwise. A query key's segments are now also `/`-delimited, so the `/token` of a `#/token=fakesecret` fragment is recognized as the credential name it is. This widens redaction for every key with a `/` in it, in queries as well as fragments (`a/token`, `sort/key`) — the same over-redaction trade the segment rule already documents. How tested Rows added for the route with a `=`, its byte-for-byte counterpart without a `?`, and the leading-slash field list; the existing route, field-list and plain-fragment rows are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 407 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 378 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer `#/token=fakesecret` comes back as `#%2Ftoken=%5Bredacted%5D`, not `#/token=%5Bredacted%5D`: re-serializing a field percent-encodes a `/` in its key. `URLSearchParams` does the same, so the two SDKs still agree byte for byte. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): split an adjacent address wherever it sits What changed `_split_addresses` stopped looking at the first `?` or `#`, so an address that followed one — `[a](https://public.test/?download)[b](https://alice:pw@private.test/doc)` — was absorbed into a query KEY, and keys are never sanitized. The boundary is gone: every authority start past index 0 splits, except one preceded by `=`. An address in value position belongs to the field that holds it and the nested pass sanitizes it there; anywhere else it is simply the next address. The single-pass structure and `_sanitize_single_url` are unchanged, and the docstring now explains value position vs adjacent address. How tested Three rows added (query key, after a comma inside a query, straight after the `?`). Every existing row is unchanged, including both gateway `?url=https://...` rows, `#access_token=…&next=https://…` and `?token=…+https://…` — all value position, none split — and the two pathological inputs still finish in ~10ms. - `.venv/bin/pytest posthog/test/mcp -q` -> 410 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 381 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): judge each half of a fragment on its own What changed `#/token=fakesecret&next=https://other.test/?page=1` starts with `/` and holds a `?`, so the route heuristic took everything before that `?` — the token included — as text to keep verbatim. Shape cannot separate that from `#/docs/id=1?token=x`, so the heuristics are gone: a fragment splits at its first `?`, and each half gets the field pass when it holds a `=` and the text pass when it does not. The halves are reassembled around the `?`, each keeping its own encoding. `_split_fragment_route` is replaced by `_sanitize_fragment_part`, which also reports whether it rewrote its last field, and `_rewrote_the_last_field` keeps the trailing-punctuation rule readable now that it has two halves to consider. How tested The new fragment row asserts the leading-slash field list with a nested address; the existing `#access_token=...&next=...` row changes as expected (its half is re-serialized on its own, so the `?page=1` after it stays verbatim rather than being encoded into a value). Every other fragment, route, markdown and pathological row keeps its expectation, and the long inputs still run in ~15ms. - `.venv/bin/pytest posthog/test/mcp -q` -> 411 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 382 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): decide value position by the nearest structural character What changed `?token=foo%20https://secret.test/private` split at the inner authority because only the character right before it was checked, cutting the token's value in two and publishing the tail beside the redaction. The decision now scans back to the nearest character of `=&;?#/`: a `=` means the authority is inside a field's value, so the nested pass handles it there; a field separator, a `/`, or nothing at all means a new address begins. `/` is in the set so a `=` inside a path (`/a=b/c,https://...`) does not read as a field. How tested Rows added for the split value, for the path-with-`=`, and for the comma inside a query — that last one is now value position, so it goes through the nested pass and is re-encoded as part of its field (`q=see%2Chttps%3A%2F%2F%255Bredacted...`) rather than being split off. Every other row is unchanged, gateway and pathological inputs included; the long inputs still run in ~16ms, since each backward scan stops at the previous address's own `/`. - `.venv/bin/pytest posthog/test/mcp -q` -> 413 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 384 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): bound value position to the fields region, fail closed on a split credential What changed - A `=` in the path is not a field. `https://example.com/redirect=https://user:pw@host/doc` kept the inner address attached, so nothing sanitized it. Value position now only exists past the first `?` or `#`: before that, an authority always starts its own address. With the region check in place, `/` leaves `_FIELD_STRUCTURE_CHARACTERS` — the path case it was there for is covered. - A `?` can be a character of a credential. `#password=prefix?fakesecret` split into a redacted head and a tail that published the rest of the password. Nothing can tell that `?` from a real boundary, so when the head ends in a value just redacted, the whole tail is replaced with the marker rather than sanitized. How tested Rows added for the path `=`, and for a fragment credential split by a `?` both with and without field-shaped text after it; `#/docs/id=1?token=...` still sanitizes its tail normally, since that head's last field is not sensitive. Every other row is unchanged and the pathological inputs still run in ~15ms. - `.venv/bin/pytest posthog/test/mcp -q` -> 416 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 387 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): find value position in one forward pass What changed `_in_value_position` scanned backwards from every authority to the nearest structural character. Once `/` left that set, a value whose addresses all sit after the same `?` made every scan run back to it: `"https://a.test/?" + "https://b.test/x," * 4000` (68 KB) took 2.2 s, synchronous on the server's event loop. `_authority_starts` now walks the value once, pairing each authority start with the last structural character seen before it, and `_split_addresses` reads value position off that. Same decisions, linear time — the same input is 7 ms. How tested `test_sanitize_url_is_not_quadratic_on_many_addresses` asserts that input comes back unchanged in under a second, next to the existing quadratic-PII test. Every row in the URL table keeps its expectation, and the two other pathological inputs still run in ~10ms and ~19ms. - `.venv/bin/pytest posthog/test/mcp -q` -> 417 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 388 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): tell a URL's delimiters from the same characters inside a value What changed - `?token=prefix?https://secret.example/private` split at the inner authority, because the forward pass counted every `?` as structural — including the one inside the token's own value. Only three positions actually divide a URL: the query's `?`, the fragment's `#`, and the `?` that splits the fragment into head and tail. `_structural_delimiters` computes them per value, and every other `?`/`#` is ordinary text, so the address after it stays with its field and the field pass redacts the lot. - `#password=phx_...?private-suffix` kept its tail: the PostHog-token pass had already rewritten that value, so comparing before and after found no change and the fail-closed rule never fired. A field list now reports whether its last field is SENSITIVE — its key is a credential name, or its value changed — and that flag drives both the fail-closed fragment tail and the trailing-punctuation rule. How tested A row for each: the value-internal `?`, and the already-redacted password whose head is byte-identical. Every existing row keeps its expectation, and all three pathological inputs still run in ~20ms or less. - `.venv/bin/pytest posthog/test/mcp -q` -> 419 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 390 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): stop splitting a credential at a semicolon What changed Every `;` was normalized to `&` before parsing, so `?password=prefix;remainingsecret` was captured as `password=%5Bredacted%5D&remainingsecret=` — the tail of the password published as a field of its own. A `;` is a legacy field separator to some servers and an ordinary character to others, and the value cannot say which, so parsing now splits on `&` alone and the decision fails closed: a sensitive key redacts its whole value (the `;` tail with it), and a value under a non-sensitive key is redacted whole when any `;`-separated piece of it names a credential. The 128-field bound counts `&` only, which is what `parse_qsl` was already doing. How tested Rows for the split password, the legacy `a=1;token=...` field (now redacted as one value rather than re-serialized into two) and a benign `a=1;b=2` that stays byte-for-byte. - `.venv/bin/pytest posthog/test/mcp -q` -> 421 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 392 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Notes for the reviewer The README's URL paragraph never described the `;` normalization, so nothing there needed changing. Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): redact an intent's credentials before its PII What changed `sanitize_intent` ran `redact_pii` first, and a PII pattern can cut a token in half: the phone pattern reads the middle of `phx_AAAAAAAA-415-555-0142-AAAAAAAAAAAAAAAAAAAA` as a number, so the intent kept `phx_AAAAAAAA-[redacted]-AAAAAAAAAAAAAAAAAAAA` where the token pass would have taken the whole thing. `_sanitize_text` is split into `_redact_credentials` (the PostHog-token pass, then the entropy detector) and the URL pass, and the intent now runs binary gate, credentials, PII, URLs. Every other captured string keeps credentials-then-URLs, unchanged. The ordering comment moves onto `_redact_credentials` and `sanitize_intent`: credentials before PII because a PII pattern can cut a token in half, PII before the URL pass because a rewritten URL percent-encodes the `@` the email pattern needs. How tested A row for that token; the email-inside-a-URL, PostHog-token-plus-email, phone-and-email and binary-blob intent rows are unchanged. - `.venv/bin/pytest posthog/test/mcp -q` -> 422 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 393 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7 * fix(mcp): keep a credential whole across `;` and a fragment's `?` What changed - A `;` inside a KEY named nothing. `?download;token=fakesecret` parses to the key `download;token`, and only values were checked for `;`-separated credentials. `;` joins the segment separators of the sensitive-key pattern, so such a key is recognized and its field redacted. - A credential's suffix was split away before anything could fail closed. `#password=prefix?https://private.example/remainingsecret` treated the fragment's `?` as structure and the address behind it as adjacent, so the fail-closed tail rule never saw it; `?password=prefix;https://...` did the same through `;`. In `_authority_starts`, `;` is no longer a field separator (fields parse on `&` alone, so a `;` belongs to whatever value holds it), and the fragment's `?` divides only when no value is already open — straight after a `=` the address stays with its field, where the fragment split and its fail-closed rule can see the whole of it. How tested Five rows: the `;` key, the `;` and `?` suffixes, a fragment head with no credential (so its tail is sanitized as text rather than dropped), and a `?` that opens no value, which still divides. Every existing row is unchanged and all three pathological inputs still run in ~20ms or less. - `.venv/bin/pytest posthog/test/mcp -q` -> 427 passed - `.venv-mcp-v2/bin/pytest posthog/test/mcp -q` -> 398 passed, 19 skipped - `.venv/bin/ruff check posthog/mcp posthog/test/mcp` -> All checks passed! - `.venv/bin/ruff format --check posthog/mcp posthog/test/mcp` -> 60 files already formatted - `uv run mypy --no-site-packages --config-file mypy.ini . | uv run mypy-baseline filter` -> Success: no issues found in 231 source files Claude-Session: https://claude.ai/code/session_01VGVQTsHUk5dmQC2rGPEgc7
1 parent a1002c5 commit 97ea5fd

15 files changed

Lines changed: 1821 additions & 45 deletions
Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
---
2+
pypi/posthog: minor
3+
---
4+
5+
Capture MCP resource discovery and reads from instrumented servers. URL credential redaction (userinfo, credential-named query and fragment parameters) now applies to every captured string, including existing `$mcp_tool_call` parameters, responses and error messages, so URLs in existing tool-call data will show `%5Bredacted%5D` values after upgrading.

‎posthog/mcp/README.md‎

Lines changed: 16 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,22 @@
11
# PostHog MCP analytics
22

33
Product analytics for Model Context Protocol servers. Wrap a Python MCP server so
4-
every tool call, agent intent, and failure is captured to PostHog as a `$mcp_*` event.
4+
tool calls, agent intent, resource discovery and reads, and failures are captured
5+
to PostHog as `$mcp_*` events.
6+
7+
Resource bodies are not captured. Resource and resource-template listings are:
8+
a listing is metadata (names, uris, mime types), so `$mcp_resources_list` carries
9+
it as `$mcp_response`.
10+
11+
Captured URLs redact usernames, passwords, and credential-named query and
12+
fragment parameters, including signed URL credentials. This applies to every
13+
captured string, tool call parameters, responses and error messages included, so
14+
it also covers a failed read that repeats the URL in its error message. URLs
15+
longer than 8,192 characters or with more than 128 query fields are redacted
16+
entirely to bound parsing work. Other query and fragment parameters can still
17+
contain application-specific sensitive data. Requests and responses keep their
18+
original addresses. Use `before_send` to remove any additional
19+
application-specific sensitive data.
520

621
```python
722
from posthog import Posthog

‎posthog/mcp/__init__.py‎

Lines changed: 7 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -4,10 +4,10 @@
44

55
"""PostHog MCP analytics SDK — product analytics for Model Context Protocol servers.
66
7-
Wrap a Python MCP server so every tool call, agent intent, and failure is
8-
captured to PostHog as a ``$mcp_*`` event. Works with the MCP Python SDK 1.x
9-
*and* 2.x (the 2026-07-28 spec revision) — the high-level server class moved
10-
between majors, but ``instrument()`` is the same::
7+
Wrap a Python MCP server so tool calls, agent intent, resource discovery and
8+
reads, and failures are captured to PostHog as ``$mcp_*`` events. Works with
9+
the MCP Python SDK 1.x *and* 2.x (the 2026-07-28 spec revision) — the high-level
10+
server class moved between majors, but ``instrument()`` is the same::
1111
1212
from posthog import Posthog
1313
from posthog.mcp import instrument
@@ -219,9 +219,9 @@ def instrument(
219219
posthog_client: Optional[Client] = None,
220220
options: Optional[MCPAnalyticsOptions] = None,
221221
) -> McpAnalytics:
222-
"""Instrument an MCP server so PostHog auto-captures tool calls, tool listings,
223-
initialize, identity, and exceptions. Returns a handle whose ``capture()``
224-
records custom events.
222+
"""Instrument an MCP server so PostHog auto-captures tool calls, tool and
223+
resource listings, resource reads, initialize, identity, and exceptions.
224+
Returns a handle whose ``capture()`` records custom events.
225225
226226
Idempotent per server instance — a second call reuses the existing tracking
227227
state instead of double-wrapping. Degrades to a no-op handle on any failure so

‎posthog/mcp/_instrument_fastmcp.py‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -27,6 +27,7 @@
2727
import mcp.types as mcp_types
2828

2929
from ._conversation_id import build_prompt_back
30+
from ._instrument_lowlevel import _wrap_resource_requests
3031
from ._instrumentation import (
3132
_to_jsonable,
3233
append_get_more_tools,
@@ -58,6 +59,9 @@ def instrument_fastmcp(server: Any, data: MCPAnalyticsData) -> None:
5859
data.server_version = getattr(getattr(server, "_mcp_server", None), "version", None)
5960
_wrap_tool_manager_call(server, data)
6061
_wrap_list_tools_handler(server, data)
62+
low_level = getattr(server, "_mcp_server", None)
63+
if low_level is not None:
64+
_wrap_resource_requests(low_level, data)
6165

6266

6367
# --- tool call seam ----------------------------------------------------------

‎posthog/mcp/_instrument_lowlevel.py‎

Lines changed: 93 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -22,13 +22,17 @@
2222

2323
from ._context_parameters import schema_has_param
2424
from ._conversation_id import build_prompt_back
25+
from ._event_types import MCPAnalyticsEventType
2526
from ._instrumentation import (
2627
_to_jsonable,
2728
append_get_more_tools,
2829
collect_listed_tools,
2930
extract_tools,
3031
mutate_tool_schema,
32+
prepare_request,
33+
record_resource_request,
3134
request_to_dict,
35+
resource_listing_response,
3236
resolve_session_and_client,
3337
start_tool_call_lifecycle,
3438
start_tools_list_lifecycle,
@@ -50,6 +54,7 @@ def instrument_low_level(server: Any, data: MCPAnalyticsData) -> None:
5054
data.server_version = getattr(server, "version", None)
5155
_wrap_call_tool(server, data, strip_injected=False)
5256
_wrap_list_tools(server, data, context_required=False)
57+
_wrap_resource_requests(server, data)
5358

5459

5560
def instrument_fastmcp_v2(server: Any, data: MCPAnalyticsData) -> None:
@@ -73,6 +78,94 @@ def instrument_fastmcp_v2(server: Any, data: MCPAnalyticsData) -> None:
7378
# sees: under `FastMCP(strict_input_validation=True)` every call fails with
7479
# "'context' is a required property".
7580
_wrap_list_tools(low_level, data, context_required=False)
81+
_wrap_resource_requests(low_level, data)
82+
83+
84+
def _wrap_resource_requests(server: Any, data: MCPAnalyticsData) -> None:
85+
for request_type, event_type in (
86+
(mcp_types.ListResourcesRequest, MCPAnalyticsEventType.MCP_RESOURCES_LIST),
87+
# Templates are listings too: the captured request method separates
88+
# `resources/templates/list` from `resources/list` on the same event.
89+
(
90+
mcp_types.ListResourceTemplatesRequest,
91+
MCPAnalyticsEventType.MCP_RESOURCES_LIST,
92+
),
93+
(mcp_types.ReadResourceRequest, MCPAnalyticsEventType.MCP_RESOURCES_READ),
94+
):
95+
_wrap_resource_request(server, data, request_type, event_type)
96+
97+
98+
def _wrap_resource_request(
99+
server: Any,
100+
data: MCPAnalyticsData,
101+
request_type: Any,
102+
event_type: str,
103+
) -> None:
104+
handlers = server.request_handlers
105+
original = handlers.get(request_type)
106+
if original is None or getattr(original, _WRAPPED_FLAG, False):
107+
return
108+
109+
async def handler(req: Any) -> Any:
110+
client_name, client_version = _client_info(server)
111+
protocol_version = _protocol_version(server)
112+
mcp_session_id = _mcp_session_id(server)
113+
token, client_name, client_version, protocol_version = (
114+
resolve_session_and_client(
115+
mcp_session_id, client_name, client_version, protocol_version
116+
)
117+
)
118+
request = request_to_dict(req)
119+
extra = {"session_id": mcp_session_id, "ctx": _request_context(server)}
120+
try:
121+
session_id = await prepare_request(
122+
data,
123+
mcp_session_id=mcp_session_id,
124+
client_name=client_name,
125+
client_version=client_version,
126+
protocol_version=protocol_version,
127+
request=request,
128+
extra=extra,
129+
token=token,
130+
)
131+
except Exception as error: # noqa: BLE001 - analytics must not break resources
132+
log(f"Warning: could not prepare resource analytics: {error}")
133+
return await original(req)
134+
135+
start = time.monotonic()
136+
try:
137+
result = await original(req)
138+
except Exception as error:
139+
await record_resource_request(
140+
data,
141+
session_id,
142+
event_type=event_type,
143+
request=request,
144+
error=error,
145+
duration_ms=(time.monotonic() - start) * 1000,
146+
client_name=client_name,
147+
client_version=client_version,
148+
protocol_version=protocol_version,
149+
extra=extra,
150+
)
151+
raise
152+
153+
await record_resource_request(
154+
data,
155+
session_id,
156+
event_type=event_type,
157+
request=request,
158+
response=resource_listing_response(event_type, result),
159+
duration_ms=(time.monotonic() - start) * 1000,
160+
client_name=client_name,
161+
client_version=client_version,
162+
protocol_version=protocol_version,
163+
extra=extra,
164+
)
165+
return result
166+
167+
setattr(handler, _WRAPPED_FLAG, True)
168+
handlers[request_type] = handler
76169

77170

78171
def _wrap_call_tool(

‎posthog/mcp/_instrument_v2.py‎

Lines changed: 82 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -37,11 +37,15 @@
3737

3838
from ._context_parameters import schema_has_param
3939
from ._conversation_id import build_prompt_back
40+
from ._event_types import MCPAnalyticsEventType
4041
from ._instrumentation import (
4142
_to_jsonable,
4243
collect_listed_tools,
4344
mutate_tool_schema,
4445
params_to_request_dict,
46+
prepare_request,
47+
record_resource_request,
48+
resource_listing_response,
4549
resolve_session_and_client,
4650
start_tool_call_lifecycle,
4751
start_tools_list_lifecycle,
@@ -68,6 +72,13 @@
6872
# injected `context` parameter per entry point (see _wrap_v2_list_tools).
6973
_CALL_METHOD = "tools/call"
7074
_LIST_METHOD = "tools/list"
75+
_RESOURCE_METHODS = {
76+
"resources/list": MCPAnalyticsEventType.MCP_RESOURCES_LIST,
77+
# Templates are listings too: the captured request method separates
78+
# `resources/templates/list` from `resources/list` on the same event.
79+
"resources/templates/list": MCPAnalyticsEventType.MCP_RESOURCES_LIST,
80+
"resources/read": MCPAnalyticsEventType.MCP_RESOURCES_READ,
81+
}
7182

7283

7384
def instrument_mcpserver_v2(server: Any, data: MCPAnalyticsData) -> None:
@@ -86,6 +97,8 @@ def instrument_mcpserver_v2(server: Any, data: MCPAnalyticsData) -> None:
8697
)
8798
_wrap_tool_manager_call_v2(server, data)
8899
_wrap_v2_list_tools(low_level, data, context_required=True, high_level=server)
100+
for method, event_type in _RESOURCE_METHODS.items():
101+
_wrap_v2_resource_request(low_level, data, method, event_type)
89102
_patch_add_request_handler(low_level, data, wrap_call=False, high_level=server)
90103

91104

@@ -98,6 +111,8 @@ def instrument_lowlevel_v2(server: Any, data: MCPAnalyticsData) -> None:
98111
data.server_version = getattr(server, "version", None)
99112
_wrap_v2_call_tool(server, data)
100113
_wrap_v2_list_tools(server, data, context_required=False)
114+
for method, event_type in _RESOURCE_METHODS.items():
115+
_wrap_v2_resource_request(server, data, method, event_type)
101116
_patch_add_request_handler(server, data, wrap_call=True)
102117

103118

@@ -132,6 +147,8 @@ def add_request_handler(method: str, params_type: Any, handler: Any) -> None:
132147
context_required=high_level is not None,
133148
high_level=high_level,
134149
)
150+
elif method in _RESOURCE_METHODS:
151+
_wrap_v2_resource_request(server, data, method, _RESOURCE_METHODS[method])
135152

136153
setattr(add_request_handler, _WRAPPED_FLAG, True)
137154
server.add_request_handler = add_request_handler
@@ -465,6 +482,71 @@ async def handler(ctx: Any, params: Any) -> Any:
465482
_replace_handler(server, _CALL_METHOD, handler, entry.params_type)
466483

467484

485+
def _wrap_v2_resource_request(
486+
server: Any, data: MCPAnalyticsData, method: str, event_type: str
487+
) -> None:
488+
entry = server.get_request_handler(method)
489+
if entry is None or getattr(entry.handler, _WRAPPED_FLAG, False):
490+
return
491+
original = entry.handler
492+
493+
async def handler(ctx: Any, params: Any) -> Any:
494+
token, client_name, client_version, protocol_version, mcp_session_id = (
495+
_resolve_ctx(ctx)
496+
)
497+
request = params_to_request_dict(method, params, by_alias=True)
498+
extra: Dict[str, Any] = {"session_id": mcp_session_id, "ctx": ctx}
499+
try:
500+
session_id = await prepare_request(
501+
data,
502+
mcp_session_id=mcp_session_id,
503+
client_name=client_name,
504+
client_version=client_version,
505+
protocol_version=protocol_version,
506+
request=request,
507+
extra=extra,
508+
token=token,
509+
)
510+
except Exception as error: # noqa: BLE001 - analytics must not break resources
511+
log(f"Warning: could not prepare resource analytics: {error}")
512+
return await original(ctx, params)
513+
514+
start = time.monotonic()
515+
try:
516+
result = await original(ctx, params)
517+
except Exception as error:
518+
await record_resource_request(
519+
data,
520+
session_id,
521+
event_type=event_type,
522+
request=request,
523+
error=error,
524+
duration_ms=(time.monotonic() - start) * 1000,
525+
client_name=client_name,
526+
client_version=client_version,
527+
protocol_version=protocol_version,
528+
extra=extra,
529+
)
530+
raise
531+
532+
await record_resource_request(
533+
data,
534+
session_id,
535+
event_type=event_type,
536+
request=request,
537+
response=resource_listing_response(event_type, result),
538+
duration_ms=(time.monotonic() - start) * 1000,
539+
client_name=client_name,
540+
client_version=client_version,
541+
protocol_version=protocol_version,
542+
extra=extra,
543+
)
544+
return result
545+
546+
setattr(handler, _WRAPPED_FLAG, True)
547+
_replace_handler(server, method, handler, entry.params_type)
548+
549+
468550
# --- tools/list -------------------------------------------------------------------
469551

470552

‎posthog/mcp/_instrumentation.py‎

Lines changed: 55 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -2,10 +2,9 @@
22
# Copyright (c) 2025 MCPcat
33
# Licensed under the MIT License: https://github.com/MCPCat/mcpcat-typescript-sdk/blob/main/LICENSE
44

5-
"""Shared tool-call / tools-list / initialize lifecycle used by both the FastMCP
6-
and low-level server adapters. The adapters resolve transport-specific details
7-
(client info, session id, raw result shape) and delegate the analytics flow here
8-
so both stay in sync."""
5+
"""Shared MCP request lifecycles used by both the FastMCP and low-level server
6+
adapters. The adapters resolve transport-specific details (client info, session
7+
id, raw result shape) and delegate analytics policy here so both stay in sync."""
98

109
from __future__ import annotations
1110

@@ -904,3 +903,55 @@ async def record_tools_list(
904903
fire_and_forget(capture_event(data, event), data)
905904
except Exception as err: # noqa: BLE001 - isolate analytics from the tool path
906905
log(f"record_tools_list failed (event dropped): {err}")
906+
907+
908+
def resource_listing_response(event_type: str, result: Any) -> Any:
909+
"""The result an adapter should capture as the event ``response``. A listing
910+
(``resources/list``, ``resources/templates/list``) is metadata — names, uris,
911+
mime types — so it is captured; a read's result is the resource body itself,
912+
which this SDK never captures."""
913+
if event_type != MCPAnalyticsEventType.MCP_RESOURCES_LIST:
914+
return None
915+
return _to_jsonable(result)
916+
917+
918+
async def record_resource_request(
919+
data: MCPAnalyticsData,
920+
session_id: str,
921+
*,
922+
event_type: str,
923+
request: Dict[str, Any],
924+
response: Any = None,
925+
error: Any = None,
926+
duration_ms: Optional[float] = None,
927+
client_name: Optional[str] = None,
928+
client_version: Optional[str] = None,
929+
protocol_version: Optional[str] = None,
930+
extra: Optional[Dict[str, Any]] = None,
931+
) -> None:
932+
"""Record a resources listing or read without affecting dispatch."""
933+
try:
934+
params = request.get("params")
935+
uri = params.get("uri") if isinstance(params, dict) else None
936+
event: Dict[str, Any] = {
937+
"event_type": event_type,
938+
"session_id": session_id,
939+
"resource_name": uri
940+
if event_type == MCPAnalyticsEventType.MCP_RESOURCES_READ
941+
else None,
942+
"parameters": build_captured_mcp_parameters(request),
943+
"response": _wrap_response(response) if response is not None else None,
944+
"duration": duration_ms,
945+
"client_name": client_name,
946+
"client_version": client_version,
947+
"protocol_version": protocol_version,
948+
"is_error": error is not None,
949+
"timestamp": datetime.now(timezone.utc),
950+
}
951+
if error is not None:
952+
event["error"] = capture_exception(error)
953+
await _apply_event_properties(data, event, request, extra)
954+
stamp_transport_identity(event, extra)
955+
fire_and_forget(capture_event(data, event), data)
956+
except Exception as err: # noqa: BLE001 - isolate analytics from the request path
957+
log(f"record_resource_request failed (event dropped): {err}")

0 commit comments

Comments
 (0)