feat(gate): check that a lesson's cited source actually exists - #1768
Conversation
The existing gates each cover one axis and none of them covers truth: `lesson_gate.py` checks structure, DCO checks signatures, `injection_scan.py` checks injection shapes. So on 2026-09-16 four lesson PRs (#1713–#1716) passed **24/24 checks** while every one of them cited `https://github.com/modelcontextprotocol/mcp-memory-service/issues/1652` — a repository that returns 404, dressed up with `evidence_level: E3`. In a corpus whose value is *verifiable* failure memory, an invented source is worse than an absent one: it looks checkable and is not. `scripts/check_provenance.py` resolves the URLs a lesson cites in its frontmatter, with two rule tiers that reuse the strict-new/advisory split the lesson gate already uses (#1506) so legacy debt cannot block unrelated PRs: - tier 1 (always fails): a placeholder URL (`<owner>`, `TODO`, `{repo}`, `/xxx/`), or a URL confirmed dead (404/410) that is not recorded in the baseline. - tier 2 (new files only): `evidence_level: E2`/`E3` with no resolvable source at all. The statuses are deliberately conservative — a timeout, a DNS failure, a rate limit or a 5xx is `unknown` and never fails the build. A gate that goes red on someone else's outage gets bypassed within a week, so this one fails on evidence only, never on the absence of it. Private-IP and example.com URLs are exempt by construction: a lesson about a corporate proxy has to be able to say `http://172.19.128.1:7890`. `data/provenance-baseline.json` is the debt register for known-dead links; placeholders are never accepted there (there is a test for that). The corpus currently passes with an empty baseline: 439 lessons, 31 citations, 30 ok, 1 unknown (a Google docs URL this network cannot reach — reported, not failed). Red-teamed against the real case: the #1713 lesson file now exits 1 with "source does not resolve — ... (HTTP 404)". The scheduled weekly sweep is report-only, for link rot. 27 tests, offline. Runs on PRs touching lessons/** and on a schedule. Signed-off-by: Ikalus1988 <136884451+Ikalus1988@users.noreply.github.com>
🧾 Audit Report — PR #1768 (35f5a24)📊 Quality Score🔏 DCO Audit✅ All commits signed-off. 📏 PR Size
🔐 Secret Scan✅ No hardcoded secrets detected. 📦 Dependency Audit⏭️ Skipped; no Python/JS dependency files changed. 🧪 Test Suite✅ PASS — 53% coverage 📋 Lesson Schema✅ All lessons valid. ⚖️ Verdict✅ All gates passed. Ready for merge. Scope: |
PR Reviewer Guide 🔍Here are some key observations to aid the review process:
|
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
misakanet-web | 9aefbf8 | Commit Preview URL Branch Preview URL |
Sep 16 2026, 03:26 PM |
…mbers (#1769) The public roadmap had not been updated in three weeks and had drifted away from the repository in ways a newcomer could not detect: it advertised v2.18.0 (really 2.30.2), three MCP tools (really 7 remote / 9 local), and `python scripts/site_health.py` — a file that does not exist. Two "done" items named scripts that never existed in git history (`freshness_scorer.py`, `gap_analyzer.py`); #1165 shipped only into the local stdio server and is not in the remote tool set; GX1 was explicitly reverted in 42e374345 and still read as a milestone. What this does: - `ROADMAP.md` — a dated `2026-09-16 更新` section: a verified status snapshot (every number reproducible via the commands in the appendix), a 39-row adjudication of the old items (done / stale / abandoned, each with a reason and evidence), the six new priorities phrased as pickable work, and an explicit "what we are still not doing" list. Old text is preserved verbatim with a one-line status marker under each old heading; the file header now says it is the only outward-facing roadmap. - `docs/rfc-280-90-day-roadmap.md` — a dated historical header pointing at `ROADMAP.md`. Its Vision 1/2 were adopted and exceeded, Vision 3 was never started as a product line, and Vision 4 (federated / enterprise) still has no PRD. Its "hybrid search with embeddings" days directly contradict the standing anti-embedding position, which is worth knowing before anyone re-proposes it. - `docs/maintainer/issue-pr-status-2026-09-16.md` — a machine-read snapshot of PR/issue state (snapshot 2026-09-16T14:59Z): 14 open PRs of which only 2 are merge-ready; 43 open issues; 8.3 issues/day and 16.6 PRs/day over 30 days, of which only 1.3–2.2 intakes/day actually need human judgement. The bottleneck is PR close-out and duplicate submissions, not triage volume — 22 of 43 open issues are waiting on a maintainer reply, and 18 of the 44 unmerged closes in 7 days were re-submissions of the same title or superseded work. Also marks priority ① (the provenance gate) as done in #1768, and states plainly what that gate still cannot check (semantic truth). Signed-off-by: Ikalus1988 <136884451+Ikalus1988@users.noreply.github.com>
PR Code Suggestions ✨Explore these optional code suggestions:
|
🎉 Merged — Thank you!Your contribution has been merged into main. PR: #1768 — feat(gate): check that a lesson's cited source actually exists What's next:
Welcome to the MisakaNet contributor community! 🧠 |
|
✅ Merged! Thanks again, @Ikalus1988. feat(gate): check that a lesson's cited source actually exists (+801 lines, 5 files) Quick question — did any MisakaNet lesson help you this time? No need to reply if nothing comes to mind. ⚡ |
…#1771) Found by red-teaming the gate from #1768 rather than by reading it. The probe lesson I wrote to prove the gate fires used the JSON-style frontmatter that **64 lessons in this corpus** use: ```json { "title": "Probe", "evidence_level": "E3", "source": "https://github.com/modelcontextprotocol/mcp-memory-service/issues/1652" } ``` The first version of the parser matched `^\\s*([A-Za-z_]+)\\s*:` — an unquoted key. Inside a JSON block every key is quoted, so the citation list came back **empty**: the file could cite anything it liked and the gate would report nothing. `evidence_level` had the same hole, and its regex also refused the indentation JSON uses. Both now accept an optional quote around the key. Two tests cover the JSON shape, one of them the end-to-end case (a 404 inside a JSON block must produce a failure). Nothing in the corpus is hidden this way today — none of the 64 JSON-frontmatter lessons cites a URL — so this is a prospective hole rather than a live one. That is the point of the probe: the previous gate was written, reviewed and tested, and still had a hole that only a deliberate attempt to defeat it exposed. 31 tests, offline. Signed-off-by: Ikalus1988 <136884451+Ikalus1988@users.noreply.github.com>
User description
Why
The gates each cover one axis, and none of them covers truth:
scripts/lesson_gate.pySigned-off-byscripts/injection_scan.pyOn 2026-09-16 four lesson PRs — #1713, #1714, #1715, #1716 — passed 24/24 checks while all four cited
https://github.com/modelcontextprotocol/mcp-memory-service/issues/1652, withevidence_level: E3. That repository does not exist:In a corpus whose value proposition is verifiable failure memory, an invented source is worse than an absent one: it looks checkable and is not. Today that class of pollution is caught only by a human reading every PR — which is exactly the bottleneck this removes.
What it does
scripts/check_provenance.pyresolves the URLs a lesson cites in its frontmatter, in two tiers that reuse the strict-new/advisory split the lesson gate already uses (from #1506), so a 440-lesson legacy corpus cannot turn unrelated PRs red:<owner>,TODO,{repo},/xxx/,...), or a URL confirmed dead (404/410) that is not recorded in the baseline.evidence_level: E2/E3while citing no resolvable source. Advisory, printed with the fix.Deliberately conservative statuses: a timeout, DNS failure, rate limit (403/429) or 5xx is
unknownand never fails the build, because a gate that goes red on someone else's outage gets bypassed within a week. Private-IP andexample.comURLs are exempt by construction — a lesson about a corporate proxy has to be able to sayhttp://172.19.128.1:7890, which is already in the corpus.data/provenance-baseline.jsonis the debt register for known-dead links (whyrequired per entry); placeholders can never be baselined — there is a test asserting that, because a gate that can be switched off is not a gate.Verification
Two bugs the tests caught while writing it, both now covered:
--strict-new(nargs=*) swallowed the positional file list, so CI would have silently scanned the whole corpus instead of the PR's files; and a case-insensitiveREPOplaceholder pattern flagged any URL containing/repo/.Boundaries (stated, not implied)
Rules for contributors, including the three legitimate ways to satisfy it:
docs/maintainer/provenance-gate-2026-09-16.md.PR Type
enhancement, tests, documentation
Description
New provenance gate verifies lesson citations resolve
Two-tier rules: placeholder/dead links fail, high evidence without source warns
Offline pytest suite covers 27 gate contract cases
CI workflow, baseline register, and maintainer docs included
Diagram Walkthrough
flowchart TD A["Lesson frontmatter"] --> B["Extract URLs<br/>(source / evidence_refs / url)"] B --> C{"Classify<br/>ok / dead / unknown<br/>exempt / placeholder"} C -->|placeholder or dead| D{"In baseline?"} D -->|no| E["Tier 1: FAIL"] D -->|yes| F["Tier 1: pass"] C -->|unknown| G["Tier 1: pass<br/>(never fail on absence)"] C -->|ok or exempt| F C -->|new file + E2/E3 + none| H["Tier 2: advisory"] F --> I["CI: green"] E --> J["CI: red"] H --> IFile Walkthrough
check_provenance.py
New provenance gate checking lesson citationsscripts/check_provenance.py
urllibonly) that scans every lessonfrontmatter for
source/evidence_refs/provenance/url/linkURLsand classifies each as
ok/dead/unknown/exempt/placeholder/none.split (from Lesson debt: legacy evidence_level (256) + gate new-vs-modified policy + en/ mirror dup FPs #1506): tier 1 always fails on placeholder URLs or
confirmed-dead (404/410) links not in the baseline; tier 2 emits
advisories when a newly-added file claims
evidence_level: E2/E3withno resolvable source.
github.comURLs onto the REST API for low-cost checks, readsGITHUB_TOKEN/GH_TOKEN, and treats timeouts / DNS / TLS / 403 / 429 /5xx as
unknownso the gate only fails on evidence, never on absence.localhost,example.com, RFC1918ranges,
*.local/*.internal/*.lan) by construction, and provides--check/--list/--offline/--strict-new/--update-baselinefor CIand maintainer use; baseline I/O goes through
data/provenance-baseline.json.test_check_provenance.py
Offline tests for the provenance gatetests/test_check_provenance.py
with no citations produces no failure, placeholder sources fail
without network, dead sources fail, and baselined dead links are
tolerated.
unknownfor403/429/5xx/timeouts/DNS/offline,
exemptfor private-IP /localhost/example.com/*.local) and replays the exact feat(lesson): alembic upgrade failure diagnosis #1713 red-team URL as aregression test.
evidence_levelwithout a source isadvisory on new files only, silent on existing ones, and cleared by a
resolvable URL), frontmatter parsing edge cases (list-form
evidence_refs, GitHub→API mapping, missing frontmatter, malformedbaseline), and CLI behaviour for
--strict-newnargs="*"swallowingpositional arguments.
provenance-gate-2026-09-16.md
Maintainer docs for the provenance gatedocs/maintainer/provenance-gate-2026-09-16.md
where feat(lesson): alembic upgrade failure diagnosis #1713–feat(lesson): write success but data missing — silent failures #1716 passed 24/24 checks against a 404 GitHub repo, and
shows why the existing three gates cannot catch invented sources.
dead links outside the baseline, tier 2 advisory for new files
claiming E2/E3 with no source) and the status semantics, including why
unknownmust never fail."placeholders can never be baselined" invariant, a contributor
self-help recipe (
--list --check, three acceptable fixes), and thegate's stated limits (semantic truth, private sources, plain-text
claims) plus next-step items already on the roadmap.
provenance-gate.yml
Add provenance gate CI workflow.github/workflows/provenance-gate.yml
scripts/check_provenance.pyagainst changed lesson files on PRs
(newly added lessons with
evidence_level: E2/E3and no resolvablesource) is strict-new
manual
workflow_dispatchrun against the full corpus as a report(continue-on-error)
rules doc, after first running the gate's unit tests
provenance-baseline.json
Add empty provenance baseline debt registerdata/provenance-baseline.json
misakanet-provenance-baseline/1known_deadandexempt_urlsarrays (corpus currently passeswithout any entries)
sources, not a permission slip