Skip to content

Commit f8256c7

Browse files
committed
knowledge-engine: add a grant-checked ask() that answers from retrieved context
Rebased and slimmed onto current main (post #10 rerank-budget validator, post #11 log-detail interpolation + degrade metrics). Dropped 13 files of pure Prettier reformatting that had accumulated on the old branch and were never part of this feature; nothing else changed. ask() performs its own knowledge:search grant check internally, since an in-process caller never passes through the HTTP requireGrant route guard. It then searches as the asking principal (same per-document ACL/block-list path as search()), assembles a bounded grounded-context block, and calls a host-supplied generate() function. The engine intentionally owns no generation client, matching its existing posture on embedding. Claude-Session: https://claude.ai/code/session_017GTgGzn5xAwvkU2GAPAHpF
1 parent b705853 commit f8256c7

5 files changed

Lines changed: 447 additions & 5 deletions

File tree

‎README.md‎

Lines changed: 65 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -58,6 +58,71 @@ import { runKnowledgeMigrations } from "@corbits/knowledge-engine/migrations";
5858
await runKnowledgeMigrations(process.env.KNOWLEDGE_DATABASE_URL);
5959
```
6060

61+
## ask()
62+
63+
`knowledge.ask()` answers a question from retrieved context, in-process — no
64+
HTTP hop, no separate host-side answer synthesis. It is grant-checked
65+
internally, so it is safe to call from anywhere a host has resolved a
66+
principal, even code paths that never go through the mounted HTTP routes:
67+
68+
```ts
69+
const { text, citations, evidence } = await knowledge.ask({
70+
tenantId,
71+
principalId,
72+
query,
73+
k: 6, // optional, defaults to hybridSearch's default
74+
});
75+
```
76+
77+
1. Checks the capability grant (`authorize(grantStore, principalId, tenantId,
78+
"knowledge", "search", conditionRegistry)`) and throws
79+
`KnowledgeNotPermittedError` unless the effect is an explicit `"allow"` — no
80+
matching grant (`effect: null`) denies too. This runs whether or not the
81+
call came through the HTTP route guard, so it can never be forgotten.
82+
2. Searches as that principal (`hybridSearch`), so per-document visibility and
83+
block lists apply exactly as they do for `search()`.
84+
3. Assembles a grounded context block from the hits' `snippet` text, truncated
85+
to a character budget.
86+
4. Calls the **host-supplied** `generate` function with a system prompt that
87+
instructs answering only from context and refusing explicitly when the
88+
context doesn't contain the answer.
89+
5. Returns the answer text, the citations actually used (matched to the `[N]`
90+
markers in the prompt), and the search's evidence level.
91+
92+
### The engine owns no generation client
93+
94+
`ask()` takes generation as an injected function, not as config:
95+
96+
```ts
97+
type Generate = (messages: readonly ChatMessage[]) => Promise<string>;
98+
99+
mountKnowledgeEngine(app, {
100+
config,
101+
grants,
102+
generate: async (messages) => runInferenceSomehow(messages),
103+
});
104+
```
105+
106+
This is deliberate. Interchange already has an inference layer
107+
(`@intx/inference`) with provider adapters, tenant-scoped credentials, a retry
108+
policy, audit collection and authz gates. A `fetch` client here would bypass all
109+
of it and take an API key from a raw env var — so hosts wire `generate` to that
110+
layer instead, and credentials stay in the credential store where they belong.
111+
112+
It also keeps the engine transport-free, which is the posture it already takes
113+
on embedding: never in-process, always an endpoint the owner plugs in.
114+
115+
Omit `generate` if the host only captures and searches; `ask()` then fails with
116+
a 501 naming what is missing rather than at some later point.
117+
118+
Two things that belong in the host's `generate`, learned the hard way:
119+
120+
- **Timeouts must be generous for local models.** A cold 10GB model can take
121+
over a minute to page into memory before emitting a token.
122+
- **Use a non-reasoning model.** A reasoning model that exhausts its budget
123+
returns chain-of-thought with empty content, and the host will see an empty
124+
answer. Detect it there and say so.
125+
61126
## Local development
62127

63128
The SDK is not a server — it mounts onto your app. This repo ships a

‎src/index.ts‎

Lines changed: 27 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -11,7 +11,11 @@ import { createRequireGrant, type TenantEnv } from "@intx/hub-api";
1111

1212
import type { KnowledgeConfig } from "./mount-config.ts";
1313
import { CaptureLog } from "./capture-log.ts";
14-
import { createKnowledgePlane, type KnowledgePlane } from "./knowledge.ts";
14+
import {
15+
createKnowledgePlane,
16+
type Generate,
17+
type KnowledgePlane,
18+
} from "./knowledge.ts";
1519
import {
1620
mountKnowledgeRoutes,
1721
type GrantConfig,
@@ -26,7 +30,16 @@ export { loadKnowledgeConfig } from "./mount-config.ts";
2630
export type { EngineConfig } from "./config.ts";
2731
export { RerankConfigError } from "./core/rerank-client.ts";
2832
// Knowledge plane + capture log
29-
export type { KnowledgePlane } from "./knowledge.ts";
33+
export type {
34+
KnowledgePlane,
35+
KnowledgePlaneOptions,
36+
KnowledgeAskParams,
37+
AskResult,
38+
AskCitation,
39+
ChatMessage,
40+
Generate,
41+
} from "./knowledge.ts";
42+
export { KnowledgeNotPermittedError, KnowledgeError } from "./knowledge.ts";
3043
export { CaptureLog, type CaptureEvent } from "./capture-log.ts";
3144
// Migrations
3245
export { runKnowledgeMigrations } from "./migrations.ts";
@@ -52,6 +65,15 @@ export type MountKnowledgeEngineOptions = {
5265
* unguarded.
5366
*/
5467
grants: GrantConfig;
68+
/**
69+
* How `ask()` reaches a model. Omit if this host only captures and searches;
70+
* `ask()` then fails with a 501 naming what is missing.
71+
*
72+
* The engine owns no generation client on purpose — Interchange's
73+
* `@intx/inference` already has provider adapters, tenant-scoped credentials,
74+
* retry, audit and authz gates. Wire this to that rather than to a bare fetch.
75+
*/
76+
generate?: Generate;
5577
};
5678

5779
export type MountedKnowledgeEngine = {
@@ -78,7 +100,9 @@ export function mountKnowledgeEngine(
78100
const rerankConfig = toRerankClientConfig(options.config.knowledge.rerank);
79101
if (rerankConfig) validateRerankConfig(rerankConfig);
80102

81-
const knowledge = createKnowledgePlane(options.config);
103+
const knowledge = createKnowledgePlane(options.config, options.grants, {
104+
...(options.generate ? { generate: options.generate } : {}),
105+
});
82106
const captureLog = new CaptureLog();
83107
const deps: RouteDeps = {
84108
knowledge,

‎src/knowledge.test.ts‎

Lines changed: 158 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,158 @@
1+
import { describe, expect, it, mock } from "bun:test";
2+
import { createInMemoryGrantStore } from "@intx/authz";
3+
import type { GrantRule } from "@intx/authz";
4+
5+
import {
6+
createKnowledgePlane,
7+
KnowledgeNotPermittedError,
8+
synthesizeAnswer,
9+
} from "./knowledge.ts";
10+
import type { KnowledgeConfig } from "./mount-config.ts";
11+
import type { ChatMessage } from "./knowledge.ts";
12+
import type { SearchHit } from "./core/schemas/search.ts";
13+
14+
const PRINCIPAL = "p1";
15+
const TENANT = "t1";
16+
17+
function grant(action: string): GrantRule {
18+
return {
19+
id: `g-${action}`,
20+
resource: "knowledge",
21+
action,
22+
effect: "allow",
23+
origin: "role",
24+
conditions: null,
25+
expiresAt: null,
26+
roleId: null,
27+
principalId: PRINCIPAL,
28+
};
29+
}
30+
31+
function hit(overrides: Partial<SearchHit> = {}): SearchHit {
32+
return {
33+
chunk_id: "chunk-1",
34+
document_id: "doc-1",
35+
version: 1,
36+
version_id: "v1",
37+
status: "current",
38+
score: 0.9,
39+
title: "Doc One",
40+
snippet: "the relevant snippet",
41+
kind: "note",
42+
created_by_kind: "human",
43+
citation: {
44+
adapter: "mcp",
45+
external_ref: "ref-1",
46+
open: { type: "doc", id: "doc-1" },
47+
},
48+
entity_ids: [],
49+
channels_matched: ["lexical"],
50+
...overrides,
51+
} as SearchHit;
52+
}
53+
54+
const config: KnowledgeConfig = {
55+
knowledge: {
56+
databaseUrl: "postgres://localhost:5432/nonexistent-test-db",
57+
dbPoolMax: 1,
58+
embed: {
59+
baseUrl: "http://embed",
60+
model: "m",
61+
apiStyle: "openai",
62+
apiKey: undefined,
63+
},
64+
rerank: {
65+
baseUrl: undefined,
66+
model: undefined,
67+
apiKey: undefined,
68+
maxDocChars: undefined,
69+
},
70+
},
71+
};
72+
73+
describe("ask() — grant check", () => {
74+
it("denies with KnowledgeNotPermittedError when no grant matches (effect: null)", async () => {
75+
const grants = {
76+
grantStore: createInMemoryGrantStore([]),
77+
conditionRegistry: {},
78+
};
79+
const plane = createKnowledgePlane(config, grants);
80+
await expect(
81+
plane.ask({ tenantId: TENANT, principalId: PRINCIPAL, query: "q" }),
82+
).rejects.toBeInstanceOf(KnowledgeNotPermittedError);
83+
});
84+
85+
it("denies when the only matching grant is an explicit deny", async () => {
86+
const denyGrant: GrantRule = { ...grant("search"), effect: "deny" };
87+
const grants = {
88+
grantStore: createInMemoryGrantStore([denyGrant]),
89+
conditionRegistry: {},
90+
};
91+
const plane = createKnowledgePlane(config, grants);
92+
await expect(
93+
plane.ask({ tenantId: TENANT, principalId: PRINCIPAL, query: "q" }),
94+
).rejects.toBeInstanceOf(KnowledgeNotPermittedError);
95+
});
96+
});
97+
98+
describe("synthesizeAnswer", () => {
99+
const neverCalled = mock(() =>
100+
Promise.reject(new Error("generate must not be called")),
101+
);
102+
103+
it("refuses with no citations when there are no hits", async () => {
104+
const result = await synthesizeAnswer(
105+
"q",
106+
{ hits: [], evidence: "none" },
107+
neverCalled,
108+
);
109+
expect(result.citations).toEqual([]);
110+
expect(result.evidence).toBe("none");
111+
expect(neverCalled).not.toHaveBeenCalled();
112+
});
113+
114+
it("refuses with no citations when hits have no readable snippet text", async () => {
115+
const result = await synthesizeAnswer(
116+
"q",
117+
{ hits: [hit({ snippet: " " })], evidence: "weak" },
118+
neverCalled,
119+
);
120+
expect(result.text).toContain("couldn't read any text");
121+
expect(result.citations).toEqual([]);
122+
expect(neverCalled).not.toHaveBeenCalled();
123+
});
124+
125+
it("grounds the prompt in the retrieved snippets and returns their citations", async () => {
126+
const generate = mock((messages: readonly ChatMessage[]) => {
127+
// The grounding contract: the model sees the numbered context, and only
128+
// the numbered context, for the hits that fit the budget.
129+
expect(messages[1]?.content).toContain("[1] Doc One");
130+
expect(messages[1]?.content).toContain("the relevant snippet");
131+
return Promise.resolve("Grounded answer [1].");
132+
});
133+
134+
const result = await synthesizeAnswer(
135+
"what is the answer?",
136+
{ hits: [hit()], evidence: "strong" },
137+
generate,
138+
);
139+
140+
expect(result.text).toBe("Grounded answer [1].");
141+
expect(result.evidence).toBe("strong");
142+
expect(result.citations).toEqual([
143+
{
144+
index: 1,
145+
documentId: "doc-1",
146+
title: "Doc One",
147+
citation: hit().citation,
148+
},
149+
]);
150+
});
151+
152+
it("propagates a generate failure rather than inventing an answer", async () => {
153+
const generate = mock(() => Promise.reject(new Error("model unreachable")));
154+
await expect(
155+
synthesizeAnswer("q", { hits: [hit()], evidence: "weak" }, generate),
156+
).rejects.toThrow("model unreachable");
157+
});
158+
});

0 commit comments

Comments
 (0)