Skip to content

Commit 5595092

Browse files
Make the embed endpoint optional so memory runs lexical-only (CL-6287)
loadMemoryConfig() hard-required EMBED_BASE_URL and EMBED_MODEL, so a host with a pgvector Postgres and no embeddings account got no memory plane at all — despite the engine already degrading to Postgres full-text search at runtime. That was a config gate standing in front of a capability that already worked. The embed block is now optional. With no embed config the plane constructs and serves add plus lexical search, skipping the dense channel entirely rather than attempting and failing it, and reports degraded: ["dense_unavailable", "lexical_only"] so the state is observable rather than silent. Capture stores chunks without vectors and never reaches the embed-model registry. Transform/replay still requires an embed endpoint and fails loudly and early without one. Memory.capabilities.embeddingsConfigured lets a host discover the tier at construction instead of inferring it from a search, and add's degraded field is now a reason array matching search's shape rather than a bare boolean. Migration note for consumers: dense_unavailable now also fires permanently on a deliberately lexical-only host. An alert keyed on it alone should check lexical_only too. Known gaps, filed: CL-6292 (documents captured while unembedded are never backfilled), CL-6293 (lexical_only escalates to log.error forever on an intentional config).
1 parent 7449350 commit 5595092

22 files changed

Lines changed: 824 additions & 86 deletions

‎.env.example‎

Lines changed: 7 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -6,9 +6,13 @@
66
DATABASE_URL=postgres://memory:memory-dev-password@localhost:5434/memory
77
DB_POOL_MAX=8
88

9-
# Embeddings (compose.yml → Ollama on :11434). The engine never embeds
10-
# in-process; it always calls this endpoint. Swap four vars for a hosted
11-
# provider (e.g. EMBED_BASE_URL=https://api.openai.com, EMBED_API_STYLE=openai).
9+
# Embeddings (compose.yml → Ollama on :11434). Optional — leave both
10+
# EMBED_BASE_URL and EMBED_MODEL unset to run lexical-only (no dense
11+
# retrieval; `add` and lexical `search` still work,
12+
# degraded: ["dense_unavailable", "lexical_only"]). When set, both are
13+
# required together. The engine never embeds in-process; it always calls
14+
# this endpoint. Swap four vars for a hosted provider (e.g.
15+
# EMBED_BASE_URL=https://api.openai.com, EMBED_API_STYLE=openai).
1216
EMBED_BASE_URL=http://localhost:11434
1317
EMBED_MODEL=nomic-embed-text
1418
EMBED_API_STYLE=ollama

‎CHANGELOG.md‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -25,6 +25,27 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
2525

2626
### Fixed
2727

28+
- `loadMemoryConfig` / `EngineConfig.embed` no longer requires an embed
29+
endpoint: a host with a pgvector Postgres and no embed endpoint now
30+
constructs and serves `add` + lexical `search`. Dense retrieval is skipped
31+
(not attempted-and-failed) and `search` reports
32+
`degraded: ["dense_unavailable", "lexical_only"]` so the state stays
33+
observable (CL-6287). **Migration note:** `dense_unavailable` is now also
34+
emitted on every search for a deliberately-unconfigured engine, not only on
35+
a transient dense-retrieval failure — a host with an existing alert rule
36+
keyed on `dense_unavailable` alone should also check for `lexical_only` in
37+
the same `degraded` array to distinguish "opted into lexical-only" from an
38+
actual regression.
39+
- `add`'s `degraded` is now a reason array (`["embed_unavailable"]` and/or
40+
`["embed_unavailable", "lexical_only"]`), matching `search`'s shape —
41+
previously a bare boolean, which made it impossible to write one
42+
"is this response degraded" check across both verbs (CL-6287). **Breaking
43+
if a host coded against the boolean:** `degraded: true` is now
44+
`degraded: [...]`; check array presence/length instead of truthiness (both
45+
are still falsy/omitted when the document captured cleanly).
46+
- `Memory.capabilities.embeddingsConfigured` (and the underlying
47+
`DocumentStore.capabilities`) let a host learn recall is lexical-only at
48+
construction time, without issuing a search first (CL-6287).
2849
- Feed `nextCursor` advances past the examined raw page after grant-tag
2950
post-filter (a fully denied page no longer stalls the consumer forever).
3051
- Distiller default system prompt uses the configured `agentId` for

‎IMPLEMENTATION.md‎

Lines changed: 39 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -88,14 +88,43 @@ fallback)` parses a positive integer or throws.
8888
| `DATABASE_URL` | **yes** | — | the engine's own pgvector Postgres |
8989
| `DB_POOL_MAX` | no | `8` | postgres-js pool size |
9090
| `FTS_LANGUAGE` | no | `english` | text search config for the lexical channel; fixed into the generated column at migration time — changing it later requires rebuilding the column (recipe below), and `runMemoryMigrations` fails loudly if config and column disagree. Unqualified `pg_catalog` config names only — a schema-qualified config (`myschema.mycfg`) is rejected explicitly, both when configuring and when read back from an already-migrated column. |
91-
| `EMBED_BASE_URL` | **yes** | — | embed endpoint root, no path suffix |
92-
| `EMBED_MODEL` | **yes** | — | model id/name passed to the embed endpoint |
91+
| `EMBED_BASE_URL` | no | — | embed endpoint root, no path suffix; absent (with `EMBED_MODEL` also absent) => lexical-only, see below |
92+
| `EMBED_MODEL` | no | — | model id/name passed to the embed endpoint; must be set together with `EMBED_BASE_URL` (both or neither — one without the other throws) |
9393
| `EMBED_API_STYLE` | no | `"openai"` | `"openai" \| "tei" \| "ollama"` |
9494
| `EMBED_API_KEY` | no | `undefined` | forwarded as `Authorization: Bearer <key>` |
9595
| `RERANK_BASE_URL` | no | `undefined` | absent => search degrades to fusion-only |
9696
| `RERANK_MODEL` | no | `undefined` | defaults to `bge-reranker-v2-m3` in the client |
9797
| `RERANK_API_KEY` | no | `undefined` | forwarded as Bearer token to the rerank endpoint |
9898

99+
**Lexical-only mode (CL-6287).** `EngineConfig.embed` is optional — leave both
100+
`EMBED_BASE_URL`/`EMBED_MODEL` unset and the engine still constructs and
101+
serves `add` + lexical `search` against a pgvector Postgres with no
102+
embed endpoint configured. Dense retrieval is skipped rather than
103+
attempted (no doomed HTTP call on every query), `add` still captures
104+
documents (chunks stored, no vectors), and both verbs report a `degraded`
105+
reason array — never a bare boolean, so a host can write one "is this
106+
response degraded" check across both: `search` reports
107+
`degraded: ["dense_unavailable", "lexical_only"]`; `add` reports
108+
`degraded: ["embed_unavailable", "lexical_only"]` (or `["embed_unavailable"]`
109+
alone when the endpoint IS configured but a specific embed pass failed — a
110+
client error, timeout, or rejected chunk). The embed-model registry
111+
(`ensureEmbedModel`/`activateEmbedModel`) is never reached in this mode.
112+
113+
**Discoverability.** A host does not have to run a search to learn recall is
114+
limited: `memory.capabilities.embeddingsConfigured` (on the `Memory` handle
115+
`createMemory` returns) is `false` for a lexical-only engine, `true`
116+
otherwise — known at construction, no query needed. A custom `documentStore`
117+
that doesn't report its own `capabilities` defaults to `true` (this SDK
118+
cannot introspect a vendor store it doesn't own); see
119+
`DocumentStoreCapabilities` (ports/types.ts) for how a vendor store opts in.
120+
121+
The replay/backfill pipeline (`runTransform`, `promoteGeneration` in
122+
`services/transform.ts`) still requires an embed endpoint — re-deriving a
123+
corpus is inherently a re-embedding operation — and fails loudly if run
124+
against an engine with none configured; re-embedding documents captured while
125+
lexical-only, once an endpoint is later added, is an open follow-up (not
126+
implemented).
127+
99128
The engine's `EngineConfig.rerank` carries
100129
no `apiStyle` field of its own; `search.ts`'s `toRerankClientConfig` hardcodes
101130
`apiStyle: "tei"` when building the client config, i.e. the engine currently
@@ -688,9 +717,9 @@ bun run test # unit suite (no external
688717

689718
`compose.yml` provisions the pgvector Postgres (`memory` db, host port
690719
`5434`), an Ollama embeddings server (`:11434`), and a TEI reranker (`:8085`).
691-
The engine **never embeds internally** — `EMBED_BASE_URL` must point at a real
692-
endpoint. A model endpoint is just a URL + capability options, trusted the same
693-
as `DATABASE_URL`:
720+
The engine **never embeds internally** — when `EMBED_BASE_URL` is set, it must
721+
point at a real endpoint. A model endpoint is just a URL + capability options,
722+
trusted the same as `DATABASE_URL`:
694723

695724
- **Local default**: Ollama at `http://localhost:11434`
696725
(`EMBED_API_STYLE=ollama`, `EMBED_MODEL=nomic-embed-text`).
@@ -701,6 +730,11 @@ as `DATABASE_URL`:
701730
`RERANK_BASE_URL` is optional — unset runs lexical+dense+MMR without the
702731
cross-encoder (`degraded: ["rerank_unavailable"]`, still ranked/citable hits).
703732

733+
`EMBED_BASE_URL`/`EMBED_MODEL` are optional too — unset both to run
734+
lexical-only (`degraded: ["dense_unavailable", "lexical_only"]`, no dense
735+
channel, `add` still captures documents unvectorized). See the lexical-only
736+
note above.
737+
704738
## Testing
705739

706740
`bun test ./src` (`bun run test`), coverage via `bun run test:coverage`. Every

‎src/config.ts‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,16 @@ export type EngineConfig = {
2121
// trusted the same as DATABASE_URL — including a self-hosted endpoint on
2222
// localhost or a private IP. Self-hosted or managed makes no difference:
2323
// there is no self-host flag anywhere in the engine.
24-
embed: {
24+
//
25+
// Absent entirely => no embed endpoint configured. The engine still
26+
// constructs and serves `add` + lexical `search` in that case: dense
27+
// retrieval is skipped (never attempted, so it never runs a doomed HTTP
28+
// call), capture stores chunks without vectors, and search reports
29+
// `["dense_unavailable", "lexical_only"]` so the state is observable
30+
// rather than silent (see hybridSearch in services/search.ts). A pgvector
31+
// Postgres with no embed endpoint is a legitimate, fully-capable
32+
// lexical-only deployment, not a misconfiguration.
33+
embed?: {
2534
baseUrl: string;
2635
model: string;
2736
apiStyle: string;

‎src/core/degrade-metrics.ts‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -25,13 +25,23 @@ const DEGRADE_FLAG_SET = {
2525
live_timeout: true,
2626
live_error: true,
2727
memory_unavailable: true,
28+
lexical_only: true,
2829
} satisfies Record<DegradeFlag, true>;
2930

3031

3132
// Deriving the list from a `satisfies Record<DegradeFlag, true>` object
3233
// means adding a flag in hybrid-search.ts without adding it here is a
3334
// compile error (missing property), not a silent gap that only a test
3435
// iterating this same constant could ever have caught.
36+
//
37+
// `lexical_only` sits at a permanent ~100% windowed rate for a host that has
38+
// deliberately opted out of dense retrieval (no embed endpoint configured) —
39+
// unlike every other flag here, that is the intended, steady state rather
40+
// than a regression, so it escalates to log.error and stays there for as
41+
// long as the host runs lexical-only. That is accurate (the health snapshot
42+
// should show "running degraded"), not a bug; a host that finds the
43+
// permanent log.error noisy can raise its own highWatermark via
44+
// `configureDegradeMetrics`.
3545
export const ALL_DEGRADE_FLAGS: readonly DegradeFlag[] = Object.keys(
3646
DEGRADE_FLAG_SET,
3747
) as DegradeFlag[];

‎src/core/embed-worker.ts‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,19 @@ export interface EmbeddableChunk {
1010
text: string;
1111
}
1212

13+
// Capture's counterpart to search's `DegradeFlag` (hybrid-search.ts) — an
14+
// array, never a bare boolean, so a host can write one "is this response
15+
// degraded" check across `add` and `search`. Lives here (not services/
16+
// capture.ts) so `ports/types.ts` can reference it for
17+
// `DocumentStoreAddResult` the same way it already references `DegradeFlag`
18+
// for `DocumentStoreSearchResult`. `embed_unavailable` is the embed pass's
19+
// counterpart to `dense_unavailable` — it ran and failed (client error,
20+
// timeout, or a rejected/dims-mismatched chunk); `lexical_only` is paired
21+
// with it specifically when there's no embed endpoint configured at all,
22+
// mirroring search's `dense_unavailable`/`lexical_only` pairing for the same
23+
// "configured off" state.
24+
export type CaptureDegradedReason = "embed_unavailable" | "lexical_only";
25+
1326
export interface EmbedChunksResult {
1427
embedded: number;
1528
rejected: Array<{ chunkId: string; reason: string }>;

‎src/core/engine-client-config.ts‎

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -17,10 +17,14 @@ const VALID_EMBED_API_STYLES = new Set(["openai", "tei", "ollama"]);
1717
// the trust boundary between config and the client — an invalid value is an
1818
// operator misconfiguration and must fail loudly, not silently degrade.
1919
// Built from the engine's own operator-configured embed endpoint — a trusted
20-
// URL, the same as DATABASE_URL.
20+
// URL, the same as DATABASE_URL. Absent `embed` => no embed endpoint
21+
// configured => `undefined`, the same degrade-soft precedent already used by
22+
// `toRerankClientConfig` below — a caller must skip dense retrieval / the
23+
// embed pass entirely rather than dispatch a client with no endpoint.
2124
export function toEmbedClientConfig(
2225
embed: EngineConfig["embed"],
23-
): EmbedClientConfig {
26+
): EmbedClientConfig | undefined {
27+
if (!embed) return undefined;
2428
if (!VALID_EMBED_API_STYLES.has(embed.apiStyle)) {
2529
throw new Error(
2630
`Invalid EMBED_API_STYLE "${embed.apiStyle}" — must be one of: ${[...VALID_EMBED_API_STYLES].join(", ")}`,

‎src/core/hybrid-search.ts‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,16 @@ export type DegradeFlag =
2020
| "rerank_query_too_long"
2121
| "live_timeout"
2222
| "live_error"
23-
| "memory_unavailable";
23+
| "memory_unavailable"
24+
// The engine has no embed endpoint configured at all (EngineConfig.embed
25+
// is absent) — a deliberate, structural lexical-only deployment, distinct
26+
// from `dense_unavailable`'s per-call "dense contributed nothing this
27+
// time" (which also covers a configured endpoint that's merely down, or a
28+
// tenant with no active embed model yet). Always paired with
29+
// `dense_unavailable` on the search response (see hybridSearch,
30+
// services/search.ts) so an aggregate degrade-rate consumer still sees
31+
// "dense didn't contribute" even if it only understands that one flag.
32+
| "lexical_only";
2433

2534
export interface RankedCandidate {
2635
/** Stable identifier the candidate is keyed by across channels (a chunk id). */

‎src/index.ts‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -38,6 +38,7 @@ export type {
3838
HybridSearchResult,
3939
MemoryAddParams,
4040
MemoryAddResult,
41+
MemoryCapabilities,
4142
MemorySearchParams,
4243
MemoryIdentity,
4344
Memory,
@@ -77,6 +78,7 @@ export type {
7778
DocumentStore,
7879
DocumentStoreAddParams,
7980
DocumentStoreAddResult,
81+
DocumentStoreCapabilities,
8082
DocumentStoreSearchItem,
8183
DocumentStoreSearchParams,
8284
DocumentStoreSearchResult,

‎src/memory.test.ts‎

Lines changed: 51 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -173,6 +173,57 @@ describe("createMemory — construction validation", () => {
173173
});
174174
});
175175

176+
// CL-6287 review: a consumer (settings page, health check) must be able to
177+
// learn recall is lexical-only WITHOUT issuing a search first.
178+
describe("createMemory — capabilities.embeddingsConfigured (CL-6287)", () => {
179+
it("reports true when EngineConfig.embed is configured", async () => {
180+
const plane = createMemory({
181+
config: baseConfig({
182+
baseUrl: undefined,
183+
model: undefined,
184+
apiKey: undefined,
185+
maxDocChars: undefined,
186+
timeoutMs: undefined,
187+
}),
188+
});
189+
expect(plane.capabilities.embeddingsConfigured).toBe(true);
190+
await plane.close();
191+
});
192+
193+
it("reports false when EngineConfig.embed is absent (lexical-only)", async () => {
194+
const config: MemoryConfig = {
195+
memory: {
196+
databaseUrl: "postgres://localhost:5432/nonexistent-test-db",
197+
dbPoolMax: 1,
198+
ftsLanguage: "english",
199+
rerank: {
200+
baseUrl: undefined,
201+
model: undefined,
202+
apiKey: undefined,
203+
maxDocChars: undefined,
204+
timeoutMs: undefined,
205+
},
206+
},
207+
};
208+
const plane = createMemory({ config });
209+
expect(plane.capabilities.embeddingsConfigured).toBe(false);
210+
await plane.close();
211+
});
212+
213+
it("defaults to true for a custom DocumentStore that doesn't report its own capabilities", async () => {
214+
const plane = createMemory({
215+
documentStore: {
216+
add: async () => ({ documentId: "d1", versionId: "v1" }),
217+
search: async () => ({ items: [] }),
218+
list: async () => [],
219+
close: async () => {},
220+
},
221+
});
222+
expect(plane.capabilities.embeddingsConfigured).toBe(true);
223+
await plane.close();
224+
});
225+
});
226+
176227
describe("createMemory.find — grant-tag post-filter wiring", () => {
177228
const hybridSearch = mock((): Promise<HybridSearchResult> =>
178229
Promise.resolve({

0 commit comments

Comments
 (0)