Skip to content

jlens: resolve swap tokens exactly and suggest the closest token - #249

Merged
hijohnnylin merged 1 commit into
mainfrom
jlens-swap-closest-token
Sep 23, 2026
Merged

hijohnnylin merged 1 commit into
mainfrom
jlens-swap-closest-token

Conversation

@hijohnnylin

Copy link
Copy Markdown
Owner

Problem

When a jlens swap token was not in the vocabulary, the inference server fell back to the first token with the same trimmed text. It did not tell the user. So ants could swap in ants, ants\n or \tants, whichever had the lowest id. When no token matched, the user got a bare error with no hint. The fallback also scanned the full vocabulary (~150k entries) on every miss.

Change

Inference

  • _resolve_steer_token_id takes exact matches only.
  • On a miss, /v1/lens/prompt returns a 400 LensErrorResponse with {error, token, suggestedToken}.
  • The suggestion is the longest vocab entry that is a prefix of the input. It tries the input as typed and with one leading space, and keeps a leading space that the user typed. The search is bounded: 3-54 碌s on the Qwen3 tokenizer, including a 4,000-character input.
  • openapi.json and lib/api/inference.d.ts are regenerated.

Webapp

  • /api/lens/prompt forwards token / suggestedToken explicitly, and its swagger block documents them.
  • jlens-stream.ts throws LensUnknownTokenError for that response.
  • The steer panel shows "Not a single token. Closest: 鈵ntid [Apply]" below the swap input. Apply fills in the token and does not run the swap.

Compatibility

  • New webapp with the old inference server: no change for users. Without the new fields, the plain error shows in the banner, as before.
  • New inference server with the old webapp: the silent trim fallback stops, and a plain error shows with the suggestion in its text.
  • Recommended order: deploy the webapp first.
  • Public API: callers of /api/lens/prompt that relied on the trim fallback now get a 400. The 400 includes the suggestion.

Testing

  • New apps/inference/tests/unit/test_lens_steer_token_resolve.py.
  • Inference unit suite passes (643 tests, with HF_TOKEN), plus ruff, pyright and make openapi-check.
  • Webapp npm run lint (eslint + tsc), format:check and npm test pass (154 tests).
  • Checked suggestions and timing against the real Qwen3-8B tokenizer.
  • Not tested end to end against a live inference server or in a browser.

A steer or swap token that was not in the vocabulary fell back to the first
token with the same trimmed text, with no notice. So "ants" could swap in
" ants", "ants\n" or "\tants", whichever had the lowest id, and the user did
not know. The fallback also scanned the full vocabulary on every miss.

Inference now takes exact matches only. On a miss, /v1/lens/prompt returns a
400 LensErrorResponse with the token and the closest token: the longest vocab
entry that is a prefix of the input, as typed or with one leading space. The
search is bounded, at 3-54 us on the Qwen3 tokenizer.

The webapp forwards token/suggestedToken from /api/lens/prompt, and the steer
panel shows the closest token next to the swap input with an Apply button.
The new webapp works with the old inference server: without the new fields
it shows the plain error, as before.

Co-authored-by: Cursor <cursoragent@cursor.com>
@hijohnnylin
hijohnnylin merged commit f013777 into main Sep 23, 2026
19 checks passed
@hijohnnylin
hijohnnylin deleted the jlens-swap-closest-token branch September 23, 2026 00:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant