Skip to content

webapp: add Jev (TypeSafe) explanation scorers - #246

Merged
hijohnnylin merged 1 commit into
mainfrom
jev-explanation-scorer
Sep 18, 2026
Merged

hijohnnylin merged 1 commit into
mainfrom
jev-explanation-scorer

Conversation

@hijohnnylin

Copy link
Copy Markdown
Owner

What

Three new explanation score types backed by TypeSafe's Jev, served by one module, lib/external/autointerp-scorer-jev.ts:

Type How value
jev_detection One Noul (yes/no probability) per example on the plain text: top 20 activations, up to 5 zero activations, and the 5 decoys shared with recall_alt Balanced accuracy at 0.5
jev_fuzz Same, with the tokens the feature fires on (>= 3/10 of max) wrapped in << >>; non-activating texts get one marked word so the marker is not the signal Balanced accuracy at 0.5
jev_score One 5-level Score question rating the explanation against the top 12 marked examples score / 4

Jev evaluates every question in a request in parallel and returns calibrated probabilities instead of text, so one explanation is one request (~2k input tokens, ~0.2s, ~$0.0001) with no JSON parsing. Each score also asks a separate question about whether the feature's top positive logits fit the explanation and stores it in jsonDetails.logit_fit. It never enters value: many features have byte-pair fragments as top logits, and folding it in penalizes correct explanations.

Why balanced accuracy

Measured over 8 gpt2-small res-jb features against real explanations, an explanation borrowed from another feature, "terms related to cats", and "common English words":

  • Balanced accuracy: real ~84, wrong ~50, over-broad ~55. Same formula as eleuther_fuzz / eleuther_recall, so the numbers are comparable.
  • AUC rewards "common English words" (78-88), because the marked tokens are common English words.
  • The holistic jev_score has the widest range: real ~73, wrong ~7, cats ~1; over-broad ~47 is its weak spot, and Jev reports low confidence there, which is stored.

Plumbing

  • Server key only (TYPESAFE_API_KEY in lib/env.ts), open to all users. /api/explanation/score is already rate limited at 120/hour per caller.
  • jev_* skips the ExplanationScoreModel lookup and the OpenRouter key; the model is fixed at jev-latest and the UI hides the model dropdown, like eleuther_embedding.
  • jsonDetails stores the versioned model the API reports, token usage, and per-example nouls so the metric can be recomputed later without re-querying.
  • Transport is plain fetch with a 30s timeout and backoff on 429/529; failures surface as upstreamError('typesafe').
  • The GPT-2 byte-level token decoder moves from chinese-translations.ts ('use client') into lib/utils/byte-level-tokens.ts, plus decodeMixedToken for the half-decoded tokens in stored activations (Ċ, âĢĻ). Behaviour of the translations lookup is unchanged.
  • Detail dialog gets a per-example table for fuzz/detection and a level-probability view with confidence for jev_score.

Database

No schema change. Needs ExplanationScoreType rows jev_fuzz, jev_detection, jev_score and an ExplanationScoreModel row jev-latest (featured = false, openRouterModelId = null). Already inserted in production.

Testing

  • 9 unit tests on the pure helpers (markedText, balancedAccuracy, plainText, decodeMixedToken); full suite 153/153; eslint, tsc --noEmit, prettier clean.
  • Live run of the module against the API with prisma mocked: real vs "cats" gave detection 0.65 / 0.50, fuzz 0.975 / 0.50, score 0.615 / 0.007.
  • Scored all three types end to end from the UI against production rows; detail dialog renders.

Made with Cursor

Three new ExplanationScoreType names, all served by one module:
- jev_detection: one Noul per example on plain text, value = balanced accuracy
- jev_fuzz: same, with firing tokens marked << >>, value = balanced accuracy
- jev_score: one 5-level Score rating over the top examples, value = level / 4

One TypeSafe request per explanation replaces the per-example chat
completions the OpenRouter scorers make; ~2k input tokens, ~0.2s. The
server key (TYPESAFE_API_KEY) is used for every user, since the route is
already rate limited at 120/hour and a score costs about $0.0001.

Top-logit fit is asked as a separate question and stored in jsonDetails
only: many features have byte-pair fragments as top logits, and folding
it into the value penalizes correct explanations.

Moves the byte-level token decoder out of chinese-translations.ts into a
shared module, since stored activation tokens still carry Ċ and âĢĻ.

Needs ExplanationScoreType rows jev_fuzz / jev_detection / jev_score and
an ExplanationScoreModel row jev-latest; no schema change.

Co-authored-by: Cursor <cursoragent@cursor.com>
@hijohnnylin
hijohnnylin merged commit 6c68e3b into main Sep 18, 2026
6 checks passed
@hijohnnylin
hijohnnylin deleted the jev-explanation-scorer branch September 18, 2026 23:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant