Repository navigation
webapp: add Jev (TypeSafe) explanation scorers - #246
Merged
Merged
Conversation
Three new ExplanationScoreType names, all served by one module: - jev_detection: one Noul per example on plain text, value = balanced accuracy - jev_fuzz: same, with firing tokens marked << >>, value = balanced accuracy - jev_score: one 5-level Score rating over the top examples, value = level / 4 One TypeSafe request per explanation replaces the per-example chat completions the OpenRouter scorers make; ~2k input tokens, ~0.2s. The server key (TYPESAFE_API_KEY) is used for every user, since the route is already rate limited at 120/hour and a score costs about $0.0001. Top-logit fit is asked as a separate question and stored in jsonDetails only: many features have byte-pair fragments as top logits, and folding it into the value penalizes correct explanations. Moves the byte-level token decoder out of chinese-translations.ts into a shared module, since stored activation tokens still carry Ċ and âĢĻ. Needs ExplanationScoreType rows jev_fuzz / jev_detection / jev_score and an ExplanationScoreModel row jev-latest; no schema change. Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Three new explanation score types backed by TypeSafe's Jev, served by one module,
lib/external/autointerp-scorer-jev.ts:valuejev_detectionrecall_altjev_fuzz<< >>; non-activating texts get one marked word so the marker is not the signaljev_scorescore / 4Jev evaluates every question in a request in parallel and returns calibrated probabilities instead of text, so one explanation is one request (~2k input tokens, ~0.2s, ~$0.0001) with no JSON parsing. Each score also asks a separate question about whether the feature's top positive logits fit the explanation and stores it in
jsonDetails.logit_fit. It never entersvalue: many features have byte-pair fragments as top logits, and folding it in penalizes correct explanations.Why balanced accuracy
Measured over 8 gpt2-small res-jb features against real explanations, an explanation borrowed from another feature, "terms related to cats", and "common English words":
eleuther_fuzz/eleuther_recall, so the numbers are comparable.jev_scorehas the widest range: real ~73, wrong ~7, cats ~1; over-broad ~47 is its weak spot, and Jev reports low confidence there, which is stored.Plumbing
TYPESAFE_API_KEYinlib/env.ts), open to all users./api/explanation/scoreis already rate limited at 120/hour per caller.jev_*skips theExplanationScoreModellookup and the OpenRouter key; the model is fixed atjev-latestand the UI hides the model dropdown, likeeleuther_embedding.jsonDetailsstores the versionedmodelthe API reports, token usage, and per-example nouls so the metric can be recomputed later without re-querying.fetchwith a 30s timeout and backoff on 429/529; failures surface asupstreamError('typesafe').chinese-translations.ts('use client') intolib/utils/byte-level-tokens.ts, plusdecodeMixedTokenfor the half-decoded tokens in stored activations (Ċ,âĢĻ). Behaviour of the translations lookup is unchanged.jev_score.Database
No schema change. Needs
ExplanationScoreTyperowsjev_fuzz,jev_detection,jev_scoreand anExplanationScoreModelrowjev-latest(featured = false,openRouterModelId = null). Already inserted in production.Testing
markedText,balancedAccuracy,plainText,decodeMixedToken); full suite 153/153; eslint,tsc --noEmit, prettier clean.Made with Cursor