Skip to content

bug: BuildQualityFromMappings hardcodes AssertionDetectionConfidence to 0 for external analyzers #1

Description

@gxmiranda

Problem

BuildQualityFromMappings in internal/adapter/quality.go hardcodes AssertionDetectionConfidence: 0 for all external analyzer quality reports (line 138):

AssertionDetectionConfidence: 0, // External analyzers don't provide aggregate detection confidence

This causes gaze quality --analyzer output to always show:

Assertion detection confidence: 0%

even when the external analyzer (e.g., snake-eyes) is correctly detecting and classifying assertions across all test functions.

Evidence

Running gaze quality --analyzer snake-eyes --language python . against the snake-eyes repo:

  • 326 tests analyzed
  • Assertions ARE detected — equality, error_check, generic assertion types are present in every mapping row
  • Per-test AssertionCount is correctly populated (e.g., len(testMappings))
  • But AssertionDetectionConfidence is always 0%

Root Cause

The Go-native quality pipeline computes detection confidence internally because it has access to the full assertion detection pipeline. The external analyzer path added in PR unbound-force#242 has no mechanism to receive or compute this metric — it was deferred with a comment.

Proposed Fix

Two options:

  1. Gaze-side computation (no protocol change): compute detection confidence from the data already available. For each test function, if the external analyzer returned at least one mapping row, that test has detected assertions. The metric becomes len(tests_with_assertions) / len(all_tests) * 100. This requires knowing the total test count, which could come from the discover method or be inferred from unique test_function values in the mapping data.

  2. Protocol extension: add an optional assertion_detection_confidence (int, 0-100) field to the test_mapping response. The external analyzer computes this itself since it knows which tests it scanned but found no assertions in (true negatives vs detection failures). This is more accurate but requires a protocol version bump.

Option 1 is simpler and can ship without protocol changes. Option 2 is more accurate for analyzers that want to report detection quality.

Impact

MEDIUM — the metric is misleading (shows 0% when detection is working correctly), but does not affect contract coverage scores, gap identification, or over-specification counts. Cosmetic/informational only.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions