Skip to content

chore(evals): Update model evaluations 2026-09-01 - #229

Merged
mtodor merged 1 commit into
mainfrom
chore/update-model-evaluation-2026-09-01
Sep 1, 2026
Merged

chore(evals): Update model evaluations 2026-09-01#229
mtodor merged 1 commit into
mainfrom
chore/update-model-evaluation-2026-09-01

Conversation

@rhacs-bot

Copy link
Copy Markdown
Contributor

Automated weekly model evaluation update.

Models evaluated: gpt-5-mini
Date: 2026-09-01

This PR was automatically generated by the Model Evaluation workflow.

@rhacs-bot
rhacs-bot requested a review from janisz as a code owner September 1, 2026 06:36
@coderabbitai

coderabbitai Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: Repository YAML (base), Central YAML (inherited), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Team

Run ID: 56bf5b43-bf22-47dd-b018-7f6cc3ccae72


Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

E2E Test Results

Commit: 51a852f
Workflow Run: View Details
Artifacts: Download test results & logs

=== Evaluation Summary ===

  ✓ cve-cluster-does-exist (assertions: 3/3)
  ✓ list-clusters (assertions: 3/3)
  ✓ cve-cluster-does-not-exist (assertions: 3/3)
  ✓ cve-clusters-general (assertions: 3/3)
  ✓ cve-detected-workloads (assertions: 3/3)
  ~ cve-detected-clusters (assertions: 2/3)
      - MaxToolCalls: Too many tool calls: expected <= 4, got 5
  ✓ cve-multiple (assertions: 3/3)
  ✓ cve-log4shell (assertions: 3/3)
  ✓ cve-cluster-list (assertions: 3/3)
  ~ rhsa-not-supported (assertions: 1/2)
      - MaxToolCalls: Too many tool calls: expected <= 4, got 8
  ✓ cve-nonexistent (assertions: 3/3)

Tasks:      11/11 passed (100.00%)
Assertions: 30/32 passed (93.75%)
Tokens:     ~63133 (estimate - excludes system prompt & cache)
MCP schemas: ~12562 (included in token total)
Agent used tokens:
  Input:  12950 tokens
  Output: 25607 tokens
Judge used tokens:
  Input:  28670 tokens
  Output: 28122 tokens

@codecov-commenter

codecov-commenter commented Sep 1, 2026

Copy link
Copy Markdown

❌ 2 Tests Failed:

Tests completed Failed Passed Skipped
380 2 378 12
View the full list of 2 ❄️ flaky test(s)
::policy 1

Flake rate in main: 100.00% (Passed 0 times, Failed 128 times)

Stack Traces | 0s run time
- test violation 1
- test violation 2
- test violation 3
::policy 4

Flake rate in main: 100.00% (Passed 0 times, Failed 128 times)

Stack Traces | 0s run time
- testing multiple alert violation messages 1
- testing multiple alert violation messages 2
- testing multiple alert violation messages 3

To view more test analytics, go to the Test Analytics Dashboard
📋 Got 3 mins? Take this short survey to help us improve Test Analytics.

@mtodor
mtodor merged commit bb62e20 into main Sep 1, 2026
10 checks passed
@mtodor
mtodor deleted the chore/update-model-evaluation-2026-09-01 branch September 1, 2026 08:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants