Skip to content

Can correlation analysis be applied to each of the factors of EQBench3? #2

Description

@TomLucidor

This happened recently, and it will be relevant as well sam-paech/spiral-bench#1

A lot of the factors of EQ-Bench 3 seems to be correlated to one another, maybe there are meta-factors of how LLMs can function? Feels like "intuitiveness" and "humanity" and "charisma" (maybe "complier") are all different traits of different LLMs in case people want to add this as a weight to existing performance benchmarks. Maybe there are linguistic profiles that causes hallucinations or failure to use tools? https://github.com/vectara/hallucination-leaderboard https://gorilla.cs.berkeley.edu/leaderboard.html

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions