This happened recently, and it will be relevant as well sam-paech/spiral-bench#1
A lot of the factors of EQ-Bench 3 seems to be correlated to one another, maybe there are meta-factors of how LLMs can function? Feels like "intuitiveness" and "humanity" and "charisma" (maybe "complier") are all different traits of different LLMs in case people want to add this as a weight to existing performance benchmarks. Maybe there are linguistic profiles that causes hallucinations or failure to use tools? https://github.com/vectara/hallucination-leaderboard https://gorilla.cs.berkeley.edu/leaderboard.html
This happened recently, and it will be relevant as well sam-paech/spiral-bench#1
A lot of the factors of EQ-Bench 3 seems to be correlated to one another, maybe there are meta-factors of how LLMs can function? Feels like "intuitiveness" and "humanity" and "charisma" (maybe "complier") are all different traits of different LLMs in case people want to add this as a weight to existing performance benchmarks. Maybe there are linguistic profiles that causes hallucinations or failure to use tools? https://github.com/vectara/hallucination-leaderboard https://gorilla.cs.berkeley.edu/leaderboard.html