A lightweight REST API service for orchestrating LLM evaluations across multiple backends. Written in Go, it routes evaluation requests to frameworks like lm-evaluation-harness, RAGAS, Garak, and GuideLLM orchestrated via a complementary SDK, tracks experiments via MLflow, and runs natively on OpenShift.
flowchart TB
subgraph clients[Clients]
direction LR
rest[REST clients]
agents[MCP clients / agents]
end
subgraph openshift[OpenShift cluster]
direction TB
proxy[kube-rbac-proxy<br/>authentication and authorization]
api[EvalHub API<br/>REST routes and handlers]
mcp[EvalHub MCP server<br/>tools, resources, prompts<br/>stdio or HTTP]
mlflow[MLflow]
otelCollector[OpenTelemetry Collector<br/>OTLP endpoint]
database[(SQL storage<br/>PostgreSQL typical)]
config[YAML config<br/>providers and collections]
metrics[Metrics listener<br/>:8081 /metrics]
prometheus[Prometheus]
runtime[Kubernetes runtime]
kubeapi[Kubernetes API]
subgraph job[Evaluation job pod]
direction TB
init[Optional data/source init]
adapter[Provider adapter]
sidecar[EvalHub sidecar<br/>API, model, MLflow and OCI proxies]
init -.-> adapter
adapter <--> sidecar
end
end
subgraph externalEndpoints[External endpoints]
direction LR
model[Model endpoints]
registry[OCI registry]
end
rest -->|REST| proxy
proxy --> api
agents -->|MCP| mcp
mcp -->|EvalHub REST client| proxy
api --> database
config --> api
api --> runtime
runtime -->|create jobs| kubeapi
kubeapi --> job
sidecar -->|status events| proxy
sidecar --> mlflow
api -->|tracking and results| mlflow
sidecar -->|Model traffic| model
sidecar -->|OCI traffic| registry
prometheus -->|scrape| metrics
api -. OTLP .-> otelCollector
adapter -. OTLP .-> otelCollector
sidecar -. OTLP .-> otelCollector
classDef external fill:#f5f7fa,stroke:#64748b,color:#1e293b
classDef service fill:#e8f2ff,stroke:#3973ac,color:#142b45
classDef runtimeNode fill:#eef8f1,stroke:#4b8b62,color:#193d26
class rest,agents,model,registry external
class proxy,api,mcp,mlflow,otelCollector,database,config,metrics service
class runtime,kubeapi,init,adapter,sidecar runtimeNode
OpenShift API requests pass through kube-rbac-proxy, which supplies the authenticated user and tenant identity. The API creates Kubernetes jobs; each job pod runs a provider adapter and sidecar. The sidecar proxies model, MLflow, OCI, and API callback traffic. Prometheus scrapes the dedicated metrics listener. The MCP server runs inside the cluster as a process separate from the API.
flowchart TB
rest[REST clients]
agents[MCP clients / agents]
mcp[EvalHub MCP server<br/>tools, resources, prompts<br/>stdio or HTTP]
prometheus[Prometheus]
model[Model endpoints]
mlflow[MLflow<br/>optional]
subgraph local[Local machine]
direction TB
api[EvalHub API<br/>REST routes and handlers]
database[(SQLite by default<br/>PostgreSQL optional)]
config[YAML config<br/>providers and collections]
metrics[Metrics listener<br/>:8081 /metrics]
runtime[Local runtime]
adapter[Provider adapter<br/>local process]
sidecar[Optional local sidecar<br/>API and per-job model proxy]
end
rest -->|REST| api
agents -->|MCP| mcp
mcp -->|EvalHub REST client| api
api --> database
config --> api
api --> runtime
runtime -->|launch| adapter
adapter -->|API callbacks when sidecar is off| api
adapter -. when enabled .-> sidecar
sidecar -->|API callbacks| api
adapter -. direct model calls when sidecar is off .-> model
sidecar -->|when enabled| model
api -->|tracking and results| mlflow
prometheus -->|scrape| metrics
classDef external fill:#f5f7fa,stroke:#64748b,color:#1e293b
classDef service fill:#e8f2ff,stroke:#3973ac,color:#142b45
classDef runtimeNode fill:#eef8f1,stroke:#4b8b62,color:#193d26
class rest,agents,mcp,prometheus,model,mlflow external
class api,database,config,metrics service
class runtime,adapter,sidecar runtimeNode
Local mode connects directly to the API without kube-rbac-proxy and runs each adapter as a local process. The shared sidecar is optional; when enabled, it handles API callbacks and per-job model routing. MLflow can be configured for experiment tracking and result export. Prometheus can scrape the separate metrics listener; local mode also exposes /metrics on the API port. The MCP server is a separate process and can run near the MCP client or alongside the API.
- Go 1.26+
- Make
- Python 3 (for
make test; used by scripts/grcat for colored output) - uv (manages the Python venv required by
make start-serviceand FVT tests; runmake venvto create it) - Podman (for container builds)
- Access to an OpenShift or Kubernetes cluster (for deployment)
make install-deps
make build
./bin/eval-hubNote that in some cases it may be necessary to exclude certain (newer) dependencies,
this can be done as shown below in the go.mod file:
exclude (
k8s.io/api v0.36.0
)The API is available at http://localhost:8080. Verify it is running:
curl http://localhost:8080/api/v1/healthInteractive documentation is served at /docs.
podman build -t eval-hub:latest -f Containerfile .
podman run --rm -p 8080:8080 eval-hub:latestEvalHub is managed by the TrustyAI Service Operator via a custom resource:
apiVersion: trustyai.opendatahub.io/v1alpha1
kind: EvalHub
metadata:
name: evalhub
namespace: my-namespace
spec:
replicas: 1
env:
- name: MLFLOW_TRACKING_URI
value: "http://mlflow:5000"
- name: EVALHUB_HARDWARE_PROFILES_NAMESPACE
value: "opendatahub" # or redhat-ods-applications on RHOAIApply the CR to your cluster:
oc apply -f evalhub-cr.yaml
oc get evalhub -n my-namespace # check statusmake start-service # start in background (logs to bin/service.log)
make stop-service # stop
make test # unit tests
make test-fvt # BDD functional tests (godog)
make test-all # both
make test-coverage # generate coverage.html
make lint # go vet
make fmt # go fmtFor local post-processing, the runtime launches the real adapter from a sibling
../eval-hub-contrib/adapters/evalhub-post-processor checkout. Install that
adapter's requirements.txt in its .venv before starting EvalHub from this
repository root. For another checkout or Python environment, set
EVALHUB_POST_PROCESSING_LOCAL_COMMAND to the command that runs its main.py.
The adapter also needs a reachable callback service and access to the referenced
data; see the adapter's README for local setup.
Run a single test:
go test -v ./internal/handlers -run TestHandleNameTo create a Python wheel distribution of the server for local development and testing:
make cross-compile
make build-wheelCreate a file called export_test.go in the package under test and re-export symbols needed by _test.go files in other packages.
SQLite in-memory is the default (database.driver: sqlite in config/config.yaml). To use PostgreSQL locally there are two approaches: a container or a native install. Both use targets in tests/postgres/Makefile.
Note: The credentials and auth settings below are for local development and testing only. For production deployments, use strong passwords, TLS, and appropriate authentication mechanisms.
No system-level install required. The container creates the database, user, and permissions automatically.
cd tests/postgres
POSTGRES_PASSWORD=<your-password> make start-postgres-containerTo stop and remove:
cd tests/postgres
make stop-postgres-container
make delete-postgres-containerConfigure EvalHub in config/config.yaml:
database:
driver: pgx
url: postgres://eval_hub:<your-password>@localhost:5432/eval_hubOr override via environment variables:
export DB_DRIVER=pgx
export DB_URL="postgres://eval_hub:<your-password>@localhost:5432/eval_hub"cd tests/postgres
make install-postgres
make start-postgres
make create-user
make create-database
make grant-permissionsTo stop:
cd tests/postgres
make stop-postgresConfigure EvalHub in config/config.yaml (no password needed with trust/peer auth):
database:
driver: pgx
url: postgres://eval_hub@localhost:5432/eval_hubOr override via environment variables:
export DB_DRIVER=pgx
export DB_URL="postgres://eval_hub@localhost:5432/eval_hub"Configuration is loaded from config/config.yaml, overridden by environment variables and secret files.
| Variable | Purpose | Default |
|---|---|---|
PORT |
API listen port | 8080 |
DB_DRIVER |
Database driver (sqlite or pgx) |
sqlite |
DB_URL |
Database connection string | SQLite in-memory |
MLFLOW_TRACKING_URI |
MLflow tracking server | http://localhost:5000 |
MLFLOW_CA_CERT_PATH |
PEM CA bundle for MLflow TLS verification | (system roots) |
LOG_LEVEL |
Logging level | INFO |
EVALHUB_HARDWARE_PROFILES_NAMESPACE |
Platform namespace where OpenDataHub HardwareProfile CRs are fetched (Kubernetes runtime). Required for hardware_config.hardware_profile_name evaluations; typically opendatahub or redhat-ods-applications. Set by the TrustyAI Service Operator deployment. |
(unset — hardware profile lookups fail) |
Provider configurations live in config/providers/ as YAML files. The default set includes lm-evaluation-harness (167 benchmarks), RAGAS, Garak, GuideLLM, LightEval, and MTEB.
Provider and collection definitions are maintained here and mirrored as ConfigMaps in the TrustyAI Service Operator:
- Providers:
config/providers/→config/configmaps/evalhub/provider-*.yaml - Collections:
config/collections/→config/configmaps/evalhub/collection-*.yaml
When adding or changing a provider or collection, update the source YAML in this repository and the corresponding embedded ConfigMap in the operator repository. Add new ConfigMaps to the operator's config/configmaps/evalhub/kustomization.yaml. Keep the two repositories' changes coordinated so the operator can deploy the same definitions.
The sync check compares the embedded ConfigMap YAML with the source files. Run it locally with:
python scripts/check_configmap_sync.pyThe same check runs in CI through the TrustyAI Operator ConfigMap Sync workflow.
All endpoints are versioned under /api/v1. Full specification at eval-hub.github.io/eval-hub.
| Endpoint | Methods | Description |
|---|---|---|
/api/v1/evaluations/jobs |
POST, GET | Create or list evaluation jobs |
/api/v1/evaluations/jobs/{id} |
GET, DELETE | Get status or cancel a job |
/api/v1/evaluations/collections |
GET, POST | List or create benchmark collections |
/api/v1/evaluations/providers |
GET, POST | List or create providers |
/api/v1/evaluations/providers/{id} |
GET, PUT, PATCH, DELETE | Manage a provider |
/api/v1/evaluations/jobs/{id}/events |
POST | Submit job events |
/api/v1/health |
GET | Health check (no identity headers; no build/version fields) |
/metrics |
GET | Prometheus metrics |
Detailed API documentation: eval-hub.github.io/eval-hub
EvalHub supports Bring Your Own Framework (BYOF). Extend the FrameworkAdapter class from the eval-hub-sdk and implement a single method -- EvalHub handles scheduling, status reporting, and result aggregation.
from evalhub.adapter import FrameworkAdapter, JobSpec, JobCallbacks, JobResults, EvaluationResult
class MyAdapter(FrameworkAdapter):
def run_benchmark_job(self, config: JobSpec, callbacks: JobCallbacks) -> JobResults:
# run your evaluation logic, report progress via callbacks
callbacks.report_status(JobStatusUpdate(status=JobStatus.RUNNING, progress=0.5))
score = evaluate(config.model, config.parameters)
return JobResults(
id=config.id,
benchmark_id=config.benchmark_id,
model_name=config.model.name,
results=[EvaluationResult(metric_name="accuracy", metric_value=score)],
num_examples_evaluated=100,
duration_seconds=elapsed,
)Register the new provider by adding a YAML entry to the providers ConfigMap. No additional services or TCP listeners are required -- adapters run as jobs, not servers. Once registered, the provider and its benchmarks are available through the standard /api/v1/evaluations/providers endpoint.
eval-hub/
├── cmd/eval_hub/ # Entry point (main binary)
├── internal/
│ ├── handlers/ # HTTP request handlers
│ ├── storage/ # Database abstraction (SQLite, PostgreSQL)
│ ├── mlflow/ # MLflow client
│ ├── runtimes/ # Backend execution adapters
│ ├── config/ # Viper-based configuration
│ ├── validation/ # Request validation
│ ├── metrics/ # Prometheus instrumentation
│ └── logging/ # Structured logging (zap)
├── config/ # config.yaml and provider definitions
├── docs/src/ # OpenAPI 3.1.0 specification (source of truth)
├── tests/features/ # BDD tests (godog)
├── Containerfile # Multi-stage UBI9 container build
└── Makefile # Build, test, and dev targets
EvalHub can run evaluations locally without a Kubernetes cluster. See the local mode guide for configuration, architecture details, and troubleshooting, and the local mode tutorial for a step-by-step walkthrough. A self-contained LightEval example is included in this repository.
When local mode has mlflow.tracking_uri configured, each local evaluation subprocess automatically receives that direct URI as MLFLOW_TRACKING_URI. Local subprocess environment variables are applied in this order, with later values replacing matching earlier values:
- Inherited process environment
- EvalHub and service configuration values
- Provider
runtime.local.envvalues
- API documentation -- full endpoint reference
- Local mode guide -- running evaluations without Kubernetes
- CONTRIBUTING.md -- contribution guidelines
- OpenAPI spec -- machine-readable API definition
Apache 2.0 -- see LICENSE.