Publication view · 02
Benchmark coverage
Evaluation categories
Independent evidence tracks. Metrics remain within their own evaluation category.
Agentic evidence
Campaign-weighted means · deterministic scores remain separate from judge assessment
Retrieval quality
Longitudinal overview
Benchmark evolution
Every measured campaign, ordered through time. Curves never connect different workload protocols.
Time series
One line per engine and verified protocol
Engine trajectories
P95 latency
Success rate
Average CPU
Peak memory
Run archive
Test series
| Run | Started | Protocol | Panel | State | Open |
|---|
Metric dossier
Mean and leader use valid measurements only
Corrobore against the full lab
Complete panel
Engine leaderboard
| Engine | State | Throughput | P95 latency | Success | Avg CPU | Peak memory |
|---|
Measured distribution
Performance and efficiency
Request performance
Resource efficiency
Audit trail
Reproducibility
Graph working memory
Agentic retrieval quality
Deterministic fact recovery and independent judge assessments across the six-engine panel.
Quality panel
Judge values never alter deterministic F1
Fact retrieval and judgment
| Engine | State | Precision | Recall | Micro F1 | Macro F1 | Judge |
|---|