Our benchmarking program is currently in measurement. We publish performance metrics only after they have been independently verified, documented with a reproducible methodology, and validated across representative workloads.
Document-extraction accuracy
Field-level precision/recall on public invoice/receipt datasets + methodology for testing on YOUR documents in a feasibility sprint.
RAG faithfulness
Citation-supported answer rates and honest-abstention rates for Knowledge across corpus types, with the eval harness published alongside.
Modernization throughput
Characterization-test coverage per week and slice cutover cadence across representative COBOL/VB6 estates - ranges, not averages, with estate-size context.
Cloud cost baselines
Typical waste distribution by category (rightsizing/scheduling/storage/orphans) across FinOps engagements, anonymized and aggregated.
- →Test environment and hardware specifications
- →Datasets and workload definitions
- →Software and model versions
- →Benchmark methodology
- →Reproducibility instructions
- →Raw performance results and analysis
We believe transparent and reproducible benchmarking is essential for enterprise customers evaluating platforms. Want a benchmark run against your data under NDA? That's usually the more useful conversation - book a demo.