BENCHMARK PROGRAM

Our benchmarking program is currently in measurement. We publish performance metrics only after they have been independently verified, documented with a reproducible methodology, and validated across representative workloads.

THE MEASUREMENT PROGRAM

Document-extraction accuracy

Field-level precision/recall on public invoice/receipt datasets + methodology for testing on YOUR documents in a feasibility sprint.

IN MEASUREMENT

RAG faithfulness

Citation-supported answer rates and honest-abstention rates for Knowledge across corpus types, with the eval harness published alongside.

IN MEASUREMENT

Modernization throughput

Characterization-test coverage per week and slice cutover cadence across representative COBOL/VB6 estates - ranges, not averages, with estate-size context.

COLLECTING

Cloud cost baselines

Typical waste distribution by category (rightsizing/scheduling/storage/orphans) across FinOps engagements, anonymized and aggregated.

COLLECTING

FUTURE BENCHMARK REPORTS WILL INCLUDE

  • Test environment and hardware specifications
  • Datasets and workload definitions
  • Software and model versions
  • Benchmark methodology
  • Reproducibility instructions
  • Raw performance results and analysis

We believe transparent and reproducible benchmarking is essential for enterprise customers evaluating platforms. Want a benchmark run against your data under NDA? That's usually the more useful conversation - book a demo.