Measure AI on
real business work.

Public leaderboards for language models, document-processing AI, long-term memory and voicebots, measured on Vietnamese data, business criteria and a cost of use whose calculation is stated.

  • Public leaderboardsEvery score comes with a 95% confidence interval.
  • Closed test setsQuestions and answers stay unpublished, so nobody can memorise them.
  • Transparent methodologyConfiguration, number of runs and limits are all published.

Evaluation areas

Each area answers its own question, with its own metrics and leaderboard. We do not combine the areas into one overall score.

Other benchmarks are on the roadmap, with nothing measured yet. See all benchmarks →

Leaderboards

Illustrative data — not real measurements

The public leaderboard of each test set, shown as a scorecard: every score comes with a 95% confidence interval. A shared rank means a paired comparison cannot yet separate the two sides; it does not mean they are equal.

Column leader, including sides not yet separatedCompare 2–4 side by side →
Illustrative data — not real measurements
ASIABENCH TEST SETS

One common standard for many systems

Many solutions are measured against the same standard to build the public ranking. Questions and answers are kept closed; the method and the aggregate results are published.

  • Decides the public leaderboard
  • Uses no customer data
CUSTOMER-SPECIFIC TEST SETS

Built from one organisation’s data and criteria

Data-use rights, ownership of the test set and the scope of publishing results are set in the contract.

  • Not on the public leaderboard by default
  • Never turned into training data
Keeping the questions closed is only part of the advantage. The rest is the quality and representativeness of the test set, how often it is updated, the expertise of the graders and the credibility of the results.

Compare

illustrative

Pick a benchmark and 2–4 solutions. The answer comes first, the measurements after.

Illustrative data — not real measurements. The names “ A–H” are placeholders too.
Benchmark
COST ASSUMPTIONS · NOT MEASUREMENTS

Conclusion

SHORTLIST FOR THE NEXT ROUND

How to read the labels. Each label compares one side with the side that has the best measured value, as pairs on the same set of questions, using a business margin fixed in advance () and corrected for multiple comparisons. Behind: the interval of the difference lies wholly on one side. Equivalent: the whole interval lies within the margin. Not enough evidence: no conclusion yet, which does not mean the two are equal. Staying on the shortlist does not mean good enough either. Because the data is illustrative, the interval of each difference is rebuilt from an assumed correlation of between the two systems; real results will use a paired bootstrap.

Metric by metric

Everyone sits on the same axis. The pale blue band is only the own confidence interval of the side with the best value, for reference; the label on the right comes from a paired comparison, not from whether two intervals touch.
See the full analysisQuality × cost, cost per accepted task (estimated), by task group, procurement table
Quality × cost per accepted task (estimated)Confidence intervals in both directions. Top-left corner: better and cheaper.
GlobalVietnamOpen weightsShortlist for the next round
Cost per accepted task (estimated) = machine cost + person handling errorsA cheap list price is not always cheap once a person is counted: the yellow stripes are the work of handling the machine’s errors, using the assumptions above. Scope: not included are integration, checking the tasks that passed, and the cost of errors that slip through.
machine cost (measured)person handling errors (assumed)├┤ 95% confidence interval of the total

A side is called the leader only when the paired comparison separates it from every other side.

For procurement

✓ yes · ◐ depends on the contract or on where you host · ✗ no.

Illustrative data — not real measurements. The interval of the cost per accepted task is combined from the machine-price interval and the completion rate; hourly wage and minutes to handle an error are assumptions you can change, not measurements. When real results are published, each cell will carry the number of items run, the number of runs and the date measured.

Benchmarks

Tasks that global benchmarks don’t cover yet. Every result comes with a 95% confidence interval. The benchmarks in the first group have only illustrative data for now; those on the roadmap have nothing measured yet.

CodeBenchmarkMeasuresStatusOpen
AB-VietLLMVietnamese LLMs, updated monthlyFixed core + monthly rotation · Text generation · AgentOctober 2026 · illustrative
AB-DocVietnamese documentsAccuracy · Cost per documentv1 · illustrative
AB-VoiceCall-center voice agentsListening · Speech · Dialogue · Task · Memoryv1 · illustrative
AB-MemoryLong context and memoryLong context (base model) · Recall across sessions (system)v1 · illustrative
On the roadmap · nothing measured yet
AB-LivenessFace anti-spoofingPrinted photos · Screen replays · MasksRoadmap
AB-DeepfakeDeepfakes and injection attacksFace swaps · AI-generated video · Virtual camerasRoadmap
AB-AgentOffice agentsEnd-to-end completion · Wrong actionsRoadmap
AB-LegalVietnamese lawCorrect citations · HallucinationRoadmap
AB-TrustSafety and complianceData leakage · BiasRoadmap
AB-SEAThai · Malay · KhmerBy language · By countryRoadmap

Methodology

Trust is the only thing we sell, so the rules are public from day one. The process takes ISO/IEC 17025 as a reference; AsiaBench is not accredited to that standard.

“No score stands alone.”

Every test set has a description card: goal, what is measured, task scope, sampling, sample size, data source and usage rights, rubric, how scores are combined, and limits. The overall score is only for a quick read; reports can always be traced down by task group, error type and severity.

v0.1 principles · coming soon
  1. R1Ranks are not for saleService fees never depend on scores or ranks. Nobody pays to get in, move up or leave the board; the rules for choosing systems, sponsorship, updates, appeals and withdrawing results will be published in the v0.1 principles.
  2. R2Only commercial versions are rankedThe version customers can actually buy. Models, agents and complete solutions are measured separately and never read as one another.
  3. R3Closed questions: fixed core, rotating partThe fixed part allows comparison with earlier months; the rotating part guards against memorisation. Every round has a set procedure and a preset number of runs.
  4. R4Paired comparison, always with a confidence intervalA shared rank means not enough evidence to separate, not that the two are equal. The business-relevant margin is set in advance; many pairs are corrected for.
  5. R5Every measurement records its conditionsVersion, configuration, resource budget, number of runs and date. Parts managed by the provider that we cannot observe are stated as limits.
  6. R6People check the answersGround truth and rubrics are reviewed by experts. If an LLM is used as a judge, its agreement with human graders and the way disagreements are handled are published.
  7. R7Customer data stays separateNever added to the public leaderboard and never turned into training data, unless the contract says otherwise.

Services

Choose AI by measurements on your own data, from a party that does not sell AI. The public leaderboards are always free, and service fees never depend on scores or ranks.

Why you need a test set from your own data
Document readingCall-center voicebotsCustomer-care chatbotsInternal knowledge searchOffice agents
By default: not published, no effect on the public leaderboard, not used as training data, and vendors cannot buy it. Data-use rights, ownership of the test set and the scope of publication are set in the contract.
PUBLIC LEADERBOARDYOUR PRIVATE TESTSSolution A · 1Solution B · 2Solution C · 3Solution D · 4Solution E · 52 · A ↓14 · B ↓21 · C ↑25 · D ↓13 · E ↑2 Illustration: a solution ranked 3rd on the public leaderboard can come first on your data, and the reverse. That is why you need a private test set before signing a contract.
  1. 01Sample your real dataDocuments, calls, customer questions; anonymised before use.
  2. 02Write business grading criteriaA wrong amount weighs more than a typo; native speakers grade.
  3. 03Run 3–5 solutions side by sideSame tests, same configuration, with confidence intervals.
  4. 04Reuse for acceptance and monitoringScore thresholds in the contract; re-test whenever the vendor updates.
SELECTION ADVISORY · 6 STEPSOnly buyers pay
  1. 01Scope the problemTasks, data, success metrics
  2. 02Customer-specific test setFrom your data and criteria
  3. 03Shortlist evaluationCompare 3–5 solutions on real data
  4. 04Architecture and costBuild or buy, API or self-hosted
  5. 05Score-based acceptanceScore thresholds written into the contract
  6. 06MonitoringPeriodic re-testing after rollout
NO VENDOR COMMISSIONS · NO IMPLEMENTATION WORK · EVERY RECOMMENDATION COMES WITH TEST DATA

Bring your solution into the lab.

Business, vendor or model lab: leave your email and we’ll get back to you within a few business days.