Measure AI on
real business work.
Public leaderboards for language models, document-processing AI, long-term memory and voicebots, measured on Vietnamese data, business criteria and a cost of use whose calculation is stated.
- Public leaderboardsEvery score comes with a 95% confidence interval.
- Closed test setsQuestions and answers stay unpublished, so nobody can memorise them.
- Transparent methodologyConfiguration, number of runs and limits are all published.
-
ILLUSTRATIVE · AB-VOICE V1Call-center voice agents: a previewEight voice agents tested on banking and telecom scripts, with accents from all three regions. Illustrative data.
-
ILLUSTRATIVE · OCTOBER 2026AB-VietLLM: new monthly leaderboardThe rotating part of the test set is replaced; the fixed core is kept for comparison with earlier months.
-
ILLUSTRATIVE · AB-MEMORY V1Long-term memory in VietnameseRecall across sessions, long context and consistent forms of address. Illustrative data.
Evaluation areas
Each area answers its own question, with its own metrics and leaderboard. We do not combine the areas into one overall score.
- AB-VietLLMLanguage models
Which model fits a specific Vietnamese task?
Illustrative data - AB-DocDocument processing
Does the solution process documents correctly and cut manual work?
Illustrative data - AB-VoiceVoicebot
Does the voicebot complete real calls and handle situations correctly?
Illustrative data - AB-MemoryLong-term memory
Does the system remember and use information correctly across sessions? Adds a long-context part measured on the base model.
Illustrative data
Other benchmarks are on the roadmap, with nothing measured yet. See all benchmarks →
Leaderboards
Illustrative data — not real measurementsThe public leaderboard of each test set, shown as a scorecard: every score comes with a 95% confidence interval. A shared rank means a paired comparison cannot yet separate the two sides; it does not mean they are equal.
One common standard for many systems
Many solutions are measured against the same standard to build the public ranking. Questions and answers are kept closed; the method and the aggregate results are published.
- Decides the public leaderboard
- Uses no customer data
Built from one organisation’s data and criteria
Data-use rights, ownership of the test set and the scope of publishing results are set in the contract.
- Not on the public leaderboard by default
- Never turned into training data
Compare
illustrativePick a benchmark and 2–4 solutions. The answer comes first, the measurements after.
Conclusion
How to read the labels. Each label compares one side with the side that has the best measured value, as pairs on the same set of questions, using a business margin fixed in advance () and corrected for multiple comparisons. Behind: the interval of the difference lies wholly on one side. Equivalent: the whole interval lies within the margin. Not enough evidence: no conclusion yet, which does not mean the two are equal. Staying on the shortlist does not mean good enough either. Because the data is illustrative, the interval of each difference is rebuilt from an assumed correlation of between the two systems; real results will use a paired bootstrap.
Metric by metric
Everyone sits on the same axis. The pale blue band is only the own confidence interval of the side with the best value, for reference; the label on the right comes from a paired comparison, not from whether two intervals touch.See the full analysisQuality × cost, cost per accepted task (estimated), by task group, procurement table
For procurement
✓ yes · ◐ depends on the contract or on where you host · ✗ no.Illustrative data — not real measurements. The interval of the cost per accepted task is combined from the machine-price interval and the completion rate; hourly wage and minutes to handle an error are assumptions you can change, not measurements. When real results are published, each cell will carry the number of items run, the number of runs and the date measured.
Benchmarks
Tasks that global benchmarks don’t cover yet. Every result comes with a 95% confidence interval. The benchmarks in the first group have only illustrative data for now; those on the roadmap have nothing measured yet.
Methodology
Trust is the only thing we sell, so the rules are public from day one. The process takes ISO/IEC 17025 as a reference; AsiaBench is not accredited to that standard.
“No score stands alone.”
Every test set has a description card: goal, what is measured, task scope, sampling, sample size, data source and usage rights, rubric, how scores are combined, and limits. The overall score is only for a quick read; reports can always be traced down by task group, error type and severity.
- R1Ranks are not for saleService fees never depend on scores or ranks. Nobody pays to get in, move up or leave the board; the rules for choosing systems, sponsorship, updates, appeals and withdrawing results will be published in the v0.1 principles.
- R2Only commercial versions are rankedThe version customers can actually buy. Models, agents and complete solutions are measured separately and never read as one another.
- R3Closed questions: fixed core, rotating partThe fixed part allows comparison with earlier months; the rotating part guards against memorisation. Every round has a set procedure and a preset number of runs.
- R4Paired comparison, always with a confidence intervalA shared rank means not enough evidence to separate, not that the two are equal. The business-relevant margin is set in advance; many pairs are corrected for.
- R5Every measurement records its conditionsVersion, configuration, resource budget, number of runs and date. Parts managed by the provider that we cannot observe are stated as limits.
- R6People check the answersGround truth and rubrics are reviewed by experts. If an LLM is used as a judge, its agreement with human graders and the way disagreements are handled are published.
- R7Customer data stays separateNever added to the public leaderboard and never turned into training data, unless the contract says otherwise.
Services
Choose AI by measurements on your own data, from a party that does not sell AI. The public leaderboards are always free, and service fees never depend on scores or ranks.
AsiaBench Verified
One independent run for a specific version and configuration, on the task scope stated in the report, valid for a stated period. Not a guarantee for every use case, and not a certification or accreditation. It does not change the public ranking.
The report records version, configuration, scope, date measured and validity period. It can serve as material when preparing a file under the AI Law; AsiaBench is not a conformity assessment body.
Pre-launch evaluation
A blind evaluation before launch and periodic re-runs on closed test sets, graded by native speakers. Reports are by task group and error type; questions and answers are never handed back. Private results are confidential under the contract; publishing to the leaderboard follows one common policy.
Selection advisory
From scoping the problem to score-based acceptance written into the contract. No commissions, no implementation work.
A test set from your own data
Built from a business’s real data and criteria, then reused for tenders, acceptance and monitoring. Data-use rights and ownership of the test set are set in the contract; results do not enter the public leaderboard.
Quoted by scope after a first conversation.
Why you need a test set from your own data
- 01Sample your real dataDocuments, calls, customer questions; anonymised before use.
- 02Write business grading criteriaA wrong amount weighs more than a typo; native speakers grade.
- 03Run 3–5 solutions side by sideSame tests, same configuration, with confidence intervals.
- 04Reuse for acceptance and monitoringScore thresholds in the contract; re-test whenever the vendor updates.
- 01Scope the problemTasks, data, success metrics
- 02Customer-specific test setFrom your data and criteria
- 03Shortlist evaluationCompare 3–5 solutions on real data
- 04Architecture and costBuild or buy, API or self-hosted
- 05Score-based acceptanceScore thresholds written into the contract
- 06MonitoringPeriodic re-testing after rollout
Certification
The AI Law requires high-risk AI to go through conformity assessment. We take the laboratory standard as a reference from day one.
Benchmarks and private evaluations
First test sets, paid pilots.
ISO/IEC 17025 testing lab
Accredited technical testing.
Conformity assessment
Under the AI Law, then across the region.
Bring your solution into the lab.
Business, vendor or model lab: leave your email and we’ll get back to you within a few business days.