Benchgen

HardMultiQA — Results

RankModelScore
1aliceai-foundation-80b-a3b-base67.9
H

HardMultiQA

1 phaseActive

Yandex's challenging 5-shot benchmark of Russian-language factual knowledge, released openly alongside AliceAI-Foundation-80B-A3B-Base.

Overview

HardMultiQA

Category Metric Language Saturation

Quick answer: HardMultiQA is a challenging 5-shot benchmark of factual knowledge in Russian, released by Yandex with its evaluation protocol alongside AliceAI-Foundation-80B-A3B-Base. The model scores 67.9, ahead of DeepSeek-V4-Flash-Base (65.4), Nemotron-3-Super-120B-A12B-Base (54.5) and GLM-4.5-Air-Base (48.6) in Yandex's own comparison.

At a Glance

What it tests: Harder, multi-fact questions about the Russian-language world, evaluated 5-shot on base models.

Why it matters: It separates models with deep Russian-language knowledge from those with only shallow coverage, where easier benchmarks saturate.

Known limitations: Scores are from Yandex's own evaluation infrastructure and compare base models only.

What HardMultiQA Measures

HardMultiQA is the harder companion to WikiWebFacts: questions require retrieving several facts from memory, with no tools or retrieval, in a Russian-language context.

Benchmark Specifications

FieldValue
Task categoryKnowledge (Russian)
Metric% accuracy, 5-shot
Created byYandex
Datasethuggingface.co/datasets/yandex/HardMultiQA

How HardMultiQA Is Scored

Models are prompted with five examples and scored on answer accuracy using Yandex's published evaluation protocol.

State-of-the-Art Results

RankModelScoreSourceDate
1AliceAI-Foundation-80B-A3B-Base67.9Yandex model card2026-09

Yandex's own comparison (not on Benchgen): DeepSeek-V4-Flash-Base 65.4, Nemotron-3-Super-120B-A12B-Base 54.5, GLM-4.5-Air-Base 48.6, Qwen3.5-35B-A3B-Base 47.2.

HardMultiQA on Benchgen

No Benchgen results yet — be the first to run HardMultiQA.

HardMultiQA vs Other Benchmarks

BenchmarkWhat it testsLanguage
HardMultiQAHard multi-fact knowledge, 5-shotRussian
WikiWebFactsFactual knowledge, 5-shotRussian
TriviaQATrivia knowledgeEnglish

Run HardMultiQA on Your Model

Download the dataset and evaluation protocol from the Hugging Face dataset page.

Frequently Asked Questions

What is HardMultiQA? A challenging 5-shot benchmark from Yandex that measures multi-fact knowledge in Russian.
What is a good HardMultiQA score? In Yandex's comparison, 67.9 (AliceAI-Foundation-80B-A3B-Base) leads; other strong open base models score about 47 to 65.
Who created HardMultiQA? Yandex, released openly on Hugging Face with the AliceAI-Foundation-80B-A3B-Base model.