Benchgen

WikiWebFacts — Results

RankModelScore
1aliceai-foundation-80b-a3b-base86.5
W

WikiWebFacts

1 phaseActive

Yandex's 5-shot benchmark of Russian-language factual knowledge, released openly alongside AliceAI-Foundation-80B-A3B-Base.

Overview

WikiWebFacts

Category Metric Language Saturation

Quick answer: WikiWebFacts is a 5-shot benchmark of factual knowledge in Russian, published by Yandex together with its evaluation protocol and AliceAI-Foundation-80B-A3B-Base. The model scores 86.5, ahead of DeepSeek-V4-Flash-Base (83.2), Nemotron-3-Super-120B-A12B-Base (72.8) and GLM-4.5-Air-Base (70.2) in Yandex's own comparison.

At a Glance

What it tests: Recall of facts about entities in a Russian-language context, evaluated 5-shot on base models.

Why it matters: Most factual-knowledge benchmarks are English-centric; WikiWebFacts measures how much Russian-language world knowledge a model has absorbed during pre-training.

Known limitations: Published scores come from Yandex's own evaluation infrastructure (vLLM, temperature 0) and compare base models only.

What WikiWebFacts Measures

WikiWebFacts probes factual knowledge that appears on Russian Wikipedia and the Russian web. It is run 5-shot, so it suits base (pre-trained, not instruction-tuned) models, and is reported together with HardMultiQA, a harder multi-hop sibling released at the same time.

Benchmark Specifications

FieldValue
Task categoryKnowledge (Russian)
Metric% accuracy, 5-shot
Created byYandex
Datasethuggingface.co/datasets/yandex/WikiWebFacts

How WikiWebFacts Is Scored

Models are prompted with five examples and scored on factual accuracy using Yandex's published evaluation protocol.

State-of-the-Art Results

RankModelScoreSourceDate
1AliceAI-Foundation-80B-A3B-Base86.5Yandex model card2026-09

Yandex's own comparison (not on Benchgen): DeepSeek-V4-Flash-Base 83.2, Nemotron-3-Super-120B-A12B-Base 72.8, GLM-4.5-Air-Base 70.2, Qwen3.5-35B-A3B-Base 62.4.

WikiWebFacts on Benchgen

No Benchgen results yet — be the first to run WikiWebFacts.

WikiWebFacts vs Other Benchmarks

BenchmarkWhat it testsLanguage
WikiWebFactsFactual knowledge, 5-shotRussian
HardMultiQAHarder multi-fact questions, 5-shotRussian
TriviaQATrivia knowledgeEnglish

Run WikiWebFacts on Your Model

Download the dataset and evaluation protocol from the Hugging Face dataset page.

Frequently Asked Questions

What is WikiWebFacts? A 5-shot benchmark from Yandex that measures factual knowledge in Russian.
What is a good WikiWebFacts score? In Yandex's comparison, 86.5 (AliceAI-Foundation-80B-A3B-Base) leads; strong open base models range from about 62 to 83.
Who created WikiWebFacts? Yandex, released openly on Hugging Face with the AliceAI-Foundation-80B-A3B-Base model.