Benchgen
Models/yandex/

AliceAI-Foundation-80B-A3B-Base

DraftPublic

Model Details

AliceAI-Foundation-80B-A3B-Base

Organization License Params Context Type

Quick answer: AliceAI-Foundation-80B-A3B-Base is Yandex's open pre-trained (base) language model: 80B total and 3B active parameters, a hybrid KDA/gated-attention MoE with 512 experts, a 262,144-token context, trained entirely from scratch and released September 12, 2026 under Apache 2.0. Yandex reports particularly strong Russian factual knowledge (WikiWebFacts 86.5, HardMultiQA 67.9) and 91.1 on MATH-500.

At a Glance

Where it leads: Russian-language factual knowledge and exams (WikiWebFacts 86.5, CultCat 86.5, EGE CoT 90.5) and math (MATH-500 91.1) among the open base models Yandex compared.

Where it lags: English trivia (TriviaQA 79.0 vs 89.8 for Nemotron-3-Super-120B-A12B-Base) and SuperGPQA (44.3 vs 46.6).

Best for: Research and fine-tuning for Russian-language assistants; it is a base model with no post-training or alignment.

What AliceAI-Foundation-80B-A3B-Base Is

Yandex built a new training corpus, chose the architecture and hyperparameters through a series of 2-trillion-token from-scratch runs, and prepared data for complex reasoning and tool use. The 48-layer network repeats a pattern of three KDA-MoE layers followed by one gated-attention-MoE layer, with 512 experts (top-10 routed plus one shared) and a 1-layer MTP head.

The model is a pre-trained checkpoint intended for research, experimentation and further tuning rather than direct use in products. Alongside the weights Yandex published two Russian factual benchmarks, WikiWebFacts and HardMultiQA, with their evaluation protocols.

Specifications

FieldValue
OrganizationYandex
Hugging Faceyandex/AliceAI-Foundation-80B-A3B-Base
ArchitectureHybrid KDA + gated-attention MoE, 80B total / 3B active
LanguagesRussian, English
Context length262,144 tokens
LicenseApache 2.0
Release date2026-09-12

Public Benchmark Scores

BenchmarkScoreSourceDate
WikiWebFacts86.5Model card2026-09
HardMultiQA67.9Model card2026-09
MMLU-Pro (5-shot CoT, base)66.8Model card2026-09
SuperGPQA (5-shot CoT, base)44.3Model card2026-09
MATH-500 (5-shot, base)91.1Model card2026-09

Self-reported by Yandex on a pre-trained base model; not independent Benchgen measurements. Also reported but not added: Yandex-internal benchmarks (CultCat, EduBench, ExpertFactsQA, EGE CoT, FinQA 128k, LongMemEval 128k), TriviaQA (LLM-as-judge instead of exact match), BigCodeBench (Yandex's own implementation with improved tests), LiveCodeBench v5-6 (not v6) and pass@k reasoning results (AIME 2026 pass@32 96.7, HMMT 2026 Feb pass@32 96.9, IMO AnswerBench pass@8 88.7).

Frequently Asked Questions

What is AliceAI-Foundation-80B-A3B-Base? Yandex's open base language model with 80B total and 3B active parameters, trained from scratch and released September 12, 2026 under Apache 2.0.
Is it an instruction-tuned chat model? No. It is a pre-trained base model with no post-training or alignment; Yandex recommends fine-tuning it before production use.
What is its context length? 262,144 tokens.