Quick answer: Hermes 3 70B is NousResearch's August 2024 instruction fine-tune of Meta's Llama 3.1 70B. It scores 47.2% on MMLU-Pro. Under the Llama 3.1 License with open weights on Hugging Face, it is designed for strong general reasoning, long-context tasks, and agentic deployment patterns.
Where Hermes 3 70B leads
Where it lags
Best for: Agentic and tool-use pipelines requiring open-weight models; teams specifically using NousResearch's Hermes instruction style; legacy Hermes-based deployments.
Hermes 3 70B is NousResearch's third generation Hermes fine-tune, built on Meta's Llama 3.1 70B base. Released August 12, 2024, it specialises in advanced instruction following, function calling, and agentic capability — areas where the base Llama 3.1 70B was strong but could be improved through targeted fine-tuning.
The Hermes series is known in the open-source community for producing models with reliable tool use, structured output generation, and consistent behaviour with complex system prompts. Hermes 3 70B brings these properties to the Llama 3.1 70B scale.
Its 47.2% MMLU-Pro score reflects its 2024 vintage and 70B base — newer open-weight models like Llama 4 Maverick (80.5%), Qwen3 32B (no MMLU-Pro score but strong AetherCode), and Phi-4 (70.4% at only 14B) have since surpassed it on academic knowledge benchmarks. Hermes 3 70B remains relevant for use cases that specifically benefit from NousResearch's instruction-tuning approach and are already integrated with Hermes-style prompting.
| Field | Value |
|---|---|
| Organization | NousResearch |
| Base model | Meta Llama 3.1 70B |
| Parameters | 70B (dense) |
| Context window | 128,000 tokens |
| License | Llama 3.1 License |
| HuggingFace | NousResearch/Hermes-3-Llama-3.1-70B |
| Release date | August 12, 2024 |
| Knowledge cutoff | March 2024 |
| Modality | Text only |
| Fine-tune type | Instruction + agentic |
Hermes 3 70B is available as open weights under the Llama 3.1 License — free to self-host. Available via third-party providers (Together AI, Replicate, etc.) at market rates.
Hermes 3 70B has a 128,000-token context window — the full Llama 3.1 70B context, suitable for long document analysis and multi-turn agentic tasks.
| Benchmark | Score | Source | Date |
|---|---|---|---|
| MMLU-Pro | 47.2% | Benchgen evaluation | 2025-07 |
| Model | MMLU-Pro | Context | License | Type |
|---|---|---|---|---|
| Hermes 3 70B | 47.2% | 128K | Llama 3.1 | Instruction FT |
| Llama 4 Maverick | 80.5% | 1M | Llama 4 | Native |
| Phi-4 | 70.4% | 16K | Apache 2.0 | Native |
| Qwen3 32B | — | 131K | Apache 2.0 | Native |
For new deployments requiring open-weight general capability, Llama 4 Maverick, Qwen3 32B, or Phi-4 provide significantly better benchmark performance. Hermes 3 70B is most relevant when the specific Hermes instruction-tuning style or existing Hermes-based production integrations are the primary consideration.
Specs from NousResearch's official Hermes 3 release (August 2024) and Benchgen evaluations. Last updated 2026-07-24.
This model isn’t on any benchmark leaderboard yet.