Benchgen
Models/nousresearch/

Hermes 3 70B

DraftPublic

Model Details

Hermes 3 70B

Organization Parameters License Weights Released

Quick answer: Hermes 3 70B is NousResearch's August 2024 instruction fine-tune of Meta's Llama 3.1 70B. It scores 47.2% on MMLU-Pro. Under the Llama 3.1 License with open weights on Hugging Face, it is designed for strong general reasoning, long-context tasks, and agentic deployment patterns.

At a Glance

Where Hermes 3 70B leads

  • NousResearch instruction-tuning — improved function calling and tool use over base Llama 3.1 70B
  • 128K context window — full Llama 3.1 70B context
  • Open weights — self-hostable for agentic pipelines
  • Strong at complex system prompts and multi-turn conversation
  • Agent-optimised: designed explicitly for agentic and tool-using deployment

Where it lags

  • 47.2% MMLU-Pro — below leading 2025 open-weight models (Qwen3 32B, Llama 4 Maverick)
  • Base Llama 3.1 70B from August 2024 — older than Llama 4, Qwen3, Phi-4
  • Knowledge cutoff March 2024 — outdated
  • No reasoning/thinking mode

Best for: Agentic and tool-use pipelines requiring open-weight models; teams specifically using NousResearch's Hermes instruction style; legacy Hermes-based deployments.

What Hermes 3 70B Is

Hermes 3 70B is NousResearch's third generation Hermes fine-tune, built on Meta's Llama 3.1 70B base. Released August 12, 2024, it specialises in advanced instruction following, function calling, and agentic capability — areas where the base Llama 3.1 70B was strong but could be improved through targeted fine-tuning.

The Hermes series is known in the open-source community for producing models with reliable tool use, structured output generation, and consistent behaviour with complex system prompts. Hermes 3 70B brings these properties to the Llama 3.1 70B scale.

Its 47.2% MMLU-Pro score reflects its 2024 vintage and 70B base — newer open-weight models like Llama 4 Maverick (80.5%), Qwen3 32B (no MMLU-Pro score but strong AetherCode), and Phi-4 (70.4% at only 14B) have since surpassed it on academic knowledge benchmarks. Hermes 3 70B remains relevant for use cases that specifically benefit from NousResearch's instruction-tuning approach and are already integrated with Hermes-style prompting.

Specifications

FieldValue
OrganizationNousResearch
Base modelMeta Llama 3.1 70B
Parameters70B (dense)
Context window128,000 tokens
LicenseLlama 3.1 License
HuggingFaceNousResearch/Hermes-3-Llama-3.1-70B
Release dateAugust 12, 2024
Knowledge cutoffMarch 2024
ModalityText only
Fine-tune typeInstruction + agentic

Pricing

Hermes 3 70B is available as open weights under the Llama 3.1 License — free to self-host. Available via third-party providers (Together AI, Replicate, etc.) at market rates.

Context Window

Hermes 3 70B has a 128,000-token context window — the full Llama 3.1 70B context, suitable for long document analysis and multi-turn agentic tasks.

Public Benchmark Scores

BenchmarkScoreSourceDate
MMLU-Pro47.2%Benchgen evaluation2025-07

Hermes 3 70B vs Alternatives

ModelMMLU-ProContextLicenseType
Hermes 3 70B47.2%128KLlama 3.1Instruction FT
Llama 4 Maverick80.5%1MLlama 4Native
Phi-470.4%16KApache 2.0Native
Qwen3 32B131KApache 2.0Native

For new deployments requiring open-weight general capability, Llama 4 Maverick, Qwen3 32B, or Phi-4 provide significantly better benchmark performance. Hermes 3 70B is most relevant when the specific Hermes instruction-tuning style or existing Hermes-based production integrations are the primary consideration.

Frequently Asked Questions

What is Hermes 3 70B? Hermes 3 70B is NousResearch's August 2024 instruction fine-tune of Llama 3.1 70B, scoring 47.2% MMLU-Pro. It is optimised for agentic tasks, function calling, and complex instruction following.
Is Hermes 3 70B open source? Yes. Hermes 3 70B is available on Hugging Face under the Llama 3.1 License — permitting commercial use within the Llama 3.1 license terms.
What is the difference between Hermes 3 and the base Llama 3.1? Hermes 3 adds NousResearch's instruction fine-tuning on top of Llama 3.1 70B, improving function calling, tool use, structured output generation, and complex system prompt handling beyond the base model.

Specs from NousResearch's official Hermes 3 release (August 2024) and Benchgen evaluations. Last updated 2026-07-24.

Benchmark Leaderboards

This model isn’t on any benchmark leaderboard yet.