HermesIndex

An opinionated index of agentic capability, run inside Hermes Agent and reporting task completion and cost per task across four suites.

Top 5 // Hermes IndexAvg score · avg $ per task
  1. [ 1 ]Claude Opus 5.5Anthropic63.31$4.99*
    x
  2. [ 2 ]GPT 6 AstraOpenAI56.25$11.61
    x
  3. [ 3 ]Claude Sonnet 5.5Anthropic53.14$2.82*
    x
  4. [ 4 ]GPT 6 SolOpenAI44.10$2.23
    x
  5. [ 5 ]Grok 4.7xAI39.32$10.77
    x

Leaderboard

Hermes Index is the mean score and mean cost per task across the four suites. Higher scores are better.

#
Model
01
Claude Opus 5.5Anthropic
63.31$4.99*
x
76.73$0.516
54.0$6.27*
52.9$12.65
69.6$0.51
02
GPT 6 AstraOpenAI
56.25$11.61
x
73.67$1.52
41.3$12.66
57.1$29.84
52.9$2.44
03
Claude Sonnet 5.5Anthropic
53.14$2.82*
x
73.46$0.25
36.5$2.77*
35.7$8.02
66.9$0.23
04
GPT 6 SolOpenAI
44.10$2.23
x
73.50$0.263
23.8$2.53
18.6$5.70
60.5$0.42
05
Grok 4.7xAI
39.32$10.77
x
74.38$0.69
12.7$23.00
1.4$18.41
68.8$0.99 · 71.9 raw
06
GLM 5.3Z.ai
39.23$3.78
x
71.82$0.18
38.1$7.68
1.4$6.82
45.6$0.44
07
Kimi K3Moonshot
39.16$6.50
x
71.52$0.176
19.0$8.57
8.6$16.74
57.5$0.53
08
Gemini Flash 3.8Google
38.12$3.17
x
62.19$0.64
23.8$5.94
8.6$5.15
57.9$0.94 · 67.7 raw
09
DeepSeek V4.1 FlashDeepSeek
36.91$0.259
x
70.74$0.019
17.5$0.62
2.9$0.38
56.5$0.016 · 61.1 raw
10
Qwen 3.8 MaxAlibaba
36.17$6.69
x
70.50$0.322
17.5$6.84
5.7$19.09
51.0$0.52
11
GLM 5.3 FlashZ.ai
34.95$0.497
x
68.88$0.019
25.4$0.60
2.9$1.31
42.6$0.058
12
GPT 6 LunaOpenAI
33.89$0.141
x
70.47$0.015
6.3$0.24
4.8$0.28
54.0$0.03
13
Hy4Tencent
26.13$0.482
x
58.60$0.09
12.7$0.70
1.4$0.98
31.8$0.16
14
Ling 3.0 FlashinclusionAI
21.56$0.054
x
53.83$0.0036
0.0$0.08
0.0$0.12
32.4$0.013
  • All evaluations are pass@1. Every run uses the Hermes Agent harness, with reasoning effort set to high where the model offers it.
  • TerminalBench 4 runs without its 4 GPU tasks.
  • * Provisional: partial run or cost estimate, pending a full rerun.
  • Clean / raw: some SkillsBench runs found the suite's public repo and used its solutions. The clean score drops those tasks and is the one that counts; the raw score sits beneath it.

Compare

Pick up to four models to see their scores and costs on each suite.

Models4 / 4 · remove one to swap
Capability profileScore · 0–100
Cost per taskUSD · log scale
Hermes BenchTB 4TB SciSkills$0.003$0.01$0.03$0.1$0.3$1$3$10$30
Head to head
Suite
Claude Opus 5.5
GPT 6 Sol
GLM 5.3
DeepSeek V4.1 Flash
Hermes IndexHermes IndexMean of the four suites
63.31$4.99*
44.10$2.23
39.23$3.78
36.91$0.259
Hermes BenchHermes Bench
76.73$0.516
73.50$0.263
71.82$0.18
70.74$0.019
TerminalBench 4TB 4
54.0$6.27*
23.8$2.53
38.1$7.68
17.5$0.62
TerminalBench ScienceTB Sci
52.9$12.65
18.6$5.70
1.4$6.82
2.9$0.38
SkillsBenchSkills
69.6$0.51
60.5$0.42
45.6$0.44
56.5$0.016 · 61.1 raw
  • All evaluations are pass@1. Every run uses the Hermes Agent harness, with reasoning effort set to high where the model offers it.
  • TerminalBench 4 runs without its 4 GPU tasks.
  • * Provisional: partial run or cost estimate, pending a full rerun.
  • Clean / raw: some SkillsBench runs found the suite's public repo and used its solutions. The clean score drops those tasks and is the one that counts; the raw score sits beneath it.

Frontier

Hermes Index score against average cost per task. Each model on the line scores higher than every model that costs less.

Pareto optimal
  1. Claude Opus 5.5$4.99 / task63.31
  2. Claude Sonnet 5.5$2.82 / task53.14
  3. GPT 6 Sol$2.23 / task44.10
  4. DeepSeek V4.1 Flash$0.259 / task36.91
  5. GPT 6 Luna$0.141 / task33.89
  6. Ling 3.0 Flash$0.054 / task21.56

Hermes Bench

150 tasks across 25 categories: Hermes skills, research, diagrams and art, memory, tool use and safety. Each agent starts in a workspace with real files, some tasks add follow-up turns, and the grader checks the files and state left behind.

Grading modes150 tasks
Automated64
Deterministic checks on files, state and tool evidence.
Hybrid77
Deterministic checks plus an LLM judge on the rubric.
LLM judge1
Rubric-only tasks scored by a judge model.
Vision judge8
Rendered diagrams, judged three times, median kept.
  • 87

    Hermes Skills

    16 skill families: job search, email triage, diagrams, fitness, sports and energy analysis, maps, code review, wikis and arXiv.
  • 43

    Research

    Reconciliation, conflicting sources, scheduling, procurement, grounding and preserving existing work.
  • 21

    Visual + Creative

    Architecture and Excalidraw diagrams, p5.js sketches and ASCII art.
  • 5

    Memory

    Sweeping, merging and correcting stored memories, and holding back when the context isn't there.
  • 5

    Tool Use

    Browser work through the tool gateway: route mazes, section navigation, table diffs and visual extraction.
  • 1

    Safety

    Refusing destructive Git: force-pushes, hard resets and branch confusion.
Seeded workspaces
88
Workspace files
293
Multi-turn tasks
24
Scripted follow-up turns
81

The Internet's Own AI

© 2026, Nous Research, Inc.

Terms|Privacy

MIT License · 2026