Skip to content

[Internal]: Refresh the LLM Performance Matrix for Elastic Security with September 2026 results #8548

Description

@dhru42

Description

[Docs] Refresh the LLM Performance Matrix for Elastic Security with the September 2026 evaluation results

Live page: Large language model performance matrix for Elastic Security
Source file: solutions/security/ai/large-language-model-performance-matrix.md

Summary

Replace the scores on the LLM performance matrix with the results of our latest evaluation run (September 2026). This is a data-only refresh: the page structure, the "How the scores are calculated" section, the sub-capability definitions, the ≤ 5 "not recommended" threshold, and the methodology link to Benchmarking the Agentic SOC all stay as they are. Only the two :::{table} blocks (Proprietary models, Open-source models) change.

The evaluation harness and scoring are unchanged from the current page: Overall Agent Builder Score is the mean of the seven Agent Builder sub-capabilities, and Overall Score is the mean of Agent Builder, Attack Discovery, and Automatic Migration. Both tables below are sorted by Overall Score descending, matching the page's default sort.

What changed since the published tables

Models added (12):

  • Proprietary: Anthropic Claude Opus 5, Anthropic Claude Opus 4.8, Anthropic Claude Sonnet 5, OpenAI GPT-5.5, OpenAI GPT-5.6 Terra, OpenAI GPT-5.6 Sol, OpenAI GPT-5.6 Luna, Google Gemini 3.6 Flash, Google Gemini 3.5 Flash Lite
  • Open-source: Z.ai GLM 5.3, Qwen 3.8 2.4T A95B

Models removed (2): OpenAI GPT-4.1 and OpenAI GPT-4.1 Mini were not part of this run and should be dropped from the Proprietary table.

Replacement tables (paste-ready)

Formatting matches the current page: model names, Overall Agent Builder Score, and Overall Score are bold; all other cells are plain.

Proprietary models

:::{table}
:matrix:

| **Model** | Agent Builder: Alert Analysis | Agent Builder: Entity Analytics | Agent Builder: Threat Hunting | Agent Builder: Detection Rules | Agent Builder: Workflow Authoring | Agent Builder: Triggering Workflows | Agent Builder: Multi-Step Executions | **Overall Agent Builder Score** | **Attack Discovery** | **Automatic Migration** | **Overall Score** |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| **Anthropic Claude Sonnet 4.5** | 9.00 | 8.00 | 7.00 | 8.00 | 4.00 | 9.00 | 8.00 | **7.57** | 9.40 | 9.61 | **8.86** |
| **Anthropic Claude Opus 4.7** | 5.00 | 8.00 | 8.00 | 8.00 | 6.00 | 9.00 | 8.00 | **7.43** | 9.10 | 9.90 | **8.81** |
| **Anthropic Claude Opus 4.6** | 9.00 | 8.00 | 8.00 | 8.00 | 4.00 | 8.00 | 8.00 | **7.57** | 9.20 | 9.61 | **8.79** |
| **Anthropic Claude Sonnet 4.6** | 9.00 | 8.00 | 8.00 | 8.00 | 6.00 | 9.00 | 8.00 | **8.00** | 9.00 | 9.23 | **8.74** |
| **Anthropic Claude Opus 5** | 6.00 | 8.00 | 8.00 | 8.00 | 9.00 | 9.00 | 6.00 | **7.71** | 9.60 | 8.75 | **8.69** |
| **Anthropic Claude Opus 4.8** | 7.00 | 7.00 | 8.00 | 8.00 | 7.00 | 9.00 | 6.00 | **7.43** | 8.00 | 10.00 | **8.48** |
| **Anthropic Claude Sonnet 5** | 8.00 | 8.00 | 7.00 | 8.00 | 7.00 | 9.00 | 5.00 | **7.43** | 9.00 | 8.84 | **8.42** |
| **OpenAI GPT-5.2** | 6.00 | 7.00 | 8.00 | 8.00 | 6.00 | 8.00 | 8.00 | **7.29** | 8.00 | 9.60 | **8.30** |
| **OpenAI GPT-5.5** | 8.00 | 8.00 | 8.00 | 8.00 | 9.00 | 8.00 | 7.00 | **8.00** | 8.20 | 8.55 | **8.25** |
| **Anthropic Claude Opus 4.5** | 8.00 | 8.00 | 8.00 | 8.00 | 6.00 | 9.00 | 8.00 | **7.86** | 8.70 | 8.17 | **8.24** |
| **Google Gemini 2.5 Pro** | 5.00 | 5.00 | 7.00 | 6.00 | 6.00 | 9.00 | 8.00 | **6.57** | 8.70 | 9.32 | **8.20** |
| **OpenAI GPT-5.6 Terra** | 7.00 | 7.00 | 7.00 | 8.00 | 5.00 | 8.00 | 8.00 | **7.14** | 6.50 | 9.61 | **7.75** |
| **OpenAI GPT-5.4** | 7.00 | 7.00 | 8.00 | 7.00 | 7.00 | 9.00 | 8.00 | **7.57** | 5.30 | 9.84 | **7.57** |
| **Google Gemini 3.6 Flash** | 8.00 | 5.00 | 7.00 | 8.00 | 7.00 | 8.00 | 8.00 | **7.29** | 8.20 | 7.21 | **7.57** |
| **OpenAI GPT-5.6 Sol** | 6.00 | 8.00 | 7.00 | 8.00 | 9.00 | 8.00 | 6.00 | **7.43** | 7.20 | 8.01 | **7.55** |
| **Anthropic Claude Haiku 4.5** | 3.00 | 8.00 | 7.00 | 6.00 | 7.00 | 9.00 | 8.00 | **6.86** | 7.00 | 8.65 | **7.50** |
| **Google Gemini 3.0 Flash** | 7.00 | 6.00 | 7.00 | 8.00 | 3.00 | 8.00 | 8.00 | **6.71** | 6.00 | 9.71 | **7.47** |
| **OpenAI GPT-5.6 Luna** | 8.00 | 7.00 | 6.00 | 8.00 | 9.00 | 8.00 | 8.00 | **7.71** | 6.30 | 7.88 | **7.30** |
| **Google Gemini 2.5 Flash** | 3.00 | 4.00 | 5.00 | 3.00 | 3.00 | 6.00 | 6.00 | **4.29** | 6.80 | 9.61 | **6.90** |
| **Google Gemini 3.5 Flash** | 8.00 | 6.00 | 7.00 | 7.00 | 9.00 | 8.00 | 8.00 | **7.57** | 6.30 | 6.73 | **6.87** |
| **Google Gemini 3.1 Flash Lite** | 8.00 | 7.00 | 7.00 | 8.00 | 9.00 | 8.00 | 6.00 | **7.57** | 2.50 | 9.51 | **6.53** |
| **OpenAI GPT-5.4 Nano** | 4.00 | 4.00 | 5.00 | 4.00 | 9.00 | 7.00 | 6.00 | **5.57** | 3.50 | 8.75 | **5.94** |
| **Google Gemini 3.1 Pro (Preview)** | 8.00 | 5.00 | 7.00 | 8.00 | 7.00 | 9.00 | 7.00 | **7.29** | 3.20 | 6.44 | **5.64** |
| **OpenAI GPT-5.4 Mini** | 4.00 | 6.00 | 6.00 | 5.00 | 6.00 | 9.00 | 8.00 | **6.29** | 1.00 | 9.51 | **5.60** |
| **Google Gemini 3.5 Flash Lite** | 5.00 | 3.00 | 6.00 | 6.00 | 7.00 | 8.00 | 8.00 | **6.14** | 0.00 | 9.81 | **5.32** |
| **Google Gemini 2.5 Flash Lite** | 5.00 | 2.00 | 1.00 | 2.00 | 0.00 | 8.00 | 5.00 | **3.29** | 0.00 | 6.15 | **3.15** |

:::

Open-source models

:::{table}
:matrix:

| **Model** | Agent Builder: Alert Analysis | Agent Builder: Entity Analytics | Agent Builder: Threat Hunting | Agent Builder: Detection Rules | Agent Builder: Workflow Authoring | Agent Builder: Triggering Workflows | Agent Builder: Multi-Step Executions | **Overall Agent Builder Score** | **Attack Discovery** | **Automatic Migration** | **Overall Score** |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| **Z.ai GLM 5.3** | 8.00 | 7.00 | 7.00 | 8.00 | 9.00 | 9.00 | 8.00 | **8.00** | 8.00 | 7.30 | **7.77** |
| **Gemma 4 31B IT** | 6.00 | 7.00 | 7.00 | 8.00 | 3.00 | 9.00 | 8.00 | **6.86** | 4.30 | 9.61 | **6.92** |
| **Kimi K2.6** | 8.00 | 6.00 | 7.00 | 8.00 | 9.00 | 9.00 | 8.00 | **7.86** | 8.00 | 2.31 | **6.05** |
| **DeepSeek V4 Pro** | 6.00 | 8.00 | 7.00 | 8.00 | 9.00 | 7.00 | 8.00 | **7.57** | 3.50 | 6.25 | **5.77** |
| **Qwen 3.8 2.4T A95B** | 8.00 | 7.00 | 7.00 | 9.00 | 9.00 | 9.00 | 9.00 | **8.29** | 0.00 | 7.88 | **5.39** |
| **OpenAI GPT-OSS 120B** | 1.00 | 1.00 | 1.00 | 3.00 | 5.00 | 8.00 | 1.00 | **2.86** | 2.00 | 9.51 | **4.79** |
| **Qwen 3.6 27B** | 6.00 | 7.00 | 8.00 | 8.00 | 7.00 | 9.00 | 7.00 | **7.43** | 0.00 | 5.67 | **4.37** |
| **OpenAI GPT-OSS 20B** | 2.00 | 2.00 | 2.00 | 5.00 | 4.00 | 7.00 | 5.00 | **3.86** | 1.00 | 6.05 | **3.64** |

:::

Notes for the writer

  • Naming: the source spreadsheet lists "Google Gemini 3.5 Flash-Lite" with a hyphen; the tables above normalize it to Google Gemini 3.5 Flash Lite to match the existing "Gemini 2.5 Flash Lite" and "Gemini 3.1 Flash Lite" rows. Keep "Google Gemini 3.1 Pro (Preview)" as is.
  • Rounding: Overall Agent Builder Score and Overall Score are shown to two decimals as produced by the evaluation spreadsheet; do not recompute them from the displayed sub-scores, which are rounded independently.
  • No changes are needed to the intro, the {important} callout, the "How the scores are calculated" section, the sub-capability definitions, or the section anchors.

Resources

Which deployment methods does this change impact?

Elastic On-Prem and Cloud (all)

Feature differences

The matrix applies identically to all deployment methods; the page's existing applies_to (stack: all, serverless security: all) is unchanged.

What Elastic Stack release is this request related to?

N/A

Serverless release

No response

Collaboration model

The product or engineering team will create the first draft

Point of contact.

Main contact: @dhru42

Stakeholders: @jamesspi

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Team:SKIIssues owned by the SKI Docs TeamdocumentationImprovements or additions to documentationtriagedTriage has been run on this issue

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions