Spring til indhold
AiNews.dkDanske nyheder om Kunstig Intelligens
ModelradarenHentet 1. oktober 2026 kl. 18.30

Intelligens, under lup

Top 100
AI-modeller

Fra GPT og Claude til åbne modeller.
Se resultaterne. Følg udviklingen.

Ranglisten

Gennemsnitlig benchmarkpercentil
Modeller sorteret efter normaliseret benchmarkscore
Nr.ModelGrundlagScore
01
GPT-6-astraOpenAILukket model · 21 tests
21 benchmarksUændret96,0
02
Claude Fable-5.1AnthropicLukket model · 20 tests
20 benchmarksUændret94,9
03
Claude Opus-5.5AnthropicLukket model · 18 tests
18 benchmarksUændret93,9
04
GPT-6.1-solOpenAILukket model · 13 tests
13 benchmarksUændret92,7
05
Claude Opus-5AnthropicLukket model · 22 tests
22 benchmarksUændret88,9
06
Claude Sonnet-5.5AnthropicLukket model · 11 tests
11 benchmarksUændret87,9
07
Claude Fable-5AnthropicLukket model · 23 tests
23 benchmarksUændret87,4
08
GPT-6-solOpenAILukket model · 18 tests
18 benchmarksUændret86,6
09
GPT-5.6-solOpenAILukket model · 25 tests
25 benchmarksUændret84,9
10
Muse Spark 1.3MetaLukket model · 12 tests
12 benchmarksUændret81,4
11
Qwen3.8 2.4T A95BAlibabaÅbne vægte · 7 tests
7 benchmarksUændret80,7
12
mimo-v2.6-proXiaomiÅbne vægte · 9 tests
9 benchmarksUændret77,9
13
GPT-5.6-terraOpenAILukket model · 22 tests
22 benchmarksUændret77,8
14
Qwen3.8 Max 0902AlibabaLukket model · 10 tests
10 benchmarksUændret76,9
15
Gemini 3.8 FlashGoogleLukket model · 18 tests
18 benchmarksUændret76,8
16
Qwen3.8 MaxAlibabaLukket model · 8 tests
8 benchmarksUændret76,6
17
GPT-5.5OpenAILukket model · 24 tests
24 benchmarksUændret76,2
18
kimi-k3KimiÅbne vægte · 21 tests
21 benchmarksUændret75,7
19
GLM-5.3Z.aiÅbne vægte · 14 tests
14 benchmarksUændret74,6
20
muse-spark-1.1MetaLukket model · 12 tests
12 benchmarksUændret74,3
21
gemini-3.7-flashGoogleLukket model · 17 tests
17 benchmarksUændret74,2
22
Claude Opus-4.8AnthropicLukket model · 22 tests
22 benchmarksUændret73,8
23
Qwen3.8 FlashAlibabaÅbne vægte · 7 tests
7 benchmarksUændret72,1
24
grok-4.7xAILukket model · 9 tests
9 benchmarksUændret70,8
25
Agnes 2.5 Pro BetaWorld (Other)Lukket model · 6 tests
6 benchmarksUændret68,7
26
grok-4.6xAILukket model · 15 tests
15 benchmarksUændret68,7
27
Claude Sonnet-5AnthropicLukket model · 18 tests
18 benchmarksUændret68,5
28
qwen3.7-maxAlibabaLukket model · 12 tests
12 benchmarksUændret67,2
29
muse-sparkMetaLukket model · 11 tests
11 benchmarksUændret66,2
30
DeepSeek V4.1 FlashDeepSeekÅbne vægte · 10 tests
10 benchmarksUændret66,2
31
gemini-3.1-proGoogleLukket model · 25 tests
25 benchmarksUændret66,0
32
GPT-5.3-codexOpenAILukket model · 11 tests
11 benchmarksUændret65,5
33
GPT-5.4OpenAILukket model · 23 tests
23 benchmarksUændret65,0
34
deepseek-v4-flash-vision-expDeepSeekLukket model · 7 tests
7 benchmarksUændret64,7
35
GPT-5.2OpenAILukket model · 19 tests
19 benchmarksUændret64,2
36
deepseek-v4-proDeepSeekÅbne vægte · 21 tests
21 benchmarksUændret64,0
37
Muse Spark 1.2MetaLukket model · 10 tests
10 benchmarksUændret63,5
38
GPT-5.6-lunaOpenAILukket model · 20 tests
20 benchmarksUændret63,5
39
GPT-6-lunaOpenAILukket model · 17 tests
17 benchmarksUændret63,2
40
gemini-3.5-flashGoogleLukket model · 20 tests
20 benchmarksUændret62,6
41
gemini-3.6-flashGoogleLukket model · 18 tests
18 benchmarksUændret62,6
42
grok-4.5xAILukket model · 21 tests
21 benchmarksUændret62,2
43
Claude Opus-4.7AnthropicLukket model · 23 tests
23 benchmarksUændret61,6
44
minimax-m3MiniMaxÅbne vægte · 13 tests
13 benchmarksUændret61,2
45
gemini-3-proGoogleLukket model · 17 tests
17 benchmarksUændret61,0
46
mimo-v2.6-flashXiaomiÅbne vægte · 10 tests
10 benchmarksUændret59,7
47
Claude Opus-4.6AnthropicLukket model · 22 tests
22 benchmarksUændret59,7
48
nex-n2-proChina (Other)Åbne vægte · 7 tests
7 benchmarksUændret59,3
49
GLM-5.3-FlashZ.aiÅbne vægte · 14 tests
14 benchmarksUændret57,7
50
glm-5.2Z.aiÅbne vægte · 18 tests
18 benchmarksUændret57,4
51
qwen3.6-maxAlibabaLukket model · 8 tests
8 benchmarksUændret57,3
52
qwen3.7-plusAlibabaLukket model · 10 tests
10 benchmarksUændret56,6
53
kimi-k2.6KimiÅbne vægte · 19 tests
19 benchmarksUændret56,4
54
deepseek-v4-flashDeepSeekÅbne vægte · 19 tests
19 benchmarksUændret56,0
55
mimo-v2.5-proXiaomiÅbne vægte · 10 tests
10 benchmarksUændret54,3
56
gemini-3-flashGoogleLukket model · 19 tests
19 benchmarksUændret53,9
57
grok-4.20xAILukket model · 14 tests
14 benchmarksUændret52,6
58
Qwen3.8 27BAlibabaÅbne vægte · 13 tests
13 benchmarksUændret51,3
59
kimi-k2.5KimiÅbne vægte · 21 tests
21 benchmarksUændret50,0
60
Claude Opus-4.5AnthropicLukket model · 17 tests
17 benchmarksUændret49,9
61
qwen3.5-397bAlibabaÅbne vægte · 14 tests
14 benchmarksUændret49,4
62
kimi-k2.7-codeKimiÅbne vægte · 11 tests
11 benchmarksUændret48,3
63
Solar Pro 4World (Other)Lukket model · 6 tests
6 benchmarksUændret48,2
64
grok-4.3xAILukket model · 15 tests
15 benchmarksUændret47,8
65
Claude Sonnet-4.6AnthropicLukket model · 20 tests
20 benchmarksUændret46,9
66
glm-5Z.aiÅbne vægte · 18 tests
18 benchmarksUændret46,7
67
GPT-5OpenAILukket model · 17 tests
17 benchmarksUændret46,4
68
qwen3.6-plusAlibabaLukket model · 13 tests
13 benchmarksUændret45,5
69
glm-5.1Z.aiÅbne vægte · 17 tests
17 benchmarksUændret44,8
70
Muse GlimmerMetaÅbne vægte · 8 tests
8 benchmarksUændret44,0
71
Inkling-SmallUS (Other)Åbne vægte · 18 tests
18 benchmarksUændret43,4
72
nemotron-3-ultraUS (Other)Åbne vægte · 11 tests
11 benchmarksUændret43,3
73
GPT-5.1OpenAILukket model · 17 tests
17 benchmarksUændret43,1
74
hy3China (Other)Åbne vægte · 7 tests
7 benchmarksUændret42,5
75
GPT-5.4-miniOpenAILukket model · 18 tests
18 benchmarksUændret42,4
76
minimax-m2.1MiniMaxÅbne vægte · 9 tests
9 benchmarksUændret42,2
77
mimo-v2-proXiaomiÅbne vægte · 6 tests
6 benchmarksUændret40,6
78
step-3.7-flashChina (Other)Åbne vægte · 8 tests
8 benchmarksUændret39,6
79
ring-2.6-1tChina (Other)Åbne vægte · 7 tests
7 benchmarksUændret37,5
80
gemma-4-31bGoogleÅbne vægte · 12 tests
12 benchmarksUændret37,4
81
grok-4xAILukket model · 16 tests
16 benchmarksUændret37,1
82
gemma-4-26bGoogleÅbne vægte · 9 tests
9 benchmarksUændret37,0
83
inklingUS (Other)Åbne vægte · 16 tests
16 benchmarksUændret36,9
84
minimax-m2.7MiniMaxÅbne vægte · 11 tests
11 benchmarksUændret36,9
85
gemini-3.5-flash-liteGoogleLukket model · 15 tests
15 benchmarksUændret36,4
86
GPT-5.4-nanoOpenAILukket model · 16 tests
16 benchmarksUændret36,2
87
nova-2.0-proUS (Other)Lukket model · 9 tests
9 benchmarksUændret36,1
88
mimo-v2.5XiaomiÅbne vægte · 11 tests
11 benchmarksUændret35,9
89
Ling 3.0 FlashChina (Other)Åbne vægte · 7 tests
7 benchmarksUændret35,6
90
glm-4.7Z.aiÅbne vægte · 15 tests
15 benchmarksUændret35,5
91
o3OpenAILukket model · 18 tests
18 benchmarksUændret35,5
92
gemini-3.1-flash-liteGoogleLukket model · 13 tests
13 benchmarksUændret35,4
93
nemotron-3-superUS (Other)Åbne vægte · 9 tests
9 benchmarksUændret35,4
94
minimax-m2.5MiniMaxÅbne vægte · 11 tests
11 benchmarksUændret35,0
95
qwen3.6-27bAlibabaÅbne vægte · 10 tests
10 benchmarksUændret34,9
96
k-exaoneWorld (Other)Åbne vægte · 8 tests
8 benchmarksUændret33,7
97
kat-coder-pro-v1China (Other)Lukket model · 8 tests
8 benchmarksUændret33,5
98
mimo-v2-flashXiaomiÅbne vægte · 8 tests
8 benchmarksUændret33,5
99
gemini-2.5-proGoogleLukket model · 16 tests
16 benchmarksUændret32,3
100
deepseek-v3.2DeepSeekÅbne vægte · 15 tests
15 benchmarksUændret30,3
100 rangerede · 367 modeller med resultaterOm tallene ↓
Modeller med begrænset dokumentation 230

Disse modeller har endnu ikke seks kvalificerede benchmarks. Se deres tal og kilder uden en samlet placering.

Fugu Ultra v1.11 kvalificerede testsFugu Max v1.01 kvalificerede testsFugu Ultra v2.01 kvalificerede testsFugu Ultra v1.04 kvalificerede testsFugu v1.04 kvalificerede testsQwen3.5-122B-A10B1 kvalificerede testsIntern-S11 kvalificerede testsStep-3.5-Flash2 kvalificerede testsQwen3-235B-A22B-Instruct-25071 kvalificerede testsSeed-OSS-36B-Instruct1 kvalificerede testsLongCat-Flash-Chat1 kvalificerede testsQwen3.5-397B-A17B2 kvalificerede testsMiniMax-M1-40k1 kvalificerede testsKimi-K2-Thinking1 kvalificerede testsMiniMax-M2.51 kvalificerede testsERNIE-4.5-300B-A47B-PT1 kvalificerede testsHy4 Preview3 kvalificerede testsMiniMax-Text-011 kvalificerede testsQwen2.5-72B1 kvalificerede testsphi-41 kvalificerede testsERNIE-4.5-300B-A47B-Base-PT1 kvalificerede testsQwen2.5-32B1 kvalificerede testsGLM-4.7-FP81 kvalificerede testsQwen3-235B-A22B1 kvalificerede testsMistral-Large-Instruct-24111 kvalificerede testsHunyuan-A13B-Instruct1 kvalificerede testsMistral-Large-Instruct-24071 kvalificerede testsDeepSeek-V2.51 kvalificerede testsSeed-OSS-36B-Base1 kvalificerede testsNVIDIA-Nemotron-3-Nano-30B-A3B-Base-BF161 kvalificerede testsMiniMax-M2.13 kvalificerede testsQwen2.5-14B1 kvalificerede testsQwen3-30B-A3B-Base1 kvalificerede testsLlama-3.1-405B1 kvalificerede testsNemotron-H-56B-Base-8K1 kvalificerede testsQwen3.5-27B3 kvalificerede testsSeed-OSS-36B-Base-woSyn1 kvalificerede testsTencent-Hunyuan-Large1 kvalificerede testsEXAONE-3.5-32B-Instruct1 kvalificerede testsMiMo-7B-RL1 kvalificerede testsKimi-K2-Instruct2 kvalificerede testsinternlm3-8b-instruct1 kvalificerede testsMiniMax-M21 kvalificerede testsGLM-5-FP81 kvalificerede testsERNIE-4.5-21B-A3B-Base-PT1 kvalificerede testsPhi-3-medium-4k-instruct1 kvalificerede testsGLM-4.52 kvalificerede testsDeepSeek-V2-Chat1 kvalificerede testsMistral-Small-24B-Base-25011 kvalificerede testsPhi-4-mini-instruct1 kvalificerede testsMeta-Llama-3-70B1 kvalificerede testsLlama-3.1-70B1 kvalificerede testsYi-1.5-34B-Chat1 kvalificerede testsPhi-3-medium-128k-instruct1 kvalificerede testsMAmmoTH2-8x7B-Plus1 kvalificerede testsMiMo-V2-Flash1 kvalificerede testsGLM-5.1-FP81 kvalificerede testsQwen1.5-110B1 kvalificerede testsAI21-Jamba-Large-1.51 kvalificerede testsMistral-Small-Instruct-24091 kvalificerede testsglm-4-9b1 kvalificerede testsPhi-3.5-mini-instruct1 kvalificerede testsEXAONE-3.5-7.8B-Instruct1 kvalificerede testsGLM-4.6-FP81 kvalificerede testsYi-1.5-9B-Chat1 kvalificerede testsPhi-3-mini-4k-instruct1 kvalificerede testsaya-expanse-32b1 kvalificerede testsgemma-2-9b1 kvalificerede testsNVIDIA-Nemotron-3-Super-120B-A12B-BF162 kvalificerede testsGLM-4.5-Air3 kvalificerede testsQwen2.5-7B1 kvalificerede testsPhi-3-mini-128k-instruct1 kvalificerede testsQwen2.5-3B1 kvalificerede testsMAmmoTH2-8B-Plus1 kvalificerede testsYi-34B1 kvalificerede testsDeepSeek-R1-05281 kvalificerede testsMathstral-7B-v0.11 kvalificerede testsMiMo-7B-Base1 kvalificerede testsQwen3-30B-A3B-Thinking-25074 kvalificerede testsDeepSeek-Coder-V2-Lite-Instruct1 kvalificerede testsMixtral-8x7B-v0.11 kvalificerede testsMeta-Llama-3-8B-Instruct1 kvalificerede testsMAmmoTH2-7B-Plus1 kvalificerede testsNVIDIA-Nemotron-3-Super-120B-A12B-FP81 kvalificerede testsQwen2-7B1 kvalificerede testsMistral-Nemo-Base-24071 kvalificerede testsEXAONE-3.5-2.4B-Instruct1 kvalificerede testsYi-1.5-6B-Chat1 kvalificerede testsQwen1.5-14B-Chat1 kvalificerede testsMinistral-8B-Instruct-24101 kvalificerede testsc4ai-command-r-v011 kvalificerede testsinternlm2-math-plus-20b1 kvalificerede testsLLaDA-8B-Instruct1 kvalificerede testsLlama-3-Smaug-8B1 kvalificerede testsLlama-3.1-8B1 kvalificerede testsMeta-Llama-3-8B1 kvalificerede testsdeepseek-math-7b-instruct1 kvalificerede testsDeepSeek-Coder-V2-Lite-Base1 kvalificerede testsaya-expanse-8b1 kvalificerede testsgrok-25 kvalificerede testsGranite 4.2 30B5 kvalificerede testsinternlm2-math-plus-7b1 kvalificerede testsgranite-3.1-8b-base1 kvalificerede testsGPT-4-turbo5 kvalificerede testsQwen2.5-1.5B1 kvalificerede testsClaude Opus-35 kvalificerede testsgranite-3.0-8b-base1 kvalificerede testsMistral-7B-Instruct-v0.21 kvalificerede testsMistral-7B-v0.21 kvalificerede testsGLM-4.61 kvalificerede testsQwen3.7 Flash1 kvalificerede testsQwen1.5-7B-Chat1 kvalificerede testslaguna-xs.23 kvalificerede testsYi-6B-Chat1 kvalificerede testslaguna-m.12 kvalificerede testsYi-6B1 kvalificerede testsDeepSeek-V31 kvalificerede testsgranite-3.1-2b-base1 kvalificerede testsllemma_7b1 kvalificerede testsQwen2-1.5B-Instruct1 kvalificerede testsQwen2-1.5B1 kvalificerede testsgemini-2.0-pro4 kvalificerede testsLlama-3.2-3B1 kvalificerede testsgranite-3.0-2b-base1 kvalificerede testsc4ai-command-a-03-20251 kvalificerede testsgranite-3.1-3b-a800m-base1 kvalificerede testsSmolLM2-1.7B1 kvalificerede testsLlama-3.2-90B-Vision-Instruct1 kvalificerede testsQwen2-0.5B1 kvalificerede testsQwen2.5-0.5B1 kvalificerede testsgranite-3.1-1b-a400m-base1 kvalificerede testsLlama-3.2-1B1 kvalificerede testsc4ai-command-r-plus-08-20241 kvalificerede testsSmolLM2-360M1 kvalificerede testsc4ai-command-r-08-20241 kvalificerede testsc4ai-command-r7b-12-20241 kvalificerede testsSmolLM2-135M1 kvalificerede testsQwen3-4B-Thinking-25072 kvalificerede testsLlama-3.2-1B-Instruct1 kvalificerede testsQwen2.5-VL-72B-Instruct1 kvalificerede testsQwen3-Coder-480B-A35B-Instruct1 kvalificerede testsNorth-Mini-Code-1.00 kvalificerede testsDeepSeek-V20 kvalificerede testsDeepSeek-V3-03240 kvalificerede testsdots3-note-prev0 kvalificerede testscwm0 kvalificerede testsDarwin-180B-RSI0 kvalificerede testsDarwin-27B-Opus0 kvalificerede testsDarwin-28B-REASON0 kvalificerede testsDarwin-31B-Opus0 kvalificerede testsDarwin-36B-Opus0 kvalificerede testsDarwin-397B-ZTC0 kvalificerede testsDarwin-398B-JGOS0 kvalificerede testsDarwin-4B-David0 kvalificerede testsDarwin-60B-DUO0 kvalificerede testsDarwin-9B-NEG0 kvalificerede testsOurbox-35B-JGOS0 kvalificerede testsDhanishtha-2.0-01260 kvalificerede testsgranite-4.1-30b0 kvalificerede testsgranite-4.1-3b0 kvalificerede testsgranite-4.1-8b0 kvalificerede testsLing-2.6-1T0 kvalificerede testsLing-2.6-flash0 kvalificerede testsLing-3.0-flash0 kvalificerede testsRing-2.6-1T0 kvalificerede testsGLM-5.1-MXFP4-Mixed-CT-AutoRound0 kvalificerede testsIntern-S2-Preview0 kvalificerede testsinternlm2_5-7b-chat0 kvalificerede testsinternlm2-chat-20b0 kvalificerede testsAgents-A10 kvalificerede testsJoyAI-LLM-Flash0 kvalificerede testsMellum2-12B-A2.5B-Base0 kvalificerede testsMellum2-12B-A2.5B-Base-Pretrain0 kvalificerede testsMellum2-12B-A2.5B-Instruct0 kvalificerede testsMellum2-12B-A2.5B-Instruct-SFT0 kvalificerede testsMellum2-12B-A2.5B-Thinking0 kvalificerede testsMellum2-12B-A2.5B-Thinking-SFT0 kvalificerede testsJGOS-31B-Citizen0 kvalificerede testsEXAONE-4.5-33B0 kvalificerede testsK-EXAONE-236B-A23B0 kvalificerede testsLongCat-Flash-Lite0 kvalificerede testsLongCat-Flash-Thinking-26010 kvalificerede testsMuse-Glimmer-30B0 kvalificerede testsMacaron-V1-Coding-Venti0 kvalificerede testsMacaron-V1-Tall0 kvalificerede testsMacaron-V1-Venti0 kvalificerede testsMiniMax-M2.70 kvalificerede testsMiniMax-M30 kvalificerede testsMistral-Medium-3.5-128B0 kvalificerede testsMistral-Small-4-119B-26030 kvalificerede testsIndustrialCoder0 kvalificerede testsLaguna-XS.20 kvalificerede testsNex-N2.5-Pro0 kvalificerede testsNemotron-Cascade-2-30B-A3B0 kvalificerede testsNemotron-Terminal-14B0 kvalificerede testsNemotron-Terminal-32B0 kvalificerede testsNemotron-Terminal-8B0 kvalificerede testsNVIDIA-Nemotron-3-Ultra-550B-A55B-BF160 kvalificerede testsNVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP40 kvalificerede testsNVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF160 kvalificerede testsNVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP40 kvalificerede testsopenPangu-2.0-Flash0 kvalificerede testsOrnith-1.0-35B0 kvalificerede testsOrnith-1.0-397B0 kvalificerede testsOrnith-1.0-9B0 kvalificerede testsOrnith-1.5-35B-A3B0 kvalificerede testsOrnith-1.5-397B0 kvalificerede testsOrnith-1.5-9B0 kvalificerede testsOpenSeeker-v2-30B-SFT0 kvalificerede testsLaguna-M.10 kvalificerede testsLaguna-S-2.10 kvalificerede testsLaguna-XS-2.10 kvalificerede testsLaguna-XS.20 kvalificerede testsQwen2-72B0 kvalificerede testsQwen3-4B-Instruct-25070 kvalificerede testsQwen3-Coder-Next0 kvalificerede testsQwen3.8-Flash-Next0 kvalificerede testsNVIDIA-Nemotron-3-Super-120B-A12B-BF160 kvalificerede testsA.X-K10 kvalificerede testsA.X-K20 kvalificerede testsStep-3.7-Flash0 kvalificerede testsHy30 kvalificerede testsHy3-preview0 kvalificerede testsHy4-preview0 kvalificerede testsInkling0 kvalificerede testsInkling-Small0 kvalificerede testsSolar-Open2-250B0 kvalificerede testsMiMo-V2.50 kvalificerede testsMiMo-V2.5-Pro0 kvalificerede testsGLM-4.7-Flash0 kvalificerede tests
Hvad betyder scoren?

Scoren er gennemsnittet af modellens percentilplaceringer i dokumenterede benchmarks: 100 er øverst, 50 er midten. Resultater sammenlignes inden for samme benchmark, før gennemsnittet beregnes. Mindst seks benchmarks med mindst fem sammenligningsmodeller kræves for rangering. Selvindberettede modelkort indgår som dokumentation, men ikke i rangeringen. Manglende tal estimeres ikke. Forskellig dækning og testopsætning giver fortsat usikkerhed.

Scoren er en relativ placering, ikke procent rigtige svar. 90 betyder en høj gennemsnitlig placering blandt de målte modeller. Benchmarkdækning, reasoning-niveau og brug af værktøjer kan variere. De faktiske resultater og opsætninger står på profilen.

Vi fylder aldrig manglende målinger ud for at nå 100 modeller.