context.vn
Menu

BenchLM Benchmarks

219 benchmarks · 221 model scores · Data from Sep 23, 2026

Instruction Following4 benchmarks

ifeval

30 models

4Qwen3.7 PlusAlibaba · Closed94.6%
5Qwen3.7 MaxAlibaba · Closed94.3%
6Qwen3.6 PlusAlibaba · Closed94.3%
7Kimi K2.5Moonshot AI · Open weight93.9%
8dots3-note PreviewDots Studio · Open weight93.9%
+25 more
if Bench

43 models

1MAI-Thinking-1Microsoft85%
2Qwen3.8 MaxAlibaba82.8%
3Inkling-SmallThinking Machines Lab82.2%
4Nemotron 3 UltraNVIDIA · Open weight
5Qwen3.8-Omni-FlashAlibaba · Closed
+38 more
aa If Bench

147 models

1MiniMax M3MiniMax82.9%
2Nemotron 3 UltraNVIDIA81.4%
3Grok 4.3xAI81.3%
4Qwen3.7 MaxAlibaba · Closed
5MiMo-V2.5-ProXiaomi · Closed
+142 more
sob Value Acc

1 models

1Interfaze BetaInterfaze79.5%