---
title: "AlmanBench"
language: en
date: 2026-07-25
canonical: https://alman.ai/almanbench/
---

AlmanBench measures language and grammatical reasoning capabilities, specifically in German. Its 1,029 public items test whether a language model can apply the complete Alman specification to real German prose, from canonical literature to contemporary text.

AlmanBench also enjoys a rare form of benchmark integrity. No lab is going to burn a training run on gaming a German simplification dataset, so the scores measure what the models can actually do. Obscurity is our contamination policy.

[Dataset](https://huggingface.co/datasets/osolmaz/almanbench) · [Results](https://huggingface.co/datasets/osolmaz/almanbench-results) · [GitHub](https://github.com/osolmaz/alman) · [Specification](https://alman.ai/#spec)

## Results

| Nr. | Model | Score ↑ |
| --- | --- | --- |
| 1 | GPT-5.5 xhigh | 94.8% (976/1029) |
| 2 | Claude Opus 5 max | 94.5% (972/1029) |
| 3 | GPT-5.6 Sol max | 94.4% (971/1029) |
| 4 | Claude Fable 5 high | 94.1% (968/1029) |
| 5 | GPT-5.6 Sol xhigh | 93.5% (962/1029) |
| 6 | Claude Fable 5 max | 93.4% (961/1029) |
| 7 | GLM 5.2 | 90.4% (930/1029) |
| 8 | DeepSeek V4 Flash | 89.7% (923/1029) |
| 9 | Inkling max | 89.2% (918/1029) |
| 10 | GPT-5.6 Luna xhigh | 87.7% (902/1029) |
| 11 | GPT-5.6 Terra xhigh | 87.5% (900/1029) |
| 12 | Kimi K2.7 Code | 87.0% (895/1029) |
| 13 | DeepSeek V4 Pro | 85.3% (878/1029) |
| 14 | Claude Sonnet 5 xhigh | 83.3% (857/1029) |
| 15 | Qwen3.6 27B | 81.6% (840/1029) |
| 16 | MiniMax M3 | 79.6% (819/1029) |
| 17 | Gemma 4 26B A4B IT | 76.6% (788/1029) |
| 18 | Claude Opus 4.8 max | 74.8% (770/1029) |
| 19 | Nemotron 3 Ultra | 74.6% (768/1029) |
| 20 | Qwen3.6 35B A3B | 74.3% (765/1029) |
| 21 | Step 3.5 Flash | 69.7% (717/1029) |
| 22 | GPT-OSS 120B high | 68.8% (708/1029) |
| 23 | Gemma 4 31B IT | 64.8% (667/1029) |
| 24 | MiMo V2.5 Pro | 64.2% (661/1029) |
| 25 | Ternary Bonsai 27B | 48.2% (496/1029) |
| 26 | Laguna S 2.1 | 40.5% (417/1029) |

### Score by tier

| Model | naturalistic | targeted | guards | curated |
| --- | --- | --- | --- | --- |
| GPT-5.5 xhigh | 94.2% (565/600) | 100.0% (216/216) | 87.5% (105/120) | 96.8% (90/93) |
| Claude Opus 5 max | 94.5% (567/600) | 95.8% (207/216) | 88.3% (106/120) | 98.9% (92/93) |
| GPT-5.6 Sol max | 93.3% (560/600) | 100.0% (216/216) | 87.5% (105/120) | 96.8% (90/93) |
| Claude Fable 5 high | 93.7% (562/600) | 98.1% (212/216) | 87.5% (105/120) | 95.7% (89/93) |
| GPT-5.6 Sol xhigh | 92.2% (553/600) | 99.5% (215/216) | 87.5% (105/120) | 95.7% (89/93) |
| Claude Fable 5 max | 92.7% (556/600) | 96.3% (208/216) | 89.2% (107/120) | 96.8% (90/93) |
| GLM 5.2 | 88.3% (530/600) | 96.3% (208/216) | 86.7% (104/120) | 94.6% (88/93) |
| DeepSeek V4 Flash | 86.5% (519/600) | 97.2% (210/216) | 88.3% (106/120) | 94.6% (88/93) |
| Inkling max | 85.3% (512/600) | 100.0% (216/216) | 85.8% (103/120) | 93.5% (87/93) |
| GPT-5.6 Luna xhigh | 83.3% (500/600) | 98.1% (212/216) | 85.8% (103/120) | 93.5% (87/93) |
| GPT-5.6 Terra xhigh | 83.0% (498/600) | 96.3% (208/216) | 88.3% (106/120) | 94.6% (88/93) |
| Kimi K2.7 Code | 83.0% (498/600) | 94.9% (205/216) | 87.5% (105/120) | 93.5% (87/93) |
| DeepSeek V4 Pro | 79.7% (478/600) | 97.7% (211/216) | 87.5% (105/120) | 90.3% (84/93) |
| Claude Sonnet 5 xhigh | 77.0% (462/600) | 95.4% (206/216) | 89.2% (107/120) | 88.2% (82/93) |
| Qwen3.6 27B | 74.2% (445/600) | 97.2% (210/216) | 86.7% (104/120) | 87.1% (81/93) |
| MiniMax M3 | 71.7% (430/600) | 95.4% (206/216) | 85.0% (102/120) | 87.1% (81/93) |
| Gemma 4 26B A4B IT | 70.5% (423/600) | 93.5% (202/216) | 77.5% (93/120) | 75.3% (70/93) |
| Claude Opus 4.8 max | 65.8% (395/600) | 89.8% (194/216) | 84.2% (101/120) | 86.0% (80/93) |
| Nemotron 3 Ultra | 66.0% (396/600) | 92.1% (199/216) | 77.5% (93/120) | 86.0% (80/93) |
| Qwen3.6 35B A3B | 64.2% (385/600) | 92.6% (200/216) | 84.2% (101/120) | 84.9% (79/93) |
| Step 3.5 Flash | 58.5% (351/600) | 88.4% (191/216) | 79.2% (95/120) | 86.0% (80/93) |
| GPT-OSS 120B high | 60.3% (362/600) | 85.2% (184/216) | 75.8% (91/120) | 76.3% (71/93) |
| Gemma 4 31B IT | 50.3% (302/600) | 92.6% (200/216) | 81.7% (98/120) | 72.0% (67/93) |
| MiMo V2.5 Pro | 50.5% (303/600) | 88.0% (190/216) | 80.8% (97/120) | 76.3% (71/93) |
| Ternary Bonsai 27B | 34.8% (209/600) | 75.5% (163/216) | 65.8% (79/120) | 48.4% (45/93) |
| Laguna S 2.1 | 26.5% (159/600) | 61.6% (133/216) | 65.8% (79/120) | 49.5% (46/93) |

## Composition

The public set contains 1,029 items in four tiers. A private held-out set of about 200 further items, reviewed to the same standard, is never published and prices training contamination over time. Every published row carries a canary GUID so the data can be filtered from training corpora.

| Tier | Items | Purpose |
| --- | --- | --- |
| naturalistic | 600 | Real prose with interacting rules. Half canonical literature (1500 to 1955), half contemporary German from Wikipedia, Tatoeba, and hand-authored sentences. |
| targeted | 216 | Hand-authored items that lift rare rules to at least 25 observations each. |
| guards | 120 | Overcorrection traps in eight families. Forms that Alman keeps, which a surface-form stripper would wrongly change. |
| curated | 93 | Hand-translated demonstrative core. Every specification rule is the designated target of at least one item. |

## Example item

A naturalistic item from German Wikipedia. The genitive after *Familie* may keep *der* or take the *von* periphrasis, so the reference lists both valid renderings.

**Standard German:** Der Braunbär gehört zu den Säugetieren aus der Familie der Bären.

**Valid renderings:**

- Die Braunbär gehört zu die Säugetiere aus die Familie der Bären. (canonical)
- Die Braunbär gehört zu die Säugetiere aus die Familie von die Bären.

## Running the benchmark

The dataset is published on Hugging Face. The harness ships with the alman repository and runs against any model supported by Inspect AI.

```python
from datasets import load_dataset

ds = load_dataset("osolmaz/almanbench", split="test")
```

```console
uv run inspect eval alman/bench/task.py --model openai/gpt-5-codex
```

## Citation

```bibtex
@misc{almanbench2026,
  title        = {AlmanBench: A Standard German to Alman Translation Benchmark},
  author       = {Solmaz, Onur},
  year         = {2026},
  howpublished = {\url{https://alman.ai/almanbench/}}
}
```
