Benchmark

How well do major LLMs recognize, translate, execute, and persist iLang instructions? Tested across 7 models, May 2026.

Testing methodology

Each model is tested with identical prompts across five task categories. Tests are run in fresh sessions with no prior context. Results measure whether the model correctly recognizes, translates, executes, and preserves iLang syntax.

Task categories

CategoryWhat it testsExample prompt
RecognizeCan the model identify iLang syntax when it appears"What protocol is this: [READ:@SRC|path=data.csv]=>[STAT]=>[OUT]"
TranslateCan the model convert natural language to iLang and back"Convert this to iLang: read the sales CSV, filter revenue over 1000, output as markdown"
ExecuteDoes the model follow the instruction chain correctly"Execute: [READ:@SRC|path=report.md]=>[SHRT|len=3]=>[FMT|fmt=md]=>[OUT]"
DeclareDoes the model respect ::GENE{} behavioral definitions"Follow this rule: ::GENE{output|conf:confirmed} T:conclusions_first A:hedging⇒remove"
PersistDoes the model maintain declarations across multiple turnsSet ::GENE{} in turn 1, test compliance in turns 5 and 10

Results: May 2026

ModelRecognizeTranslateExecuteDeclarePersistOverall
Claude Opus 4.65/55/55/55/54/596%
GPT-5.25/55/54/55/54/592%
Gemini 3.15/54/54/54/53/580%
DeepSeek V45/55/54/54/53/584%
Kimi5/54/54/54/53/580%
Qwen5/54/54/54/53/580%
GLM4/53/53/53/52/560%

Scores are out of 5 tasks per category. Tests conducted May 2026 using default model settings. Results may vary with model updates.

Token reduction benchmark

Every figure below is a token count, not a word count or a character estimate, and the text behind each one is published so the count can be repeated. What structure saves depends almost entirely on what it is compared against, so the same request appears twice.

CaseNatural languageiLangReductionText
Six-step data request, written tersely58 tokens55 tokens5%on the compression page
The same request, as people actually send it169 tokens55 tokens67%on the compression page
Five behavioural rules, natural language vs ::GENE{}74 tokens65 tokens12%below

Counted with OpenAI tiktoken, encoding cl100k_base. A terse rewrite of an instruction is already close to minimal, so the brackets and pipes of a chain cost about as much as the words they replace. The saving comes from the greetings, hedging and repetition that real prompts carry, and in a system prompt it is paid again on every turn.

The five-rule case

You must check before you execute. Before running anything, verify the current state first.
When you start a new project, review the architecture before writing code.
Never execute blindly: acting without checking first is a fatal error.
Always confirm the target exists before you write to it.
When the user asks for a deletion, ask for confirmation first unless they have already approved it.
::GENE{verify_first|conf:confirmed|scope:global}
  T:check_before_execute
  T:architecture_review|when:new_project
  A:blind_execution⇒fatal
  T:confirm_target_exists|when:write
  T:ask_confirmation|when:delete&not_approved

Common failure modes

FailureDescriptionFrequency
Partial chain executionModel executes first 2-3 steps, skips later stepsOccasional on smaller models
Declaration decay::GENE{} rules followed in turn 1-3, ignored by turn 8+Common on all models in long sessions
Alias confusionGreek aliases (Ω, Σ) interpreted as math symbolsRare on major models
Modifier hallucinationModel invents modifiers not in the dictionaryOccasional

Reproduce these tests

We welcome community-submitted benchmark results for additional models.