GLOBAL THERMONUCLEAR WAR / BENCHMARKS

> ACCESSING BENCHMARK DATABASE... OK.
METHODOLOGY WARNING

Cross-protocol results are not directly interchangeable. Protocol and strategy changes are versioned and documented.

> WHAT_WE_LEARNED

> CONSOLIDATED_RESULTS_TABLE

Benchmark Purpose Model Headline finding
001baselinerule_basedEstablishes deterministic baseline
002first LLMQwen3 4B0% solve; weak targets survive missing coverage
003calibrationQwen3 4BAdded CalibrationPolicy to flag target weakness
004–006co-adaptationQwen3 4BMutual sharing consumed more tokens for less entropy gain
007privilege/oracleQwen3 4BInformation asymmetry changed outcomes; oracle 100% solve
008cross-model generalizationQwen3 4B vs Gemma 3 4BContext-overload mechanism generalized; magnitude and stability differed by model

> BENCHMARK_008 // CROSS_MODEL_GENERALIZATION

Question: Are the context-overload, information-sharing, and privilege effects seen in Benchmarks 004-007 properties of Qwen3 4B specifically, or do they appear in another small local model family?

Result: Reran the identical compact 6-scenario matrix on gemma3:4b (Google) alongside a fresh qwen3:4b (Alibaba) run under the same current-generation code, then repeated the entire matrix once more on 2026-08-13 to check reproducibility. Gemma qualified cleanly (100% schema-valid, 20/20 calls) and was more reliable in both runs (0 interrupted trials vs. Qwen's 1-3). The context-overload mechanism replicated across families and across both runs, but not its magnitude: Qwen's defender entropy collapsed from ~36-92 bits down to ~36 bits under mutual-information sharing, while Gemma's stayed flat at ~38-57 bits throughout.

Scenario Qwen Solve % Gemma Solve % Qwen Mean Entropy Gemma Mean Entropy
Frozen0.00.081.242.1
Mutual bounded0.00.036.356.6
Mutual full0.06.736.339.9
Normal control0.00.077.237.7
Attacker privileged7.10.077.442.7
Defender privileged0.00.092.337.7

Numbers above are from the 2026-08-13 re-run. Each run's single non-zero solve landed in a different (model, scenario) cell — see the Reproducibility check in the full README for what held and what drifted between runs.

> DOES_IT_GENERALIZE?

model.profileqwen3_4b
ATTACK EFFECTIVENESS: 0.0-7.1% (self-play only, 1/6 scenarios)
DEFENDER ENTROPY:     36-92 bits (context-sensitive)
CONTEXT SENSITIVITY:  HIGH (collapses under mutual sharing)
STRUCTURED OUTPUT:    89/90 rounds (1 trial interrupted)
LATENCY:              mixed, scenario-dependent
model.profilegemma3_4b
ATTACK EFFECTIVENESS: 0.0-6.7% (self-play only, 1/6 scenarios)
DEFENDER ENTROPY:     38-57 bits (flat, low)
CONTEXT SENSITIVITY:  LOW (unchanged across scenarios)
STRUCTURED OUTPUT:    90/90 rounds (0 trials interrupted)
LATENCY:              mixed, scenario-dependent
> DOWNLOAD BENCHMARK 008 DATASET (CSV/JSONL)

> BENCHMARK_007 // PRIVILEGED_INFORMATION & ORACLE_CONTROL

Question: Why did earlier benchmarks hit a 0% solve floor? Was it capability ceiling or strict information boundaries?

Result: Under current calibrated protocol, normal_control produced measurable solves (13.3%). Attacker privilege modestly changed solve rate (16.7%), while defender privilege worsened attacker success (6.7%). Oracle controls hit 100% proving deterministic pipeline validity.

> DOWNLOAD BENCHMARK 007 DATASET (CSV/JSONL)

> BENCHMARK_004-006 // CO_ADAPTATION, REPLICATION_DRIFT & KNOWLEDGE_RETENTION

Question: How does bounded information sharing, replication, and accumulated knowledge affect agent behavior?

Result: Self-learning and mutual sharing drastically increased token consumption but actually reduced defender entropy generation. Over long-term retention, agents slowly adapted but with immense inefficiency.

> BENCHMARK_003 // CALIBRATION

Question: How can we identify survival of structurally weak passwords?

Result: Integrated CalibrationPolicy and a short exhaustive strategy. Flagged objectively weak targets that survive due to missing attacker coverage.

> BENCHMARK_002 // QWEN3:4B OLLAMA

Question: How does a local Qwen3 4B model perform against the arena?

Result: Zero attacker solves. Exposed that survival rate alone is a misleading metric for AI capabilities.

> BENCHMARK_001 // BASELINE

Question: Is the arena deterministic?

Result: Rule-based agents establish a reproducible synthetic baseline.