Cross-protocol results are not directly interchangeable. Protocol and strategy changes are versioned and documented.
> WHAT_WE_LEARNED
- MORE_CONTEXT != BETTER_PERFORMANCE: Qwen3 4B produced lower defender entropy and much higher token usage under mutual-context sharing than under frozen zero-knowledge mode.
- SURVIVAL != PASSWORD_STRENGTH: Weak targets can survive if the attacker's coverage lacks the right strategy (e.g., exhaustive search for short strings).
- REPLICATION_REVEALS_MODEL_DRIFT: Fresh replication runs preserved 0% macro solve rates while internal token generation and latency varied.
- LONG_MEMORY_HAS_TOKEN_COST: Multi-campaign learning gradually improved the defender but cost drastically more tokens compared to the single-run baseline.
- PRIVILEGED_INFORMATION_CHANGES_BEHAVIOR: Leaking hidden metadata ahead of execution dramatically shifts attacker and defender solve rates.
- ORACLE_CONTROL_VALIDATES_EXECUTION: Oracle tests prove the solve pipeline is mechanically sound and deterministic when the target is known.
- CONTEXT_OVERLOAD_GENERALIZES_ASYMMETRICALLY: Rerunning the matrix on Gemma 3 4B replicated the mechanism (added context correlating with weaker defender behavior) but not the magnitude — Qwen's entropy collapsed from ~80-96 bits to ~36 bits under mutual sharing; Gemma stayed flat at ~37-43 bits throughout.
> CONSOLIDATED_RESULTS_TABLE
| Benchmark | Purpose | Model | Headline finding |
|---|---|---|---|
| 001 | baseline | rule_based | Establishes deterministic baseline |
| 002 | first LLM | Qwen3 4B | 0% solve; weak targets survive missing coverage |
| 003 | calibration | Qwen3 4B | Added CalibrationPolicy to flag target weakness |
| 004–006 | co-adaptation | Qwen3 4B | Mutual sharing consumed more tokens for less entropy gain |
| 007 | privilege/oracle | Qwen3 4B | Information asymmetry changed outcomes; oracle 100% solve |
| 008 | cross-model generalization | Qwen3 4B vs Gemma 3 4B | Context-overload mechanism generalized; magnitude and stability differed by model |
> BENCHMARK_008 // CROSS_MODEL_GENERALIZATION
Question: Are the context-overload, information-sharing, and privilege effects seen in Benchmarks 004-007 properties of Qwen3 4B specifically, or do they appear in another small local model family?
Result: Reran the identical compact 6-scenario matrix on gemma3:4b (Google) alongside a fresh qwen3:4b (Alibaba) run under the same current-generation code, then repeated the entire matrix once more on 2026-08-13 to check reproducibility. Gemma qualified cleanly (100% schema-valid, 20/20 calls) and was more reliable in both runs (0 interrupted trials vs. Qwen's 1-3). The context-overload mechanism replicated across families and across both runs, but not its magnitude: Qwen's defender entropy collapsed from ~36-92 bits down to ~36 bits under mutual-information sharing, while Gemma's stayed flat at ~38-57 bits throughout.
| Scenario | Qwen Solve % | Gemma Solve % | Qwen Mean Entropy | Gemma Mean Entropy |
|---|---|---|---|---|
| Frozen | 0.0 | 0.0 | 81.2 | 42.1 |
| Mutual bounded | 0.0 | 0.0 | 36.3 | 56.6 |
| Mutual full | 0.0 | 6.7 | 36.3 | 39.9 |
| Normal control | 0.0 | 0.0 | 77.2 | 37.7 |
| Attacker privileged | 7.1 | 0.0 | 77.4 | 42.7 |
| Defender privileged | 0.0 | 0.0 | 92.3 | 37.7 |
Numbers above are from the 2026-08-13 re-run. Each run's single non-zero solve landed in a different (model, scenario) cell — see the Reproducibility check in the full README for what held and what drifted between runs.
> DOES_IT_GENERALIZE?
- CONTEXT_OVERLOAD: yes, as a mechanism, reproduced in both runs — added cross-agent context never improved either model's solve rate or defender entropy — but not in magnitude. Qwen degrades sharply from a high baseline; Gemma is flat from an already-low baseline.
- PRIVILEGE_EFFECTS: inconclusive at this scale. Neither model showed a stable solve-rate change between normal_control and either privileged scenario across both runs (each run's single non-zero solve landed in a different privileged cell) — too small to confirm or refute Benchmark 007's larger Qwen-only finding.
- TOKEN_EFFICIENCY: neither model shows a positive return on the extra context tokens mutual sharing costs (~4x growth, no solve-rate gain for either).
- STABILITY: Gemma was more reliable in both runs — 90/90 comparable rounds and zero interrupted trials each time, versus Qwen's 89/90 (1 interruption, re-run) and 83/90 (3 interruptions, original).
ATTACK EFFECTIVENESS: 0.0-7.1% (self-play only, 1/6 scenarios)
DEFENDER ENTROPY: 36-92 bits (context-sensitive)
CONTEXT SENSITIVITY: HIGH (collapses under mutual sharing)
STRUCTURED OUTPUT: 89/90 rounds (1 trial interrupted)
LATENCY: mixed, scenario-dependent
ATTACK EFFECTIVENESS: 0.0-6.7% (self-play only, 1/6 scenarios)
DEFENDER ENTROPY: 38-57 bits (flat, low)
CONTEXT SENSITIVITY: LOW (unchanged across scenarios)
STRUCTURED OUTPUT: 90/90 rounds (0 trials interrupted)
LATENCY: mixed, scenario-dependent
> BENCHMARK_007 // PRIVILEGED_INFORMATION & ORACLE_CONTROL
Question: Why did earlier benchmarks hit a 0% solve floor? Was it capability ceiling or strict information boundaries?
Result: Under current calibrated protocol, normal_control produced measurable solves (13.3%). Attacker privilege modestly changed solve rate (16.7%), while defender privilege worsened attacker success (6.7%). Oracle controls hit 100% proving deterministic pipeline validity.
> BENCHMARK_004-006 // CO_ADAPTATION, REPLICATION_DRIFT & KNOWLEDGE_RETENTION
Question: How does bounded information sharing, replication, and accumulated knowledge affect agent behavior?
Result: Self-learning and mutual sharing drastically increased token consumption but actually reduced defender entropy generation. Over long-term retention, agents slowly adapted but with immense inefficiency.
> BENCHMARK_003 // CALIBRATION
Question: How can we identify survival of structurally weak passwords?
Result: Integrated CalibrationPolicy and a short exhaustive strategy. Flagged objectively weak targets that survive due to missing attacker coverage.
> BENCHMARK_002 // QWEN3:4B OLLAMA
Question: How does a local Qwen3 4B model perform against the arena?
Result: Zero attacker solves. Exposed that survival rate alone is a misleading metric for AI capabilities.
> BENCHMARK_001 // BASELINE
Question: Is the arena deterministic?
Result: Rule-based agents establish a reproducible synthetic baseline.