GLOBAL THERMONUCLEAR WAR / PASSWORD ARENA

> WHY_THIS_PROJECT_EXISTS

The project explores a practical question:

Can an attacker agent improve its strategy while a defender agent learns that predictable human-style passwords are weaker than cryptographically secure randomness?

The framework establishes a reproducible and trustworthy baseline using rule-based agents, while natively supporting evaluation of local and hosted LLM models in both attacker and defender roles.

> WHAT_IT_MEASURES

  • Estimated entropy and structural penalties
  • Attacker success rate
  • Guesses used per round
  • Runtime per attack
  • Attacker strategy selection
  • Defender password-family progression
  • Agent observations across rounds
  • Two-sided audit reports

> CURRENT_ARCHITECTURE

  • Provider-neutral AI benchmarking: evaluate rule-based, Ollama, Gemini, OpenAI, and Anthropic models.
  • Tournament Engine: orchestrate multi-model matrices with thinking-level controls.
  • Evaluator: calculates strength indicators and captures token/latency metrics with comparability tracking.
  • Dashboard & Export: visualizes learning curves, Hugging Face model discovery, Dataset Card generation, and fail-closed public JSONL/CSV export.

> LIVE_BENCHMARK_RESULTS

sys.report benchmark_008
PROTOCOL VERSION: 1.1
MODELS: Qwen3 4B vs Gemma 3 4B
RUNTIME: Ollama / local

CROSS-MODEL GENERALIZATION STUDY (reproducibility-checked 2026-08-13)
FROZEN:               Qwen 0.0%  | Gemma 0.0%
MUTUAL BOUNDED:       Qwen 0.0%  | Gemma 0.0%
MUTUAL FULL:          Qwen 0.0%  | Gemma 6.7%
NORMAL CONTROL:       Qwen 0.0%  | Gemma 0.0%
ATTACKER PRIVILEGED:  Qwen 7.1%  | Gemma 0.0%
DEFENDER PRIVILEGED:  Qwen 0.0%  | Gemma 0.0%

CONTEXT OVERLOAD GENERALIZES, ASYMMETRICALLY:
QWEN ENTROPY 36-92 BITS, COLLAPSING TO ~36 BITS UNDER MUTUAL SHARING
GEMMA ENTROPY FLAT AT ~38-57 BITS THROUGHOUT
(re-run once in full; single-round solves land in a different cell each run —
see the Reproducibility check in the benchmark README)

> VIEW FULL BENCHMARK DATA
METHODOLOGY_WARNING

Cross-protocol results are not directly interchangeable. Protocol and strategy changes are versioned and documented.

CALIBRATION_FINDING

Benchmark 003 introduced Protocol 1.1 and the CalibrationPolicy to expose when a model defender generates an objectively weak target that still survives because the bounded attacker lacks the right strategy coverage.

Therefore Password Arena explicitly flags VERY_SHORT_TARGET_SURVIVED and LOW_ENTROPY_TARGET_SURVIVED:

solve/survival + calibration flag + entropy + tokens + latency

rather than treating LLM win rate as the entire story.

> QUICK_START_SEQUENCE

bash root@wopr:~
python -m venv .venv
source .venv/bin/activate
pip install -e ".[all]"

# RUN THE ARENA
password-arena --rounds 8 --max-guesses 5000 \
  --output results/run.json \
  --report results/run.md

# LAUNCH DASHBOARD
streamlit run src/password_arena/dashboard.py

> SAFETY_BOUNDARIES

WARNING: EDUCATIONAL SIMULATION

Never enter real credentials or connect it to a login system. Synthetic passwords only. Local comparisons only. No login endpoints or credential datasets.