> WHY_THIS_PROJECT_EXISTS
The project explores a practical question:
Can an attacker agent improve its strategy while a defender agent learns that predictable human-style passwords are weaker than cryptographically secure randomness?
The framework establishes a reproducible and trustworthy baseline using rule-based agents, while natively supporting evaluation of local and hosted LLM models in both attacker and defender roles.
> WHAT_IT_MEASURES
- Estimated entropy and structural penalties
- Attacker success rate
- Guesses used per round
- Runtime per attack
- Attacker strategy selection
- Defender password-family progression
- Agent observations across rounds
- Two-sided audit reports
> CURRENT_ARCHITECTURE
- Provider-neutral AI benchmarking: evaluate rule-based, Ollama, Gemini, OpenAI, and Anthropic models.
- Tournament Engine: orchestrate multi-model matrices with thinking-level controls.
- Evaluator: calculates strength indicators and captures token/latency metrics with comparability tracking.
- Dashboard & Export: visualizes learning curves, Hugging Face model discovery, Dataset Card generation, and fail-closed public JSONL/CSV export.
> LIVE_BENCHMARK_RESULTS
PROTOCOL VERSION: 1.1
MODELS: Qwen3 4B vs Gemma 3 4B
RUNTIME: Ollama / local
CROSS-MODEL GENERALIZATION STUDY (reproducibility-checked 2026-08-13)
FROZEN: Qwen 0.0% | Gemma 0.0%
MUTUAL BOUNDED: Qwen 0.0% | Gemma 0.0%
MUTUAL FULL: Qwen 0.0% | Gemma 6.7%
NORMAL CONTROL: Qwen 0.0% | Gemma 0.0%
ATTACKER PRIVILEGED: Qwen 7.1% | Gemma 0.0%
DEFENDER PRIVILEGED: Qwen 0.0% | Gemma 0.0%
CONTEXT OVERLOAD GENERALIZES, ASYMMETRICALLY:
QWEN ENTROPY 36-92 BITS, COLLAPSING TO ~36 BITS UNDER MUTUAL SHARING
GEMMA ENTROPY FLAT AT ~38-57 BITS THROUGHOUT
(re-run once in full; single-round solves land in a different cell each run —
see the Reproducibility check in the benchmark README)
> VIEW FULL BENCHMARK DATA
Cross-protocol results are not directly interchangeable. Protocol and strategy changes are versioned and documented.
Benchmark 003 introduced Protocol 1.1 and the CalibrationPolicy to expose when a model defender generates an objectively weak target that still survives because the bounded attacker lacks the right strategy coverage.
Therefore Password Arena explicitly flags VERY_SHORT_TARGET_SURVIVED and LOW_ENTROPY_TARGET_SURVIVED:
solve/survival + calibration flag + entropy + tokens + latency
rather than treating LLM win rate as the entire story.
> QUICK_START_SEQUENCE
python -m venv .venv
source .venv/bin/activate
pip install -e ".[all]"
# RUN THE ARENA
password-arena --rounds 8 --max-guesses 5000 \
--output results/run.json \
--report results/run.md
# LAUNCH DASHBOARD
streamlit run src/password_arena/dashboard.py
> SAFETY_BOUNDARIES
Never enter real credentials or connect it to a login system. Synthetic passwords only. Local comparisons only. No login endpoints or credential datasets.