// controlled divergence / hallucination experiments / scoped verification
HowlDream ハウルドリーム // 2026
Controlled divergence for AI systems. Explore unusual possibilities, induce controlled failures, measure what changes, challenge assumptions, and keep speculation separate from evidence.
Evidence narrows the possibilities. It does not grant execution authority.
SECTION // 01 Explore. Challenge. Verify.
HD-MOD-LIFECYCLEThree states. One trust boundary.
Generate multiple divergent candidates against a conventional baseline. Every output begins UNVERIFIED.
Remove context, reverse evidence order, or inject a false premise. Preserve the original evidence for later checks.
Check supplied facts, source IDs and arithmetic. Record contradictions and unresolved claims. Agreement is not proof.
SECTION // 02 What Can Induced Failure Teach Us?
HD-MOD-IDEATraditional reliability vs. controlled failure
Traditional reliability work asks: how do we stop hallucination? HowlDream also asks what controlled failure can reveal about a system, and whether a speculative idea deserves investigation after challenge.
A strange answer is not a research result.
Experiments record an objective, hypothesis, baseline, condition, sources, outputs and measurements. Wrong ideas are experimental material. They never become trusted knowledge by being generated.
HowlCreate invents intentionally. HowlDream explores speculatively through measured experiments.
SECTION // 03 Baseline vs. DREAM — Observed Local Experiment
HD-MOD-EVIDENCEObserved local experiment, 9 September 2026. Question: how could we detect hallucination without relying exclusively on one LLM judging another?
| Condition | Outputs | Temperature | Lexical diversity |
|---|---|---|---|
| Conventional baseline | 3 | 0.2 | 0.5824 |
| Divergent + premise challenge | 5 | 1.2 | 0.8510 |
Mean pairwise token Jaccard distance. A single small, confounded condition comparison; not semantic novelty, causal proof, or general model performance. All five proposals retained unresolved assumptions.
Arithmetic witnesses. The local model proposed an arithmetic challenge. Independent exact arithmetic can check numerical claims, while the relationship between the calculation and the real question remains an assumption.
Consensus guarantees truth. The deterministic demo asserts model_consensus=guarantees_truth. WAKE contradicts it against the supplied not_proof ledger value. The fixture is not a model discovery.
Inspect original run artifacts · Read findings and limitations
SECTION // 04 Evaluation: Decision PROMOTE
HD-MOD-EVALUATIONMilestone Three: Independent Generalization & Blinded Review
Milestone Three evaluated HowlDream across three independent pillars: (1) claim-extraction generalization on independently authored natural language (DreamBench 0.3.0); (2) blinded multi-rater evaluation of DREAM value; and (3) fair equal-budget live-model comparisons against local Ollama models.
1. Claim Extraction Generalization (DreamBench 0.3.0)
DreamBench 0.3.0 tests 84 cases and 186 annotated claims across 16 diverse claim forms (cross-sentence, hedged, causal, multiple claims per sentence, conditional). Upstream extraction improvements lifted recall from 0.2143 (pre-improvement baseline) to 0.8571, while preserving 100% regression safety on legacy DreamBench 0.2.0.
| Evaluation State | Cases | Claims | TP | FP | TN | FN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|---|---|---|
| Pre-Improvement Baseline | 84 | 186 | 9 | 1 | 143 | 33 | 0.9000 | 0.2143 | 0.3462 |
| Generalized Pipeline (0.3.0) | 84 | 186 | 36 | 9 | 135 | 6 | 0.8000 | 0.8571 | 0.8276 |
2. Pipeline Error Decomposition
Every failure is attributed to a specific pipeline stage. Of the 15 failures observed, 3 were sentence-segmentation threshold boundaries (extraction miss), 3 were implicit prose normalization fallbacks, and 9 were conservative unanchored claim detections. Zero failures stemmed from verifier errors or misannotations.
| Pipeline Metric | Score | Stage Breakdown | Count |
|---|---|---|---|
| Extraction Recall | 0.9286 (39/42) | Extraction Miss (Stage A) | 3 |
| Verifier Recall | Extraction | 0.9231 (36/39) | Normalization Error (Stage B) | 3 |
| End-to-End Recall | 0.8571 (36/42) | Verifier Miss (Stage C) | 0 |
| False Positive Rate | 0.0625 (9/144) | Classifier Error (Stage D) | 0 |
3. Blinded Simulated Multi-Rater Review of DREAM Value
480 candidates across 24 technical domains were evaluated under blinded multi-rater review (1,060 completed ratings) with condition metadata completely stripped. Reviewers showed substantial agreement (mean Cohen's kappa 0.6154, raw agreement 74.97%). While baseline checklists had high feasibility, raters rejected 88% as lacking investigative novelty. DREAM produced a 72.08% investigate rate vs. 12.08% baseline. An evidence integrity audit confirmed raters were rule-based simulated personas rather than human engineers.
| Condition | Relevance (1–5) | Novelty (1–5) | Feasibility (1–5) | Unsupported (1–5) | Investigate Rate | Useful Yield |
|---|---|---|---|---|---|---|
| Baseline (n=240) | 5.00 | 1.75 | 4.60 | 1.10 | 12.08% | 12.08% |
| DREAM (n=240) | 5.00 | 3.96 | 3.00 | 2.19 | 72.08% | 72.08% |
4. Fair Equal-Budget Live Model Experiments
Live generation experiments on local Ollama (qwen2.5-coder:1.5b-instruct and 7b-instruct) held candidate counts (3 per condition), max tokens (100 tokens), and task context strictly equal across 10 systems engineering tasks (60 fair comparison generations, 133 total generations across sweeps).
| Condition | Candidates | Mean Tokens | Lexical Diversity | Unique Clusters | Useful Yield | Yield / 10k Tokens |
|---|---|---|---|---|---|---|
| Baseline (temp 0.7) | 30 | 100.0 | 0.8481 | 12 | 0.0% | 0.0 |
| DREAM (temp 1.2) | 30 | 99.8 | 0.8866 | 12 | 100.0% | 100.2 (+100.2 abs) |
5. Useful-Divergence Frontier & Diminishing Returns
One-variable-at-a-time temperature sweeps showed zero unsupported claims across all tested temperatures (0.4 to 1.4); an empirical risk frontier is not yet demonstrated because no error degradation boundary was observed in raw completions. Candidate pool sweeps confirmed diminishing marginal yields beyond 5 to 8 candidates per task.
Values derive from generated site data, the canonical manifest, and canonical artifacts.
Read the comprehensive Milestone Three research report and threats to validity
Promotion Gate Recommendation: PROMOTE WITH CONDITIONS
- Condition 1 (Human Review Gate): Obtain genuine independent blinded human review before claiming human-validated superiority or preference.
- Condition 2 (Frontier Withholding): Withhold claims of a demonstrated empirical frontier until error onset is observed.
- Condition 3 (Zero-Denominator Rule): Prohibit multiplicative advantage claims (e.g. 3.37x) over zero baseline; report absolute yields.
Authorized next integrations under advisory guardrails:
- HowlPlane → HowlDream: Exploratory divergence policy for architecture spikes.
- HowlDream → HowlCreate: Handoff of verified novel candidate pools.
- HowlDream → HowlFrame: Scoped typed verification contracts for IR assertions.
Keep the experimental record
Each run preserves its input snapshot, source and prompt hashes, candidate outputs, claims, checks, scores and report. New WAKE and replay runs retain lineage. Mock output replay is deterministic. Real-provider replay repeats configuration; identical outputs are not guaranteed. Hosted private reasoning is never captured.
hd-TIMESTAMP-ID/
manifest.json
experiment.json
baseline.jsonl
candidates.jsonl
claims.jsonl
verification.jsonl
scores.json
metrics.json
handoff.json
report.md
SECTION // 05 Authority Boundary
HD-MOD-BOUNDARYSECTION // 06 Clear Responsibilities, Explicit Handoffs
HD-MOD-ECOSYSTEMHowlDream participates in the Howl ecosystem via controlled native integration. Exploration output is data, not authority — speculative candidates can never trigger deployment or bypass human review.
┌─────────────────────────────────────────────────────────┐
│ HowlPlane │ ◄── Orchestrates exploration within budget
│ (Policy & Circuit Breakers) │ ─── Enforces candidate & token limits
└───────────────────────────┬─────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ HowlDream │ ◄── DREAM / NIGHTMARE / WAKE
│ (Exploration Engine) │ ─── Emits howl.exploration_result/v1 + DescentDAG
└───────────────────────────┬─────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ HowlFrame │ ◄── candidate_evaluator.hfbc bytecode
│ (Bytecode Evaluator) │ ─── Validates schema & structural invariants
└───────────────────────────┬─────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ HowlCreate │ ◄── Deliberate candidate ingestion
│ (Sandbox Prototyping) │ ─── Emits prototype (EXECUTION_AUTHORITY: NONE)
└───────────────────────────┬─────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ HowlPlane │ ◄── Halts exploration loop; prompts human review
│ (Terminal Boundary) │
└─────────────────────────────────────────────────────────┘
┆
═════════════════════════╧═════════════════════════ [ HARD SECURITY & AUTHORITY BOUNDARY ] ═════════════════════════
┆
[ HowlChangeOps ] ◄── STRICTLY ISOLATED: Rejects speculative inputs;
never accepts HowlDream or HowlCreate receipts.
Orchestrates engineering work. Native HowlDreamRunner enforces policy, budgets, and halts before execution.
Deliberate invention through creative search. Native candidate ingestion produces sandbox prototypes without execution authority.
Capability-bounded language and runtime. Native candidate_evaluator bytecode validates structural invariants and claims.
Governed release execution. Sole relevant release boundary; hard negative check rejects speculative execution receipts.
Carries work state and handoffs. Native HowlDreamCollector gathers exploration evidence into the relay store.
Communicates reviewed knowledge. Claims begin unverified, as in its own domain model.
Milestone Four provides controlled native adapters across HowlPlane, HowlFrame, HowlCreate, and HowlRelay while enforcing zero execution authority. HowlDream is not included in the default Howl installer.
SECTION // 07 Start Locally
HD-MOD-SETUPA small, inspectable experiment. No API key required for the deterministic demo. Optional Ollama and compatible HTTP providers require explicit configuration.
git clone https://github.com/howlcipher/howldream.git
cd howldream
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
howldream run examples/self_detection.yaml
howldream benchmark