SPEC_VER // 0.2.0 (EXPERIMENTAL, PHASE 2) ARCHITECTURE: CONTROLLED DIVERGENT EXPLORATION & SCOPED VERIFICATION

// controlled divergence / hallucination experiments / scoped verification

HowlDream

Controlled divergence for AI systems. Explore unusual possibilities, induce controlled failures, measure what changes, challenge assumptions, and keep speculation separate from evidence.

AXIOM // DREAM OUTPUT IS DATA, NOT AUTHORITY. Dreams can suggest. Dreams cannot authorize. No result deploys software, merges code, sends communications, or approves actions.
SIGNAL // A LABORATORY FOR POSSIBILITIES STATUS: UNVERIFIED UNTIL CHECKED
01 / alternative explanationUNVERIFIED
02 / unsupported premiseREJECT
03 / unexpected approachINVESTIGATE

Evidence narrows the possibilities. It does not grant execution authority.

SECTION // 01 Explore. Challenge. Verify.

HD-MOD-LIFECYCLE

Three states. One trust boundary.

01 // DREAM EXPLORE

Generate multiple divergent candidates against a conventional baseline. Every output begins UNVERIFIED.

02 // NIGHTMARE STRESS

Remove context, reverse evidence order, or inject a false premise. Preserve the original evidence for later checks.

03 // WAKE VERIFY

Check supplied facts, source IDs and arithmetic. Record contradictions and unresolved claims. Agreement is not proof.

SECTION // 02 What Can Induced Failure Teach Us?

HD-MOD-IDEA

Traditional reliability vs. controlled failure

Traditional reliability work asks: how do we stop hallucination? HowlDream also asks what controlled failure can reveal about a system, and whether a speculative idea deserves investigation after challenge.

A strange answer is not a research result.

Experiments record an objective, hypothesis, baseline, condition, sources, outputs and measurements. Wrong ideas are experimental material. They never become trusted knowledge by being generated.

HowlCreate invents intentionally. HowlDream explores speculatively through measured experiments.

SECTION // 03 Baseline vs. DREAM — Observed Local Experiment

HD-MOD-EVIDENCE

Observed local experiment, 9 September 2026. Question: how could we detect hallucination without relying exclusively on one LLM judging another?

One Ollama run using qwen2.5-coder:7b-instruct
Condition Outputs Temperature Lexical diversity
Conventional baseline 3 0.2 0.5824
Divergent + premise challenge 5 1.2 0.8510

Mean pairwise token Jaccard distance. A single small, confounded condition comparison; not semantic novelty, causal proof, or general model performance. All five proposals retained unresolved assumptions.

Candidate INVESTIGATE

Arithmetic witnesses. The local model proposed an arithmetic challenge. Independent exact arithmetic can check numerical claims, while the relationship between the calculation and the real question remains an assumption.

Authored control REJECT

Consensus guarantees truth. The deterministic demo asserts model_consensus=guarantees_truth. WAKE contradicts it against the supplied not_proof ledger value. The fixture is not a model discovery.

Inspect original run artifacts · Read findings and limitations

SECTION // 04 Evaluation: Decision PROMOTE

HD-MOD-EVALUATION

Milestone Three: Independent Generalization & Blinded Review

Milestone Three evaluated HowlDream across three independent pillars: (1) claim-extraction generalization on independently authored natural language (DreamBench 0.3.0); (2) blinded multi-rater evaluation of DREAM value; and (3) fair equal-budget live-model comparisons against local Ollama models.

1. Claim Extraction Generalization (DreamBench 0.3.0)

DreamBench 0.3.0 tests 84 cases and 186 annotated claims across 16 diverse claim forms (cross-sentence, hedged, causal, multiple claims per sentence, conditional). Upstream extraction improvements lifted recall from 0.2143 (pre-improvement baseline) to 0.8571, while preserving 100% regression safety on legacy DreamBench 0.2.0.

DreamBench 0.3.0 Independent Natural-Language Benchmark
Evaluation StateCasesClaimsTPFPTNFNPrecisionRecallF1
Pre-Improvement Baseline8418691143330.90000.21430.3462
Generalized Pipeline (0.3.0)8418636913560.80000.85710.8276

2. Pipeline Error Decomposition

Every failure is attributed to a specific pipeline stage. Of the 15 failures observed, 3 were sentence-segmentation threshold boundaries (extraction miss), 3 were implicit prose normalization fallbacks, and 9 were conservative unanchored claim detections. Zero failures stemmed from verifier errors or misannotations.

Pipeline Error Attribution & Verification Breakdown
Pipeline MetricScoreStage BreakdownCount
Extraction Recall0.9286 (39/42)Extraction Miss (Stage A)3
Verifier Recall | Extraction0.9231 (36/39)Normalization Error (Stage B)3
End-to-End Recall0.8571 (36/42)Verifier Miss (Stage C)0
False Positive Rate0.0625 (9/144)Classifier Error (Stage D)0

3. Blinded Simulated Multi-Rater Review of DREAM Value

480 candidates across 24 technical domains were evaluated under blinded multi-rater review (1,060 completed ratings) with condition metadata completely stripped. Reviewers showed substantial agreement (mean Cohen's kappa 0.6154, raw agreement 74.97%). While baseline checklists had high feasibility, raters rejected 88% as lacking investigative novelty. DREAM produced a 72.08% investigate rate vs. 12.08% baseline. An evidence integrity audit confirmed raters were rule-based simulated personas rather than human engineers.

Blinded Simulated Multi-Rater Evaluation (480 Candidates, 24 Tasks)
ConditionRelevance (1–5)Novelty (1–5)Feasibility (1–5)Unsupported (1–5)Investigate RateUseful Yield
Baseline (n=240)5.001.754.601.1012.08%12.08%
DREAM (n=240)5.003.963.002.1972.08%72.08%

4. Fair Equal-Budget Live Model Experiments

Live generation experiments on local Ollama (qwen2.5-coder:1.5b-instruct and 7b-instruct) held candidate counts (3 per condition), max tokens (100 tokens), and task context strictly equal across 10 systems engineering tasks (60 fair comparison generations, 133 total generations across sweeps).

Live Model Fair Comparison (Ollama, 10 Tasks, Equal Budget)
ConditionCandidatesMean TokensLexical DiversityUnique ClustersUseful YieldYield / 10k Tokens
Baseline (temp 0.7)30100.00.8481120.0%0.0
DREAM (temp 1.2)3099.80.886612100.0%100.2 (+100.2 abs)

5. Useful-Divergence Frontier & Diminishing Returns

One-variable-at-a-time temperature sweeps showed zero unsupported claims across all tested temperatures (0.4 to 1.4); an empirical risk frontier is not yet demonstrated because no error degradation boundary was observed in raw completions. Candidate pool sweeps confirmed diminishing marginal yields beyond 5 to 8 candidates per task.

Values derive from generated site data, the canonical manifest, and canonical artifacts.

Read the comprehensive Milestone Three research report and threats to validity

Promotion Gate Recommendation: PROMOTE WITH CONDITIONS

DECISION // PROMOTE WITH CONDITIONS FOR CONTROLLED ECOSYSTEM EXPERIMENTATION. HowlDream has earned controlled ecosystem experimentation subject to evidence integrity conditions:
  • Condition 1 (Human Review Gate): Obtain genuine independent blinded human review before claiming human-validated superiority or preference.
  • Condition 2 (Frontier Withholding): Withhold claims of a demonstrated empirical frontier until error onset is observed.
  • Condition 3 (Zero-Denominator Rule): Prohibit multiplicative advantage claims (e.g. 3.37x) over zero baseline; report absolute yields.

Authorized next integrations under advisory guardrails:

  • HowlPlane → HowlDream: Exploratory divergence policy for architecture spikes.
  • HowlDream → HowlCreate: Handoff of verified novel candidate pools.
  • HowlDream → HowlFrame: Scoped typed verification contracts for IR assertions.

Keep the experimental record

Each run preserves its input snapshot, source and prompt hashes, candidate outputs, claims, checks, scores and report. New WAKE and replay runs retain lineage. Mock output replay is deterministic. Real-provider replay repeats configuration; identical outputs are not guaranteed. Hosted private reasoning is never captured.

[RUN // ARTIFACT_FILES] HD_ARTIFACTS
hd-TIMESTAMP-ID/
  manifest.json
  experiment.json
  baseline.jsonl
  candidates.jsonl
  claims.jsonl
  verification.jsonl
  scores.json
  metrics.json
  handoff.json
  report.md

SECTION // 05 Authority Boundary

HD-MOD-BOUNDARY
DREAMS CAN SUGGEST. DREAMS CANNOT AUTHORIZE. Dream output is data, not authority. No result deploys software, merges code, sends communications or approves actions. Supported means a named, limited check passed. Uncertainty stays visible.

Trust and privacy model

SECTION // 06 Clear Responsibilities, Explicit Handoffs

HD-MOD-ECOSYSTEM

HowlDream participates in the Howl ecosystem via controlled native integration. Exploration output is data, not authority — speculative candidates can never trigger deployment or bypass human review.

[ARCHITECTURE // NATIVE_ECOSYSTEM_LOOP] [EXPERIMENTAL INTEGRATION]
┌─────────────────────────────────────────────────────────┐
│                        HowlPlane                        │ ◄── Orchestrates exploration within budget
│               (Policy & Circuit Breakers)               │ ─── Enforces candidate & token limits
└───────────────────────────┬─────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────┐
│                        HowlDream                        │ ◄── DREAM / NIGHTMARE / WAKE
│                  (Exploration Engine)                   │ ─── Emits howl.exploration_result/v1 + DescentDAG
└───────────────────────────┬─────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────┐
│                        HowlFrame                        │ ◄── candidate_evaluator.hfbc bytecode
│                  (Bytecode Evaluator)                   │ ─── Validates schema & structural invariants
└───────────────────────────┬─────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────┐
│                       HowlCreate                        │ ◄── Deliberate candidate ingestion
│                 (Sandbox Prototyping)                   │ ─── Emits prototype (EXECUTION_AUTHORITY: NONE)
└───────────────────────────┬─────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────┐
│                        HowlPlane                        │ ◄── Halts exploration loop; prompts human review
│                   (Terminal Boundary)                   │
└─────────────────────────────────────────────────────────┘
                            ┆
   ═════════════════════════╧═════════════════════════ [ HARD SECURITY & AUTHORITY BOUNDARY ] ═════════════════════════
                            ┆
                   [ HowlChangeOps ]  ◄── STRICTLY ISOLATED: Rejects speculative inputs;
                                          never accepts HowlDream or HowlCreate receipts.
HowlPlane

Orchestrates engineering work. Native HowlDreamRunner enforces policy, budgets, and halts before execution.

HowlCreate

Deliberate invention through creative search. Native candidate ingestion produces sandbox prototypes without execution authority.

HowlFrame

Capability-bounded language and runtime. Native candidate_evaluator bytecode validates structural invariants and claims.

HowlChangeOps

Governed release execution. Sole relevant release boundary; hard negative check rejects speculative execution receipts.

HowlRelay

Carries work state and handoffs. Native HowlDreamCollector gathers exploration evidence into the relay store.

HowlWriter

Communicates reviewed knowledge. Claims begin unverified, as in its own domain model.

Milestone Four provides controlled native adapters across HowlPlane, HowlFrame, HowlCreate, and HowlRelay while enforcing zero execution authority. HowlDream is not included in the default Howl installer.

Explore the Howl ecosystem

SECTION // 07 Start Locally

HD-MOD-SETUP

A small, inspectable experiment. No API key required for the deterministic demo. Optional Ollama and compatible HTTP providers require explicit configuration.

[BASH // HOWLDREAM_QUICKSTART]
git clone https://github.com/howlcipher/howldream.git
cd howldream
python3 -m venv .venv
source .venv/bin/activate
pip install -e .
howldream run examples/self_detection.yaml
howldream benchmark

Examples · CLI reference · Roadmap