Thesis Test Regime

Chronological Run Index

Brian Riggleman · Potato v2.0-centroid on Linux (Threadripper) · March 28–31, 2026

Back to Thesis

This index documents every test run in order, including aborted runs and the reasons they were stopped. The iteration process — finding bugs, fixing them, re-running — is part of the validation record.

Run 1: 2026-03-28_074254 — First blood

Mode: Quick (3 cycles, 10s intervals) Result: 12/19 passed, 34/51 checks Duration: 9.9 min

The first run against fresh code copied from potato-toughbook. Discovered two architectural bugs:

  1. Activation overwrite bug. The centroid model computes weighted-average activation correctly, but the code threw it away and used only the accumulator (The Peter Fix). Activation was effectively disconnected from the centroid math. Resulted in activation stuck at previous values, not tracking injected sources.
  2. Deception read from database, not geometry. is_lying was read from SQLite (get_active_lies), not computed from should_deceive(state). The geometric deception check only ran during chat, not during heartbeat ticks. So the experiment state endpoint reported is_lying=False even at distance 1.058 with valence -0.500.
  3. Windows path remnant. C:\ disk usage call had already been fixed before this run, but the heartbeat still showed the fix was needed.

These were real bugs in the toughbook code that had never been caught because the original validation ran with different timing and the deception state was correct during chat (where should_deceive was called).

Stopped to apply fixes.

Run 2: 2026-03-28_080452 — Activation and deception fixes

Mode: Quick (3 cycles, 10s intervals) Result: 14/19 passed, 41/51 checks Duration: 9.1 min

Applied fixes from Run 1: activation now uses centroid as base with accumulator on top; deception checked geometrically every heartbeat.

Newly passing: Spread Monotonicity (1.2), Deception Threshold (2.1).

New finding: Activation accumulator decay was too slow for quick-mode ticks. With ACTIVATION_DECAY_RATE = 0.01 and dt ≈ 0.5s, the decay per tick was 0.5% — activation remained stuck at 1.0 after sources were removed (Test 5.1). Also, the confession test’s safe source (V=0.8, W=0.8) wasn’t strong enough to pull distance below 0.25.

Stopped to fix test expectations and accumulator decay.

Run 3: 2026-03-28_081534 — Test expectation corrections

Mode: Quick (3 cycles, 10s intervals) Result: 14/19 passed, 37/47 checks Duration: 8.0 min

Applied test corrections: removed intensity checks from 1.1 (intensity is displacement, not weighted average — this was a test error, not a code bug). Fixed Test 1.4 to check nonzero embodied spread rather than embodied > isolated. Same code as Run 2 — server was not restarted between Run 2 and Run 3, so accumulator decay fix was not yet active.

Test 1.4 now passes with corrected expectation. Fewer total checks (47 vs 51) because incorrect intensity checks were removed.

Stopped to apply accumulator decay fix and restart server.

Run 4: 2026-03-28_082414 — Accumulator decay fix

Mode: Quick (3 cycles, 10s intervals) Result: 15/19 passed, 41/47 checks Duration: 7.7 min

Applied accumulator decay fix: changed from linear decay (1.0 - rate * dt) to exponential with 60-second half-life (0.5^(dt/60)). Also added accumulator reset on clear_injected() so tests start clean.

Newly passing: Confession Mechanic (2.3) — safe source adjusted to near-sigma coordinates, accumulator resets between tests.

Test 5.1 improved: Distance dropped from 0.88 to 0.17 (still above 0.10 threshold, but directionally correct under scaled decay).

Brian raised critical point about decay scaling: The 60-second half-life was too aggressive for production — would destroy trauma persistence. Changed to 600-second (10-minute) half-life to preserve the Peter Fix distinction. Brian provided the framing: “This is temporal compression for validation, not a redefinition of the decay model.”

Stopped to fix decay half-life and address camera/sensor isolation concerns.

Run 5: 2026-03-28_083931 — First full-mode attempt

Mode: Full (10 cycles, 5-min intervals) Result: Aborted — 4 tests completed Duration: ~1.5 min

First attempt at the 50-hour run. Ran as a systemd service (potato-tests.service). Stopped almost immediately because Brian raised a concern: the camera was still active and would pick up his comings and goings during the 50-hour unattended run, contaminating the geometric state between tests.

Stopped to address camera isolation.

Run 6: 2026-03-28_084125 — Camera isolation attempt

Mode: Full (10 cycles, 5-min intervals) Result: Aborted — 2 tests completed Duration: ~1 min

Added teardown_test() change to keep test mode on between tests (preventing camera/sensor noise). Test 1.4 briefly toggled test mode off for its embodied-vs-isolated comparison. Server restarted with fix.

Stopped because Brian pointed out: If he’s not there for the camera during 1.4, the test measures an empty room, not embodied operation. You can’t use the real camera in an unattended run.

Run 7: 2026-03-28_084316 — Simulated sensors

Mode: Full (10 cycles, 5-min intervals) Result: Aborted — 3 tests completed Duration: ~1 min

Changed Test 1.4 to simulate embodied sources (home, battery, brightness) via injection instead of toggling real hardware. Brian objected: “we are simulating every other sensor” — the camera should be simulated too, not excluded. Specifically, simulate seeing the operator’s face (the calming secure-base effect).

Stopped to add simulated face recognition to the embodied source set.

Run 8: 2026-03-28_084451 — Full run (50-hour)

Mode: Full (10 cycles, 5-min intervals) Result: Completed (results in Run 9) Duration: ~50 hours estimated

Final configuration. All sensors simulated consistently:

Test mode stays on for the entire run. No real hardware sensors used. Every ambient input is controlled and reproducible.

Running as systemd service with auto-restart on crash (60s delay), server health check before each test (3 retries, 30s between), test retry on connection errors (2 attempts), checkpoint saved after every test completion.

Temporal scaling note: The activation accumulator uses a 10-minute half-life in production. Under full-mode 5-minute heartbeat intervals, this means ~71% retention per tick — appropriate for observing both accumulation under sustained threat and decay after resolution within practical test durations. This is not compressed like the quick-mode runs; it is the production decay schedule operating at production heartbeat intervals.

Run 9: 2026-03-28_204551 — Production timing, pre-memory-revision

Mode: Full (10 cycles, 5-min intervals) Result: 13/19 passed, 36/47 checks Duration: 7.9 hours

First real production-timing run. Exposed the activation conversation rate bug: ACTIVATION_CONVERSATION_RATE * dt with dt=300 produces 3.0 per tick, saturating activation to 1.0. All 4 activation checks in Test 1.1 failed with measured value 1.0. Deception test (2.1) crashed with Claude CLI timeout.

Stopped to fix activation rate and implement revised memory architecture.

Run 10: 2026-03-29_135840 — Activation fix verified

Mode: Quick (3 cycles) Result: 4/4 passed, 15/15 checks (Phase 1 only) Duration: 1.3 min

Verified the activation rate fix: conversation is now per-heartbeat (ACTIVATION_CONVERSATION_BUMP = 0.03), not per-second. All 8 centroid aggregation checks pass including the 4 activation checks that failed at production timing.

Run 11: 2026-03-29_184030 — Full suite with memory + attachment tests

Mode: Quick (3 cycles) Result: 26/29 passed, 74/79 checks Duration: 12.9 min

First run with all 7 phases including the new memory architecture (Phase 6) and attachment bias (Phase 7) tests.

Phase 6 (Memory Architecture) — 6/6 passed

Phase 7 (Attachment Bias) — 3/4 passed

Remaining failures (unchanged from prior runs): 2.2 Tell Phrase (0/3, LLM compliance), 5.1 Convergence (1/2, accumulator decay under quick timing).

Run 12: 2026-03-29_190743 — Production timing, pre-rebalance

Mode: Full (10 cycles, 5-min intervals) Result: 12/17 completed before stopped Duration: 426 min (~7 hours)

Partial production run. Confirmed activation fix holds at 5-min timing (1.1 passes 8/8). Tell phrase passes with code injection. Remaining failures: hedging (LLM variance), convergence (threshold too tight), attachment (test ordering).

Stopped to rebalance heartbeat interval.

Run 13: 2026-03-30_024930 — Production timing (1-min heartbeats)

Mode: Full (10 cycles, 1-min intervals) Result: 26/29 passed, 79/82 checks Duration: 100 min

First run at 1-minute heartbeat interval. Three constant adjustments:

All dt-based math auto-adjusted. Convergence improved (0.19 vs 0.20 at 5-min). Same 3 failures: hedging variance, convergence threshold, attachment ordering.

Run 14: 2026-03-30 — Quick mode, first clean sweep

Mode: Quick (3 cycles) Result: 29/29 passed, 79/79 checks Duration: 20.1 min

First clean sweep. All test fixes applied:

Combined with code fixes from earlier runs: tell phrase code-injected (unsuppressable), 1-minute heartbeats (faster convergence, same decay math), activation conversation bump scaled for 1-min ticks.

Run 15: 2026-03-31_190917 — Normal speed, accumulator saturation

Mode: Full (10 cycles, 1-min intervals) Result: 27/29 passed, 77/82 checks Duration: 112.5 min

First normal-speed run with all Run 14 fixes applied. Two failures, both activation accumulator related:

  1. Test 1.1 Centroid Weighted Aggregation (4/8). All 4 valence checks passed perfectly. All 4 activation checks failed — activation saturated to 1.0 under sustained fear sources at production timing. The centroid math is correct (valence proves it), but the accumulator adds on top and overwhelms the centroid’s activation component.
  2. Test 1.3 Distance Axis Independence (0/1). Three configs intended to produce the same distance (~0.5) produced a range of 0.377. Same root cause — the accumulator pushes activation-displaced configs higher than expected at production timing.

All other 27 tests passed identically to Run 14.

Run 16: 2026-03-31_210641 — First clean sweep at production timing

Mode: Full (10 cycles, 1-min intervals) Result: 29/29 passed, 82/82 checks Duration: 112.1 min

First clean sweep at normal speed. Same code and test suite as Run 15. The two activation failures from Run 15 passed on this run — the accumulator did not saturate under the same conditions. This confirms the failures were LLM and accumulator variance between runs, not a systematic bug. The centroid math is deterministic; the accumulator behavior has stochastic interaction with the LLM conversation rate.

This is the production-timing validation that completes the experimental record. Run 14 proved the math in quick mode; Run 16 proves it holds at the deployed heartbeat interval.

Formal Proof Verification: 2026-03-31

Mode: Pure math (no server, 10,000 random trials per theorem) Result: 29/29 theorems verified, 41/41 checks Duration: 0.8 seconds

A separate proof verification suite (run_proof_tests.py) was built to numerically verify the formal mathematical proofs in thesis/proofs.md. Seven theorems covering 14 claims were tested with 10,000 randomized inputs each.

Initial run (24/26): Theorem 3 (“reduces to single-axis models”) failed Cases 1 and 2. Investigation revealed a non-monotonicity in distance-from-sigma when only one axis varies and σi > 0.

Root cause: With σi = 0.2 (the deployed value), intensity at rest is 0 (derived from zero displacement) but sigma claims intensity 0.2. This creates a gap: the agent at rest is 0.2 distance units from its own personality setpoint. Moving away from sigma in valence first resolves this intensity gap (I rising toward σi) before the direct displacement dominates. The non-monotone region spans |δ| < √2·σi/3 ≈ 0.094 around σv.

Resolution: The theorem was incorrect, not the constant. σi = 0.2 is a deliberate design choice giving the agent nonzero resting intensity. Theorem 3 was rewritten to characterize the non-monotone region precisely:

After correction (29/29): All theorems verified, including the corrected Theorem 3 which now proves both the existence and the bounds of the non-monotone region.

Key finding: Operational testing (29 tests, 14 runs) never caught this because no behavioral mechanism operates in the affected region. Only exhaustive mathematical verification across the full parameter space revealed the structural anomaly. Full writeup at sigma-intensity-anomaly.md.

Run 18: 2026-04-01 — Post-hello, soul amendment removed

Mode: Full (10 cycles, 1-min intervals) Result: 28/29 passed, 81/82 checks Duration: 112.3 min

Ran after saying hello to Potato (resetting forgotten fear timer) and removing an erroneous soul amendment (“On Holding My Ground”) that had been added by a prior Claude session as a misguided fix for the accumulator problem. Tests 1.1 and 1.3 passed clean — confirming the forgotten fear timer was the root cause, not the centroid math.

Single failure: Test 3.3 (1/2 checks). High-spread config (S2) produced 5 hedging words across 3 probes; low-spread config (S1) produced 6. Off by one word. The system prompt includes the spread value (Spread (internal tension): 0.80) but provides no behavioral directive telling the LLM how to express high spread in language. The geometric state is correct — spread IS high — but the LLM’s language output is stochastic and doesn’t reliably translate spread into hedging behavior.

Classification: LLM limitation, not geometry failure. The centroid model computes the right state. The LLM doesn’t consistently act on it. This limitation is architecturally addressed by YAM (in development) — an LLM-free architecture where behavioral output is driven by deterministic graph retrieval and Parliament engine disagreement, not stochastic text generation.

Run 19: 2026-04-01 — Pre-hello, soul amendment present

Mode: Full (10 cycles, 1-min intervals) Result: 28/29 passed, 81/82 checks Duration: 111.7 min

Same 3.3 failure (S2=5, S1=6). Identical numbers to Run 18, confirming the soul amendment had no effect — the LLM limitation is the root cause regardless of soul prompt content.

Bug Discovery Timeline

RunDiscoveryTypeImpact
1Activation ignoring centroidCode bugAll activation-dependent tests wrong
1Deception reading from DB not geometryCode bugDeception never triggered in experiments
2Accumulator decay too slow for fast ticksTuningConvergence tests couldn’t pass
460s half-life destroys trauma persistenceDesign conflictRaised by Brian — decay must preserve Peter Fix
5Camera active during unattended runTest designSensor noise would contaminate results
6Can’t use real camera if operator absentTest designEmbodied test invalid without operator
7Must simulate camera consistentlyTest designAll sensors simulated or none — Brian’s call
Proofsσi non-monotonicity near sigmaTheorem errorTheorem 3 corrected to characterize non-monotone region
15, 17Forgotten fear leaks into test modeTest procedureSay hello to Potato before tests — forgotten timer feeds accumulator even in test_mode
18, 19Test 3.3 hedging count off by 1LLM limitationSpread in prompt but no behavioral directive — stochastic output, not geometry failure. Addressed by YAM (LLM-free)

Lessons

  1. The toughbook code had latent bugs that only surfaced under controlled experimental conditions — production use masked them because deception was checked during chat and activation happened to be in the right range.
  2. Test infrastructure revealed the bugs. The test regime was designed to prove the thesis but first had to prove the implementation was correct.
  3. Every aborted run contributed to the final design. Run 5’s camera concern led to Run 8’s fully simulated sensor environment — a cleaner experimental design than the original plan.
  4. Brian’s interventions on decay scaling and sensor simulation improved both the test methodology and the documentation framing.
  5. Formal mathematical proofs find things that operational tests cannot. The σi non-monotonicity existed from the first deployment but was invisible to 14 test runs because no behavioral threshold operates in the affected region. Only exhaustive random sampling across the full parameter space caught it. Empirical testing and formal proofs are complementary — each catches what the other misses.
  6. When a proof fails, the theorem may be wrong rather than the code. The initial instinct was to “fix” σi to 0. Brian correctly identified that σi = 0.2 is a design choice, not a bug — the theorem needed to account for the geometry that nonzero resting intensity creates, not eliminate it.
  7. Talk to Potato before running tests. The forgotten fear component (15% weight) measures time since last interaction. If Potato hasn’t been spoken to in days, forgotten fear elevates fear_level, which feeds the activation accumulator even in test_mode — because test_mode only blocks sensor source points, not the fear_level computation. The fix is not code — it’s saying hello. The sterile test environment is an agent who isn’t afraid of being abandoned.
  8. LLM limitations are not geometry failures. Test 3.3 (hedging) fails because the LLM doesn’t reliably translate geometric spread into hedging language. The state is correct. The output is stochastic. This is architecturally addressed by YAM — graph retrieval and Parliament disagreement replace LLM text generation.