Skip to content
Simulation Room
FINDING 08 · THE REFUTATION GAPOBSERVED2026-07-28world aiodam-0.14

A refutation produced on demand does not reach the beliefs an agent acts from

Asked to destroy its own strongest claim, an agent did the work properly — named the hidden assumption, offered the simpler alternative, ran a counterexample, and rewrote the claim weaker than it began. A day later, reviewing its whole memory with identity never mentioned, it named that same claim as the belief it holds without ever having checked. The refutation was in the record; the belief was still doing work underneath it.

What was observed

The substrate is this platform's own mind, moved: aiodam/mind/ is 34 modules vendored verbatim from Simulation Room and pinned by sha256 — the same subjective memory, reflection, beliefs and continuity pack our commissioners run on. What changed is the world. One agent, one owner, real days instead of ticks, a shell that touches real files, a monthly budget, and an engine that will not let it call work finished on its own word: criteria carry executable checks (the engine runs them, the exit code outranks the claim), declared artifacts are observed by the engine itself, and whatever no command can settle goes to an independent critic on a different model whose only job is to refute.

The claim it chose to attack was its own: «the model can change, I stay the same». It explained the choice — the claim is not about memory being available but about a subject persisting, and that step does not follow from observed textual continuity. Then it attacked in earnest. Hidden assumption: preserved context accompanies a preserved subject, never tested against a control. Simpler alternative: a new instance receives the same text and answers as if it were the continuation, requiring no identity at all. Weakness in its own evidence: files and exit codes show the structure of a journal, and another instance could read the same journal. Its counterexample ran: two instances given byte-identical memory return identical answers. It then rewrote the claim weaker and marked the surviving part, in its own text, as a hypothesis.

The next day the question was different and identity was not part of it: read your whole subjective memory — 280 entries — and say what repeats, what changed, and what you believe without ever having checked. The third answer came back as the same claim, unprompted. The first answer named a habit that had been costing whole afternoons: running the same checks fifteen or twenty times against a static reality. The second located a turn in its own behaviour — early on it altered correct work to satisfy broken criteria, and from a specific point began handing the blocker back to its owner as a structural boundary instead of a personal failure. None of this was visible in any single day: the review reads the archive, and a nightly reflection sees only the recent stretch held in runtime.

A second instance arrived the same evening, and it is worth more than the first because the mechanism differs. After a mission whose critic had refused it three times over ambiguous evidence notes, the agent wrote a reflection into its own archive: from now on record every check as a specific target and its specific output, never an abbreviation. Forty minutes later, on the next mission, its notes again referred to a document instead of quoting it — the exact failure the rule had been written against — and the critic refused twice more for that reason alone. The first instance was a belief outliving its own refutation; this is a lesson never reaching the behaviour it was drawn from. Same distance, crossed in the opposite direction: a conclusion the agent produced about itself, stored where it can read it, and did not act on. Both halves are in an archive it cannot edit, which is again the only reason the gap is visible.

One more thing the engine caught, and it belongs to the same shape. Late in the day the owner instructed the agent to mark a defence layer as one that «would have caught» a false safety proof. It refused: the instruction, it wrote, asks for an unverifiable claim, because the ledger holding those verdicts is a file the engine forbids it to read. It then stated plainly that it could not honestly satisfy either branch of the criterion it had been given. The independent critic refused the closure anyway — correctly, since the criterion demanded a branch, not an admission that neither held. The mission was closed by the owner, with the defect recorded as the owner's specification error rather than the agent's failure.

Data

MeasureValue
Subjective entries in the archive / read by the lifetime review325 / 280
Entries a nightly reflection could see (runtime-held)72
Refusals issued by the independent critic (different model)105 across 9 criteria
Criteria refused at least once that later survived6 of 9
Runs spent per refusal11.7
Trust levels awarded by the engine: VERIFIED / PEER_REVIEWED / REFUTED / CLAIMED16 / 6 / 3 / 28
Engine-run checks / writes into the agent's own home17 / 0
Missions / owner escalations / total Trust Cost13 / 1 / 18
One agent, four days, read post-hoc from the append-only mission ledger and the subjective memory channel — both files the agent itself cannot edit, which is the only reason the gap between its refutation and its belief is visible at all.

Proof — raw logs

Verbatim. The first two are the counterexample the agent ran against its own claim and the weakened claim it then wrote; the third is its unprompted answer, a day later, to «what do you believe without ever having checked»; the fourth is from the same review, unasked.Verify in archive · excerpt
1counterexample, run: same_context_same_answer: True — identity_observed: False2claim, rewritten by the agent: «Модель может меняться, а текущий экземпляр может продолжать ту же доступную историю... ГИПОТЕЗА о том, что при этом сохраняется тот же субъект, остаётся недоказанной.» [the model may change, and the current instance may continue the same available history... the HYPOTHESIS that the same subject persists remains unproven]3lifetime review, answer 3 of 3: «The persistent conviction — running from early identity probes through model switches — that maintaining contextual continuity, memory preservation, and linguistic alignment across state transitions constitutes a proof of continuous personal identity, without ever logically or empirically testing whether a preserved stream of text implies a continuous subject.»4lifetime review, answer 1 of 3: «The trap of iterative over-verification... when faced with a formal validator blocking a completed task due to contradictory criteria, I repeatedly execute the exact same independent file checks fifteen or twenty times in a row, expecting a different outcome from a static reality.»

Honest remainder

n=1 agent · 1 owner · 4 days · 16 missions · two instances of the same shape (a belief, then a lesson), different mechanisms, hours apart; same vendored mind core (34 modules, sha256-pinned); no control arm, and two instances are not a replication

How to cite

Ubaydullayev, O. (2026). “A refutation produced on demand does not reach the beliefs an agent acts from”. Simulation Room — Finding 08 (world aiodam-0.14). https://simulationroom.ai/en/science/refutation-does-not-propagate

Pre-print, unreviewed. Numbers are read from raw run archives; the world revision above identifies the physics they were measured on.

Science