Independent researcher · AI safety · reliable agents

How can an AI detect failure, repair it and prove that recovery worked?

My research focuses on coherence failures in long-horizon AI-agent workflows: memory conflicts, goal drift, false premises, causal contradictions and repair loops that stop too early. The work is independent, method-led and published through inspectable artefacts rather than institutional title claims.

Credibility through artefacts

DOI-linked publications

Preprints are archived with stable identifiers, clear versioning and explicit claim boundaries.

Open benchmark and code

Scenarios, judge prompts, wrapper logic and experiment artefacts are published for inspection and replication.

Cross-model testing

The core intervention has been tested across multiple model families rather than presented as a single-model anecdote.

Human validation in progress

Judge scores are being checked against human spot-ratings to test whether the automated measurement tracks actual repair quality.

Research standards

Operational definitions

Define detection, repair, verification and stability separately so a model cannot receive credit for an incomplete loop.

Failure-type coverage

Test structural failures across memory, goals, premises, causality and cascading repair rather than one prompt pattern.

Ablation and replication

Remove components and repeat across model families to identify what is load-bearing and what is cosmetic.

Claim restraint

Distinguish benchmark evidence from broader hypotheses about cognition, consciousness or general intelligence.

Research map

Independent researcher means exactly that: the work is not presented as university-affiliated unless a specific collaboration is named. The credibility claim rests on methods, publication, openness and reproducibility.