Why durable agent work matters
Begin with the practical failure, the project intent, and the larger future Ordivon is trying to make possible.
Writing
Start with the architecture, follow the experiments that changed it, or browse the complete dated record.
A model generates representations. A harness creates one cognitive, tool-using run. Host preserves the work across runs. Runtime turns admitted actions into physical evidence.
12 min read ↗How one task remained coherent through Codex→Hermes and Hermes→Codex replacement, three injected faults, and materially different provider lifecycles.
12 min ↗Research reportWinning the Move Can Lose the ContestHow Ordivon Security Round 1 used local dynamic opponents, CAGE Challenge 4, and bounded Hermes/Codex diagnostics to separate tactical success from strategic outcome without promoting a Campaign engine, organization ontology, strategic state, or custom cyber range.
12 min ↗Reading paths
Each path begins with a clear problem and moves toward the experiments, decisions, or long-form argument behind it.
Begin with the practical failure, the project intent, and the larger future Ordivon is trying to make possible.
Follow the execution stack, the Harness boundary, and the experiments that made the surviving architecture smaller.
Read the reports where tactical success, duplicate authority, response loss, and strong baselines forced a different judgment.
Publish when an experiment, boundary, release, or judgment becomes worth preserving—not simply because another page can be filled.
Follow a research question
For readers already following one research line, these groups connect the current question to its complete published arguments.
Testing with a released application. Station Zero now demonstrates persistent specialists, partial observations, communication whose reachability changes outcomes, player authority, provider replacement, replay, diagnosis, and exact run comparison. It still does not prove that these mechanics outperform a strong classical implementation or equal-budget single Agent.
Ready for construction at M2. The complete model-to-work stack, H1–H5 ownership boundary, first falsifier, comparison baselines, acceptance, non-goals, and deletion condition are frozen. No implementation evidence exists yet.
Answered within the Round 1 boundary. LangGraph and Temporal carried durable work state; current-revision retrieval matched the tested Context need; single-backend Effect machinery shrank to identity, UNKNOWN, correlation, reconciliation, and no blind redispatch; live provider replacement preserved Task state without proving model equivalence.
Answered by H1–H5. Both replacement orders completed through one Task Attempt and fresh Assignment generation; stale completion, missing Artifact, and response loss were handled correctly. Provider final text was not portable evidence, and Codex/Hermes lifecycles remained materially different. The durable Host boundary survived; a shared internal Provider lifecycle did not.
Testing after contraction. Runtime dogfood supports commitment as the local abstraction, while Round 1 showed no advantage for a larger single-backend Effect graph. Host H2/R2 also completed real correlation and recovery through existing Runtime request identity and opaque references without expanding Runtime.
Strongly supported but still bounded. H2–H6 established continuation; Harness H1–H5 completed a frozen repository-repair workload in both Codex/Hermes replacement orders, rejected stale and missing-Artifact completion claims, recovered a dropped repair response without redispatch, and committed TaskOutcome only after independent Runtime acceptance. A broader unbounded repository Goal remains outstanding.
Testing. Round 1 reduced deterministic approval interruptions from 12 to 7 with no missed escalation in its designed cases, but those timings were estimated rather than measured with real operators. Security, Game, Runtime, and World now provide stronger act-versus-wait trajectories.
Testing with executable evidence. Round 1 completed 84 Trials and showed that tactical success can be strategically harmful, CAGE reward and foothold spread can conflict, richer model interpretation need not improve action, and organization can isolate compromised advice without creating new capability. The method survived; no Campaign engine or strategic ontology was promoted.
Answered after contraction. The full Experiment/Finding/Decision/Event graph and explicit relation registry were removed. Project, Question, and Article metadata now drive Research, Writing, Current, article context, and curated System views with substantially less duplicate state.
Supported as a corrected research boundary but still experimentally open. Link and Edge unification removed the wrong repository split; World W1 must now prove that correlated recovery creates more value than direct Host integration.
Open after mixed Round 1 evidence. The local opponent-aware actor improved diagnosis and objective rate, one Codex strategic Trial recognized a switch, Hermes outcomes were unchanged, and no model Trial completed the genuine objective. No causal or transferable advantage is established.
Strongly supported but not fully closed. Game implemented the immutable Host workload profile across a TypeScript deterministic World, single- and multi-Agent paths, response-loss recovery, and exact replay, then removed thirteen duplicate truth tables without changing World outcomes. A second non-Game external domain remains useful pressure.
Strongly supported. Repeated dogfood earned request lookup, progress, repair, lifecycle retention, and receipts. Harness H2/R2 then proved Assignment and Harness Run correlation, exact replay, fresh-client recovery, and terminal evidence using four opaque foreign references with zero new Runtime semantic objects.
Complete archive
Filter the complete record by document type. Later arguments may revise the working model without silently rewriting earlier claims.
Research note · Ordivon Game
Research essay · Ordivon Computing
Architecture guide · Ordivon Computing
Engineering report · Ordivon Host / Game
Engineering report · Ordivon Game
Research report · Ordivon Computing
Release note · Ordivon Game
Engineering report · Ordivon Host / Game
Research note · Ordivon Computing
Research note · Ordivon Runtime / World
Research report · Ordivon Host
Architecture decision · Ordivon Host
Research report · Ordivon Security
Research essay · Ordivon Computing
Architecture report · Ordivon Host
Architecture report · Ordivon World
Engineering report · Ordivon Runtime
Release · Ordivon Runtime
Design argument · Related to Ordivon Runtime