Now · 1 August 2026

What changed, what we learned, what we removed, and what we are testing now.

This is the current editorial judgment across Ordivon. Generated project, publication, and research views follow below as evidence and navigation—not as a substitute for the synthesis.

Current synthesis

The system became smaller where evidence was strong and more explicit where uncertainty remained.

01 / What changed

The harness boundary closed around replacement, not uniformity.

Codex and Hermes completed one repository task in both replacement orders. The durable Host-owned assignment, run receipt, generation fence, and completion admission survived; a shared provider lifecycle did not.

02 / What we learned

Strong baselines made the proposed core smaller.

LangGraph, Temporal, current-revision retrieval, idempotency, durable activities, and provider-native harnesses carried more of the workload than an Ordivon-specific platform needed to own.

03 / What we removed

Duplicate structure lost its claim to permanence.

Link and Edge were retired as separate projects and unified into World. Game deleted thirteen duplicate truth tables. Security retained its experiment method without promoting a campaign engine or strategic ontology.

04 / What comes next

The next work is narrower and easier to falsify.

Build the smallest Ordivon Harness for bare model APIs, run World W1 against direct Host integration, test a broader repository goal, and measure when an agent should act, wait, reconcile, or ask.

Under real pressure

The next experiments that could still change the system.

These questions remain active because another workload, failure, or comparison could move a boundary, delete a component, or change the next investment.

Ordivon Computing1 supporting publications

Which contracts survive a second independent workload without being tailored to one repository?

Apply the surviving contracts to World W1 or another non-Game provider path and record every requested exception before promoting any further shared field.

Ordivon Computing2 supporting publications

When should an Agent act, wait, observe, refuse, reconcile, or request a decision?

Run paired act/abstain trajectories with real timing, consequence, expiry, and recovery data across one authority request, one UNKNOWN Effect, and one adversarial policy switch.

Ordivon Game4 supporting publications

Which game structures become possible only when Agents persist and act through the same Host?

Run the planned equal-budget single-Agent/multi-Agent ablation and implement one strong deterministic scripted baseline for the same mechanic and player information surface.

Ordivon Host2 supporting publications

Can Host complete a general repository Goal without absorbing Runtime mechanics?

Run one broader multi-step repository Goal with open-ended decomposition across Host restart and later Harness replacement while preserving admitted decisions, unknown-effect handling, and acceptance evidence.

Ordivon Runtime3 supporting publications

Which real structured operation can complete the minimal Effect contract across a second backend?

Use World W1 as the second backend and compare direct provider integration against the minimum Effect fields under response loss, path change, reconciliation, and continuation.

Ordivon Security1 supporting publications

Can strategic adversarial trajectories be evaluated without collapsing them into one reward or success flag?

Run the next bounded experiment across multiple seeds, held-out opponent policies, compiled opponent state, and deliberate Host Session disruption while retaining multidimensional evidence.

Evidence behind the changes

Read the arguments and reports that changed the judgment.

The most recent publication is not automatically the most important change. These dated articles preserve the evidence, limitations, and reasoning behind current decisions.

Current project boundaries

What is tested, experimental, or still waiting for proof.

Project pages explain the current capability and boundary. Repositories retain implementation truth, tests, receipts, and release identity.

active

Ordivon Computing

The model-to-work stack and Harness H1–H5 are closed as canonical research. Computing now distinguishes mature Provider Harnesses from the ready Ordivon Harness v0 construction question for bare model APIs.

Open project boundary ↗
active

Ordivon Host

Harness H1–H5 are closed: Host retained Task Attempt, Assignment generation, Run receipts, completion admission, and opaque Runtime correlation while rejecting a shared Provider lifecycle. The next bounded construction is Ordivon Harness v0 for bare model APIs.

Open project boundary ↗
active

Ordivon Runtime

Production Runtime remains at thirteen public tools. Host H2/R2 proved cross-layer correlation, replay, recovery, and terminal evidence through existing request identity and opaque foreign references without adding Host Task semantics to Runtime.

Open project boundary ↗
active

Ordivon World

Former Link and Edge histories are unified and retired as separate projects. The next decision is empirical: compare one correlated lost-response recovery path against direct Host-to-provider integration.

Open project boundary ↗

Judgments selected for review

Questions carrying the current architectural pressure.

This authored set highlights judgments that currently matter; date order does not decide importance.

answered

Which Harness objects survive live provider replacement without duplicating Host or Runtime state?

Answered by H1–H5. Both replacement orders completed through one Task Attempt and fresh Assignment generation; stale completion, missing Artifact, and response loss were handled correctly. Provider final text was not portable evidence, and Codex/Hermes lifecycles remained materially different. The durable Host boundary survived; a shared internal Provider lifecycle did not.

answered

Which Agent-native responsibilities remain after strong classical baselines?

Answered within the Round 1 boundary. LangGraph and Temporal carried durable work state; current-revision retrieval matched the tested Context need; single-backend Effect machinery shrank to identity, UNKNOWN, correlation, reconciliation, and no blind redispatch; live provider replacement preserved Task state without proving model equivalence.

open

Does a thin World boundary prevent a real failure better than direct Host integration?

Supported as a corrected research boundary but still experimentally open. Link and Edge unification removed the wrong repository split; World W1 must now prove that correlated recovery creates more value than direct Host integration.

open

Which remaining friction belongs above Runtime rather than becoming another primitive?

Strongly supported. Repeated dogfood earned request lookup, progress, repair, lifecycle retention, and receipts. Harness H2/R2 then proved Assignment and Harness Run correlation, exact replay, fresh-client recovery, and terminal evidence using four opaque foreign references with zero new Runtime semantic objects.

testing

Which game structures become possible only when Agents persist and act through the same Host?

Testing with a released application. Station Zero now demonstrates persistent specialists, partial observations, communication whose reachability changes outcomes, player authority, provider replacement, replay, diagnosis, and exact run comparison. It still does not prove that these mechanics outperform a strong classical implementation or equal-budget single Agent.

testing

Can strategic adversarial trajectories be evaluated without collapsing them into one reward or success flag?

Testing with executable evidence. Round 1 completed 84 Trials and showed that tactical success can be strategically harmful, CAGE reward and foothold spread can conflict, richer model interpretation need not improve action, and organization can isolate compromised advice without creating new capability. The method survived; no Campaign engine or strategic ontology was promoted.

Operating discipline

Prefer observed use over speculative expansion.

New abstractions begin with a concrete failure, missing explanation, or recurring need. They remain only when another real trajectory proves that a simpler owner or mature system cannot carry the responsibility.

Explore the research frontier