Research · 1 August 2026

Research here exists to change what we build.

Each question is tied to a decision that another experiment can still overturn: keep a boundary, narrow it, move it to a mature system, or delete it.

replacement orders completed
2 / 2
duplicate game tables deleted
13
bounded adversarial trials
84
dated public arguments
19

Start here

Three research results define the current frontier.

Begin with the work that can still change the system, the boundary most recently answered, and the experiment that removed the most proposed machinery.

Full research index

Browse every active, open, answered, and reframed question.

Use the explorer after the current priorities are clear. It groups the same research by question, project, publication date, or status.

Research atlas

The same work, organized by the question you need to answer.

01Ordivon Computingtesting

Which contracts survive a second independent workload without being tailored to one repository?

Use unrelated applications to distinguish general contracts from local accommodations.

Current judgmentStrongly supported but not fully closed. Game implemented the immutable Host workload profile across a TypeScript deterministic World, single- and multi-Agent paths, response-loss recovery, and exact replay, then removed thirteen duplicate truth tables without changing World outcomes. A second non-Game external domain remains useful pressure.
1 supporting publications2026-07-31
02Ordivon Computingtesting

When should an Agent act, wait, observe, refuse, reconcile, or request a decision?

Treat non-action as a timed, evidence-bearing decision rather than the absence of Agent capability.

Current judgmentTesting. Round 1 reduced deterministic approval interruptions from 12 to 7 with no missed escalation in its designed cases, but those timings were estimated rather than measured with real operators. Security, Game, Runtime, and World now provide stronger act-versus-wait trajectories.
2 supporting publications2026-07-31
03Ordivon Gametesting

Which game structures become possible only when Agents persist and act through the same Host?

Separate genuinely Agent-native mechanics from classical systems with an LLM attached.

Current judgmentTesting with a released application. Station Zero now demonstrates persistent specialists, partial observations, communication whose reachability changes outcomes, player authority, provider replacement, replay, diagnosis, and exact run comparison. It still does not prove that these mechanics outperform a strong classical implementation or equal-budget single Agent.
4 supporting publications2026-07-31
04Ordivon Hosttesting

Can Host complete a general repository Goal without absorbing Runtime mechanics?

Test semantic continuity on real engineering work while preserving the Host–Runtime boundary.

Current judgmentStrongly supported but still bounded. H2–H6 established continuation; Harness H1–H5 completed a frozen repository-repair workload in both Codex/Hermes replacement orders, rejected stale and missing-Artifact completion claims, recovered a dropped repair response without redispatch, and committed TaskOutcome only after independent Runtime acceptance. A broader unbounded repository Goal remains outstanding.
2 supporting publications2026-07-31
05Ordivon Runtimetesting

Which real structured operation can complete the minimal Effect contract across a second backend?

Find one operation that proves stable identity, explicit uncertainty, correlation, reconciliation, and no blind redispatch beyond local process execution.

Current judgmentTesting after contraction. Runtime dogfood supports commitment as the local abstraction, while Round 1 showed no advantage for a larger single-backend Effect graph. Host H2/R2 also completed real correlation and recovery through existing Runtime request identity and opaque references without expanding Runtime.
3 supporting publications2026-07-31
06Ordivon Securitytesting

Can strategic adversarial trajectories be evaluated without collapsing them into one reward or success flag?

Model adaptive opposition through separate World truth, actor observation, decisions, effects, and multidimensional outcomes.

Current judgmentTesting with executable evidence. Round 1 completed 84 Trials and showed that tactical success can be strategically harmful, CAGE reward and foothold spread can conflict, richer model interpretation need not improve action, and organization can isolate compromised advice without creating new capability. The method survived; no Campaign engine or strategic ontology was promoted.
1 supporting publications2026-07-31
07Ordivon Securitytesting

Does compiled opponent state retain value across seeds, held-out policies, Context loss, and Harness replacement?

Distinguish durable strategic information from fixture-specific explanation and provider-specific reasoning style.

Current judgmentOpen after mixed Round 1 evidence. The local opponent-aware actor improved diagnosis and objective rate, one Codex strategic Trial recognized a switch, Hermes outcomes were unchanged, and no model Trial completed the genuine objective. No causal or transferable advantage is established.
1 supporting publications2026-07-31
08Ordivon Webtesting

Can the public site expose changing research judgment without becoming a second fact system?

Test whether an article-centered index of Projects and active Questions improves understanding without duplicating repository truth.

Current judgmentAnswered after contraction. The full Experiment/Finding/Decision/Event graph and explicit relation registry were removed. Project, Question, and Article metadata now drive Research, Writing, Current, article context, and curated System views with substantially less duplicate state.
1 supporting publications2026-07-31
09Ordivon Worldtesting

Can one Host Task combine path evidence and provider execution, reconcile a lost response, and continue without duplicate work?

Test whether path and external action evidence can support one recoverable Task trajectory.

Current judgmentTesting. Link and Edge are now retired as separate projects and their useful planes are preserved inside World, but no end-to-end lost-response trajectory has yet reconciled a real provider action across a path change.
0 supporting publications2026-07-31
10Ordivon Computingopen

Which Task Runtime objects are required by asynchronous waiting and Join semantics?

Determine the smallest durable objects required when work waits, fans out, and later joins.

Current judgmentOpen. Round 1 showed that mature workflow engines can carry durable work mechanics, but no interrupted fan-out and fan-in workload has yet proved which semantic Join facts must remain visible above those mechanics.
0 supporting publications2026-07-31
11Ordivon Hostopen

What operational surface is required before Host becomes installable rather than architectural proof?

Identify the minimum inspect, service, backup, and recovery surface earned by real use.

Current judgmentOpen. The journal, recovery model, operator handoff, and Harness receipts exist, but long-lived installation and operator recovery have not yet established the minimum public surface.
0 supporting publications2026-07-31
12Ordivon Hostopen

What is the smallest Ordivon Harness that turns a bare model API into a verifiable Agent Run?

Build one thin first-party model/Tool Loop without copying Host Task truth, Runtime process truth, or mature Provider Harness lifecycles.

Current judgmentReady for construction at M2. The complete model-to-work stack, H1–H5 ownership boundary, first falsifier, comparison baselines, acceptance, non-goals, and deletion condition are frozen. No implementation evidence exists yet.
3 supporting publications2026-07-31
13Ordivon Runtimeopen

Which remaining friction belongs above Runtime rather than becoming another primitive?

Prevent workflow and Harness convenience from expanding the trusted execution kernel.

Current judgmentStrongly supported. Repeated dogfood earned request lookup, progress, repair, lifecycle retention, and receipts. Harness H2/R2 then proved Assignment and Harness Run correlation, exact replay, fresh-client recovery, and terminal evidence using four opaque foreign references with zero new Runtime semantic objects.
1 supporting publications2026-07-28
14Ordivon Worldopen

Does a thin World boundary prevent a real failure better than direct Host integration?

The project boundary must prove a concrete continuity gain before it expands.

Current judgmentSupported as a corrected research boundary but still experimentally open. Link and Edge unification removed the wrong repository split; World W1 must now prove that correlated recovery creates more value than direct Host integration.
1 supporting publications2026-07-30
15Ordivon Computinganswered

Which Agent-native responsibilities remain after strong classical baselines?

Use mature workflow, retrieval, idempotency, and provider systems to delete any shared abstraction that does not own a distinct failure.

Current judgmentAnswered within the Round 1 boundary. LangGraph and Temporal carried durable work state; current-revision retrieval matched the tested Context need; single-backend Effect machinery shrank to identity, UNKNOWN, correlation, reconciliation, and no blind redispatch; live provider replacement preserved Task state without proving model equivalence.
3 supporting publications2026-07-31
16Ordivon Hostanswered

Which Harness objects survive live provider replacement without duplicating Host or Runtime state?

Prove or delete a Host-local Harness boundary across provider-faithful lifecycles, replacement, faults, and completion adjudication.

Current judgmentAnswered by H1–H5. Both replacement orders completed through one Task Attempt and fresh Assignment generation; stale completion, missing Artifact, and response loss were handled correctly. Provider final text was not portable evidence, and Codex/Hermes lifecycles remained materially different. The durable Host boundary survived; a shared internal Provider lifecycle did not.
3 supporting publications2026-07-31

Research contract

Current judgments may change; dated evidence must not.

Question summaries evolve as evidence accumulates. Publications preserve the complete argument at a date. Repositories, tests, releases, receipts, and observed behavior remain authoritative.

01

Question

A bounded uncertainty whose answer can change structure, priority, or project scope.

02

Publication

A dated complete argument that records evidence, limits, conclusions, and the next test.

03

Source

The owning repository, receipt, release, or report that remains authoritative for exact technical facts.