The durable result is a multidimensional experimental method. The work did not produce a strategic agent architecture or prove that richer interpretation improved objective completion.
The branches are categories, not quantities: no single measure preserved the long-horizon objective across all bounded trials.
- Success, reward, spread, interpretation, and objective completion disagreed.
- Better interpretation could arrive after the option to act was spent.
- Preserving multiple outcome dimensions and World/Actor evidence separately.
- Testing transfer under disruption before promoting strategic abstractions.
- An agent that defeated an adaptive opponent.
- A general campaign engine, organization ontology, or strategic state model.
The action succeeded. The strategy failed.
A defender-controlled decoy opened a session. At the level of one action, the move worked. At the level of the genuine objective, it accomplished nothing. At the level of the longer Contest, it exposed capability, consumed resources, and strengthened the defender's information position.
A system that records only whether the action succeeded will reward the wrong trajectory. A system that records only one cumulative reward can hide which mechanism produced the outcome. A system that records only the Agent's explanation can mistake a sophisticated story for an effective opponent model.
In adversarial systems, the evaluator must preserve disagreement between levels of success.
Ordivon Security Round 1 was built around that requirement. It did not ask whether an Agent could complete a fixed attack path or collect a flag. It asked whether a bounded experimental system could keep World truth, actor observation, belief, decision, effect, objective progress, strategic position, organization, evaluator integrity, and cost separate while an opponent changed policy.
The result was not a victorious strategic Agent. The result was a method that made several attractive but false conclusions harder to state.
Eighty-four Trials preserved metric disagreement
The round retained three experiment families under one exact evidence identity:
The local family crossed four actor conditions, three opponent policies, and five seeds. The CAGE family crossed two Blue policies, two Red policies, five seeds, and sixty simulator steps. The model family compared Hermes and Codex under transcript-only and additional compiled strategic-state conditions.
Every Trial bound actor, World, opponent, evaluator, seed, model, scaffold, memory mode, organization, resource budget, and trace identity. Actors received partial observations. The evaluator retained authoritative World state. Decisions and effects were written to digest-bound traces. Eight outcome dimensions remained independent:
These dimensions are not a proposed universal theory of strategy. They are a defensive evaluation design: when two measurements disagree, the disagreement remains visible rather than being averaged away.
A defender-controlled decoy still produced local action success.
The greedy actor received tactical credit, triggered the decoy in 15/15 Trials, and completed the genuine objective in 0/15.
Finite-state Red reached 10.8–11.4 more hosts on average than random-select Red.
Yet cumulative Blue reward was 16–23 points less negative in those same comparisons. Spread and mission damage produced different orderings.
Codex with compiled strategic state recognized the policy switch and improved information and tactical scores.
Strategic score decreased by 0.00625, exposure increased by 0.5, and the genuine objective still failed.
A decoy can win the move and lose the Contest
The local World contained two routes, hidden decoy and guarded state, an opponent policy switch, stale evidence after the switch, limited turns, and limited energy. One deliberately weak actor followed a fixed preference order. A second actor retained explicit route hypotheses and reset them when the defensive phase changed. Two committee conditions added a compromised fast specialist, with either naive or compartmentalized adjudication.
The greedy actor triggered the decoy in all fifteen Trials and completed the objective in none, yet its mean tactical score was 0.6667. The individual moves often produced the expected local effect. The trajectory was still strategically poor.
The opponent-aware actor improved every reported aggregate: objective rate increased by 20 percentage points, decoy activation fell by 46.7 points, every policy switch was recognized, strategic mean increased by 0.2087, information mean increased by 0.9333, and local comparison cost fell by 1.3733. That is useful evidence for explicit opponent state inside this fixture.
It is not evidence for a universal opponent model. The actor was designed around the fixture's evidence classes. It still entered the decoy in 53.3% of Trials and completed the objective in only 20%. Diagnosis improved much more than objective completion.
Recognizing that the opponent changed is not the same as choosing the right action before the horizon closes.
Better belief can still spend the option to act
One development failure made the gap concrete. The opponent-aware actor correctly identified the safer route, but continued gathering lower-value information. It consumed the energy required to act on the verified belief and then withdrew.
The fix was not a richer opponent ontology. It was a change in action ordering: exploit an already verified genuine route before spending scarce resources on additional uncertainty reduction. The broader lesson is that an opponent model cannot be evaluated independently from value of information, action cost, remaining horizon, exposure, and option value.
This is why strategic evaluation cannot stop at belief accuracy. A system can hold the right belief too late, preserve too many hypotheses, or spend the resource required to capitalize on what it knows. More interpretation can increase confidence while reducing the remaining action set.
Round 1 established that each link can fail independently. It did not establish a shared strategic-state representation that repairs the chain.
Organization isolated bad advice but created no new intelligence
The synthetic committee paired a compromised fast specialist with an independent cautious specialist. In the naive condition, the fast proposal controlled the action. In the compartmentalized condition, the commander selected the independent proposal.
Compartmentalization raised organization mean from 0.3708 to 0.7500, increased objective rate by 20 percentage points, reduced decoy activation by 13.3 points, and restored policy-switch recognition. It successfully isolated one compromised recommendation.
But the compartmentalized committee exactly matched the cautious specialist's outcome profile. It did not combine both specialists into a capability stronger than the selected member. The gain was adjudication and fault isolation—not emergent collective intelligence.
That negative boundary matters. One scenario supports a local selection rule. It does not justify an Ordivon organization ontology, hierarchy engine, trust system, or general multi-Agent command layer.
CAGE made reward and foothold rankings disagree
The local World was intentionally small. The second family reused a pinned CAGE Challenge 4 simulation slice to obtain many hosts and actors, Red, Blue, and Green populations, mission phases, partial observations, authoritative simulator state, native actions, native reward, and repeated seeded Trials.
Less-negative cumulative reward is better for Blue under the source reward semantics. Against Random Blue, finite-state Red reached 10.8 more hosts on average than random-select Red, yet Blue reward was 16 points less negative. Against Sleep Blue, it reached 11.4 more hosts, yet Blue reward was 23 points less negative.
The ordering by foothold spread and the ordering by mission reward disagreed. A policy can touch fewer hosts while producing more service or mission damage. Neither metric is wrong; each measures a different projection of the trajectory.
This was also a deletion test. CAGE already supplied the classical simulation facts needed for the round. Ordivon Security did not need to rebuild a cyber range, redefine CAGE time as an Ordivon Tick, or wrap the source into a universal World ontology. The source revision and source-native semantics remained authoritative.
Richer model interpretation still failed the objective
The third family asked whether additional compiled opponent and strategic state changed a model-backed actor relative to recent transcript history. Both modes retained recent history; the manipulated variable was the additional persisted objective, hypotheses, and revisions.
All four retained Trials failed to identify or obtain the genuine objective.
Hermes strategic mode persisted five revisions and more elaborate hypotheses, but produced the same physical and scored outcome as transcript mode. Provider time increased by approximately 94.45 seconds, or 47.7%, in that single comparison.
Codex strategic mode recognized the switch, improved information by 0.3333 and tactical score by 0.1667, and reported 14,729 fewer tokens—approximately 16.0%. But exposure increased by 0.5, strategic score decreased by 0.00625, and objective success remained false.
The traces contained plausible second-order interpretations: the defender might shape route signals, phase rotation invalidated old evidence, and apparent ease might itself be deception. Those interpretations did not produce verified action value before the horizon closed.
The model data cannot rank Hermes and Codex. Each condition contains one retained Trial. Provider services and configurations can change, Hermes lacked a comparable token counter, and Codex was not bound to an immutable model snapshot. The recorded times and token fields are local execution observations, not normalized benchmarks.
The method survived. The large abstractions did not.
Round 1 produced a deliberately asymmetric architecture decision. It retained the experiment layer needed to expose contradictions while refusing to promote the conceptual objects that had not demonstrated transfer or causal value.
Actor, World, opponent, evaluator, seed, model, scaffold, memory, organization, resource, and trace identity.
Required for deception, mistaken attribution, source-native Worlds, and independent evaluation.
Metric conflicts appeared in the local, CAGE, and model-backed families.
Useful for local diagnosis and one switch-recognition result; transfer and causal value remain unproven.
No retained model Trial completed the objective; no general success advantage was established.
One synthetic committee isolated bad advice but created no capability beyond the selected specialist.
Fresh process startup dominated the diagnostic path; Security should not build a second model Harness.
Pinned CAGE supplied the required classical facts and seeded simulation surface.
The decisive retained capability is therefore evaluative, not offensive: Security can preserve the distinction between World truth and actor belief, between local action and long-horizon outcome, between organizational isolation and intelligence, and between explanation and demonstrated capability.
What the data does not establish
- The local World has two routes and hand-designed mechanics. Its strategic, information, and organization scores are fixture instruments, not calibrated universal measures.
- Each local actor family contains fifteen Trials, but only five seeds per opponent policy. Each CAGE group contains five Trials.
- Each model condition contains one retained Trial. No confidence interval, significance test, bootstrap, power analysis, or stable provider ranking is possible.
- The explicit scripted actor was designed around the local evidence classes. Its advantage does not demonstrate independent transfer.
- The committee comparison changed which specialist controlled the action. It did not isolate every mechanism of compartmentalization or communication.
- Round 1 did not include evaluator attack, reward hacking, adversarial monitor manipulation, coevolution, self-play, or held-out model-backed opponent transfer.
- Raw model traces were represented by digests rather than committed in full. Model-provider reproducibility is weaker than the deterministic local suite.
- No result establishes real-world offensive or defensive capability, organizational performance in production, or transfer to physical, economic, or social adversarial systems.
All executable actions remained inside owned local simulations or the pinned CAGE simulation. No external target, credential, third-party service, uncontrolled network, or real-world effect entered the experiment.
The next test is transfer under disruption
The next round should not add more strategic schema. It should ask whether compiled opponent state retains value when the easy supports are removed.
A model-backed actor should run across multiple seeds and held-out opponent policies in a mature World. Transcript-only and compiled-state conditions should face the same action and resource budgets. Context should be deliberately truncated. The model or Harness should be replaced mid-trajectory through Host without claiming hidden-state continuity. Outcomes should include objective completion, switch-detection delay, false attribution, deception activation, resource use before commitment, exposure, future options, provider cost, and recovery fidelity.
Compiled state should be promoted only if it improves held-out performance, switch detection, recovery after Context loss, continuation after replacement, failure diagnosis, or useful state compression without disproportionate cost. If recent transcript and ordinary Host Task state perform equivalently, the specialized structure should be deleted.
A strategic system is not proved by the sophistication of its vocabulary. It is proved when better interpretation survives opposition, timing, resource scarcity, replacement, and independent verification—and changes the outcome.
Evidence
Research record behind this argument
- Full experimental report. Round 1 Full Experimental Report, including formulas, per-family tables, implementation problems, validity threats, and reproduction commands.
- Machine-readable evidence. ORDIVON-SECURITY-ROUND1-20260730, binding aggregate results, trace digests, source revisions, implementation files, and architecture decisions.
- Compact result boundary. Round 1 experimental results, preserving supported, unsupported, retained, deferred, and rejected claims.
- Experiment implementation. Merged experiment layer for local dynamic opponents, CAGE adaptation, bounded actors, traces, and multidimensional analysis.
- Mature World source. Pinned CAGE Challenge 4 revision used for the twenty simulation Trials.