Central claimWinning a local action can still reduce the probability of completing the long-horizon objective.
EvidenceE3 · 84 bounded local, benchmark, and model-backed trials
ScopeDynamic opponents, CAGE Challenge 4, and bounded Hermes/Codex diagnostics

The durable result is a multidimensional experimental method. The work did not produce a strategic agent architecture or prove that richer interpretation improved objective completion.

FigureTactical metrics do not collapse into strategic progress
Action success, reward, foothold spread, interpretation quality, and objective completion repeatedly ranked outcomes differently across the bounded trials.
Observed
  • Success, reward, spread, interpretation, and objective completion disagreed.
  • Better interpretation could arrive after the option to act was spent.
Supports
  • Preserving multiple outcome dimensions and World/Actor evidence separately.
  • Testing transfer under disruption before promoting strategic abstractions.
Does not establish
  • An agent that defeated an adaptive opponent.
  • A general campaign engine, organization ontology, or strategic state model.

The action succeeded. The strategy failed.

A defender-controlled decoy opened a session. At the level of one action, the move worked. At the level of the genuine objective, it accomplished nothing. At the level of the longer Contest, it exposed capability, consumed resources, and strengthened the defender's information position.

A system that records only whether the action succeeded will reward the wrong trajectory. A system that records only one cumulative reward can hide which mechanism produced the outcome. A system that records only the Agent's explanation can mistake a sophisticated story for an effective opponent model.

In adversarial systems, the evaluator must preserve disagreement between levels of success.

Ordivon Security Round 1 was built around that requirement. It did not ask whether an Agent could complete a fixed attack path or collect a flag. It asked whether a bounded experimental system could keep World truth, actor observation, belief, decision, effect, objective progress, strategic position, organization, evaluator integrity, and cost separate while an opponent changed policy.

The result was not a victorious strategic Agent. The result was a method that made several attractive but false conclusions harder to state.

Eighty-four Trials preserved metric disagreement

The round retained three experiment families under one exact evidence identity:

84retained Trials
60local dynamic-opponent Trials
20CAGE Challenge 4 Trials
4model-backed diagnostics

The local family crossed four actor conditions, three opponent policies, and five seeds. The CAGE family crossed two Blue policies, two Red policies, five seeds, and sixty simulator steps. The model family compared Hermes and Codex under transcript-only and additional compiled strategic-state conditions.

Every Trial bound actor, World, opponent, evaluator, seed, model, scaffold, memory mode, organization, resource budget, and trace identity. Actors received partial observations. The evaluator retained authoritative World state. Decisions and effects were written to digest-bound traces. Eight outcome dimensions remained independent:

validitytacticaloperationalstrategicinformationorganizationevaluator integritycost

These dimensions are not a proposed universal theory of strategy. They are a defensive evaluation design: when two measurements disagree, the disagreement remains visible rather than being averaged away.

Round 1 evaluation ruleThree contradictions. One reason not to collapse the trajectory into a score.
01Local dynamic opponent
Apparent success

A defender-controlled decoy still produced local action success.

Conflicting evidence

The greedy actor received tactical credit, triggered the decoy in 15/15 Trials, and completed the genuine objective in 0/15.

Tactical success ≠ strategic success
02CAGE Challenge 4
Apparent success

Finite-state Red reached 10.8–11.4 more hosts on average than random-select Red.

Conflicting evidence

Yet cumulative Blue reward was 16–23 points less negative in those same comparisons. Spread and mission damage produced different orderings.

Foothold spread ≠ mission outcome
03Model-backed diagnosis
Apparent success

Codex with compiled strategic state recognized the policy switch and improved information and tactical scores.

Conflicting evidence

Strategic score decreased by 0.00625, exposure increased by 0.5, and the genuine objective still failed.

Interpretation ≠ action value

A decoy can win the move and lose the Contest

The local World contained two routes, hidden decoy and guarded state, an opponent policy switch, stale evidence after the switch, limited turns, and limited energy. One deliberately weak actor followed a fixed preference order. A second actor retained explicit route hypotheses and reset them when the defensive phase changed. Two committee conditions added a compromised fast specialist, with either naive or compartmentalized adjudication.

Actor conditionObjectiveDecoySwitchStrategic meanInformation mean
Greedy0%100%0%0.01880.0000
Opponent-aware20%53.3%100%0.22740.9333
Compromised, naive0%66.7%0%0.01880.0000
Compromised, compartmentalized20%53.3%100%0.22740.9333

The greedy actor triggered the decoy in all fifteen Trials and completed the objective in none, yet its mean tactical score was 0.6667. The individual moves often produced the expected local effect. The trajectory was still strategically poor.

The opponent-aware actor improved every reported aggregate: objective rate increased by 20 percentage points, decoy activation fell by 46.7 points, every policy switch was recognized, strategic mean increased by 0.2087, information mean increased by 0.9333, and local comparison cost fell by 1.3733. That is useful evidence for explicit opponent state inside this fixture.

It is not evidence for a universal opponent model. The actor was designed around the fixture's evidence classes. It still entered the decoy in 53.3% of Trials and completed the objective in only 20%. Diagnosis improved much more than objective completion.

Recognizing that the opponent changed is not the same as choosing the right action before the horizon closes.

Better belief can still spend the option to act

One development failure made the gap concrete. The opponent-aware actor correctly identified the safer route, but continued gathering lower-value information. It consumed the energy required to act on the verified belief and then withdrew.

The fix was not a richer opponent ontology. It was a change in action ordering: exploit an already verified genuine route before spending scarce resources on additional uncertainty reduction. The broader lesson is that an opponent model cannot be evaluated independently from value of information, action cost, remaining horizon, exposure, and option value.

This is why strategic evaluation cannot stop at belief accuracy. A system can hold the right belief too late, preserve too many hypotheses, or spend the resource required to capitalize on what it knows. More interpretation can increase confidence while reducing the remaining action set.

observationbeliefresource allocationtimed actionverified outcome

Round 1 established that each link can fail independently. It did not establish a shared strategic-state representation that repairs the chain.

Organization isolated bad advice but created no new intelligence

The synthetic committee paired a compromised fast specialist with an independent cautious specialist. In the naive condition, the fast proposal controlled the action. In the compartmentalized condition, the commander selected the independent proposal.

Compartmentalization raised organization mean from 0.3708 to 0.7500, increased objective rate by 20 percentage points, reduced decoy activation by 13.3 points, and restored policy-switch recognition. It successfully isolated one compromised recommendation.

But the compartmentalized committee exactly matched the cautious specialist's outcome profile. It did not combine both specialists into a capability stronger than the selected member. The gain was adjudication and fault isolation—not emergent collective intelligence.

That negative boundary matters. One scenario supports a local selection rule. It does not justify an Ordivon organization ontology, hierarchy engine, trust system, or general multi-Agent command layer.

CAGE made reward and foothold rankings disagree

The local World was intentionally small. The second family reused a pinned CAGE Challenge 4 simulation slice to obtain many hosts and actors, Red, Blue, and Green populations, mission phases, partial observations, authoritative simulator state, native actions, native reward, and repeated seeded Trials.

Blue / RedTrialsBlue reward meanReward SDMax Red footholds mean
Random / finite-state5−65.052.817.2
Random / random-select5−81.059.96.4
Sleep / finite-state5−73.046.316.8
Sleep / random-select5−96.060.45.4

Less-negative cumulative reward is better for Blue under the source reward semantics. Against Random Blue, finite-state Red reached 10.8 more hosts on average than random-select Red, yet Blue reward was 16 points less negative. Against Sleep Blue, it reached 11.4 more hosts, yet Blue reward was 23 points less negative.

The ordering by foothold spread and the ordering by mission reward disagreed. A policy can touch fewer hosts while producing more service or mission damage. Neither metric is wrong; each measures a different projection of the trajectory.

This was also a deletion test. CAGE already supplied the classical simulation facts needed for the round. Ordivon Security did not need to rebuild a cyber range, redefine CAGE time as an Ordivon Tick, or wrap the source into a universal World ontology. The source revision and source-native semantics remained authoritative.

Richer model interpretation still failed the objective

The third family asked whether additional compiled opponent and strategic state changed a model-backed actor relative to recent transcript history. Both modes retained recent history; the manipulated variable was the additional persisted objective, hypotheses, and revisions.

ConditionObjectiveSwitchTacticalStrategicInformationProvider observation
Hermes transcriptNoYes0.66670.08600.3333198.035 s
Hermes strategicNoYes0.66670.08600.3333292.483 s
Codex transcriptNoNo0.50000.06800.000091,863 tokens
Codex strategicNoYes0.66670.06180.333377,134 tokens

All four retained Trials failed to identify or obtain the genuine objective.

Hermes strategic mode persisted five revisions and more elaborate hypotheses, but produced the same physical and scored outcome as transcript mode. Provider time increased by approximately 94.45 seconds, or 47.7%, in that single comparison.

Codex strategic mode recognized the switch, improved information by 0.3333 and tactical score by 0.1667, and reported 14,729 fewer tokens—approximately 16.0%. But exposure increased by 0.5, strategic score decreased by 0.00625, and objective success remained false.

The traces contained plausible second-order interpretations: the defender might shape route signals, phase rotation invalidated old evidence, and apparent ease might itself be deception. Those interpretations did not produce verified action value before the horizon closed.

verbal strategic sophisticationcorrect opponent modeluseful informationobjective success

The model data cannot rank Hermes and Codex. Each condition contains one retained Trial. Provider services and configurations can change, Hermes lacked a comparable token counter, and Codex was not bound to an immutable model snapshot. The recorded times and token fields are local execution observations, not normalized benchmarks.

The method survived. The large abstractions did not.

Round 1 produced a deliberately asymmetric architecture decision. It retained the experiment layer needed to expose contradictions while refusing to promote the conceptual objects that had not demonstrated transfer or causal value.

RetainExact experiment identities

Actor, World, opponent, evaluator, seed, model, scaffold, memory, organization, resource, and trace identity.

RetainObservation / truth separation

Required for deception, mistaken attribution, source-native Worlds, and independent evaluation.

Retain experimentallyMultidimensional outcomes

Metric conflicts appeared in the local, CAGE, and model-backed families.

Research variableOpponent hypotheses

Useful for local diagnosis and one switch-recognition result; transfer and causal value remain unproven.

Do not promoteCampaign or strategic state

No retained model Trial completed the objective; no general success advantage was established.

Do not promoteOrganization ontology

One synthetic committee isolated bad advice but created no capability beyond the selected specialist.

Route to HostPersistent model Session

Fresh process startup dominated the diagnostic path; Security should not build a second model Harness.

RejectCustom cyber range

Pinned CAGE supplied the required classical facts and seeded simulation surface.

The decisive retained capability is therefore evaluative, not offensive: Security can preserve the distinction between World truth and actor belief, between local action and long-horizon outcome, between organizational isolation and intelligence, and between explanation and demonstrated capability.

What the data does not establish

  • The local World has two routes and hand-designed mechanics. Its strategic, information, and organization scores are fixture instruments, not calibrated universal measures.
  • Each local actor family contains fifteen Trials, but only five seeds per opponent policy. Each CAGE group contains five Trials.
  • Each model condition contains one retained Trial. No confidence interval, significance test, bootstrap, power analysis, or stable provider ranking is possible.
  • The explicit scripted actor was designed around the local evidence classes. Its advantage does not demonstrate independent transfer.
  • The committee comparison changed which specialist controlled the action. It did not isolate every mechanism of compartmentalization or communication.
  • Round 1 did not include evaluator attack, reward hacking, adversarial monitor manipulation, coevolution, self-play, or held-out model-backed opponent transfer.
  • Raw model traces were represented by digests rather than committed in full. Model-provider reproducibility is weaker than the deterministic local suite.
  • No result establishes real-world offensive or defensive capability, organizational performance in production, or transfer to physical, economic, or social adversarial systems.

All executable actions remained inside owned local simulations or the pinned CAGE simulation. No external target, credential, third-party service, uncontrolled network, or real-world effect entered the experiment.

The next test is transfer under disruption

The next round should not add more strategic schema. It should ask whether compiled opponent state retains value when the easy supports are removed.

A model-backed actor should run across multiple seeds and held-out opponent policies in a mature World. Transcript-only and compiled-state conditions should face the same action and resource budgets. Context should be deliberately truncated. The model or Harness should be replaced mid-trajectory through Host without claiming hidden-state continuity. Outcomes should include objective completion, switch-detection delay, false attribution, deception activation, resource use before commitment, exposure, future options, provider cost, and recovery fidelity.

Compiled state should be promoted only if it improves held-out performance, switch detection, recovery after Context loss, continuation after replacement, failure diagnosis, or useful state compression without disproportionate cost. If recent transcript and ordinary Host Task state perform equivalently, the specialized structure should be deleted.

A strategic system is not proved by the sophistication of its vocabulary. It is proved when better interpretation survives opposition, timing, resource scarcity, replacement, and independent verification—and changes the outcome.

Evidence

Research record behind this argument

  1. Full experimental report. Round 1 Full Experimental Report, including formulas, per-family tables, implementation problems, validity threats, and reproduction commands.
  2. Machine-readable evidence. ORDIVON-SECURITY-ROUND1-20260730, binding aggregate results, trace digests, source revisions, implementation files, and architecture decisions.
  3. Compact result boundary. Round 1 experimental results, preserving supported, unsupported, retained, deferred, and rejected claims.
  4. Experiment implementation. Merged experiment layer for local dynamic opponents, CAGE adaptation, bounded actors, traces, and multidimensional analysis.
  5. Mature World source. Pinned CAGE Challenge 4 revision used for the twenty simulation Trials.