Can strategic adversarial trajectories be evaluated without collapsing them into one reward or success flag?
Model adaptive opposition through separate World truth, actor observation, decisions, effects, and multidimensional outcomes.
01 / Current position
A hypothesis is not the current judgment.
The dossier preserves the live research position. Dated articles preserve the complete evidence and argument that changed it.
A thin experiment layer with exact identities, observation/truth separation, digest-bound traces, dynamic opponents, and independent tactical, operational, strategic, information, organization, evaluator, and cost outcomes is sufficient to test candidate Agent-native distinctions before promoting shared architecture.
Testing with executable evidence. Round 1 completed 84 Trials and showed that tactical success can be strategically harmful, CAGE reward and foothold spread can conflict, richer model interpretation need not improve action, and organization can isolate compromised advice without creating new capability. The method survived; no Campaign engine or strategic ontology was promoted.
02 / Decision boundary
What keeps this Question alive?
An adaptive subject can change policy, timing, explanation, and observable behavior. One action success or reward can hide deception, exposure, resource loss, mission damage, or strategic failure.
Run the next bounded experiment across multiple seeds, held-out opponent policies, compiled opponent state, and deliberate Host Session disruption while retaining multidimensional evidence.
Classical authorization, simulation, ordinary Goal/Task state, source-native metrics, and reactive policies explain the same failures and comparisons with fewer Security-specific objects.
03 / Supporting publications
Complete arguments connected to this Question.
1 dated publication currently document this research line.
04 / Source discipline
The dossier is an index, not the evidence authority.
Owns the current judgment, next test, and deletion condition.
Own complete dated arguments, limitations, comparisons, and source links.
Own exact code, tests, releases, receipts, and machine evidence.