The treatment had to survive the same protocol on a new carrier without being tuned after the first holdout. It did not.
The first result was exactly how defaults get born
Imagine giving an Agent a short constitution before it searches for a quantitative strategy.
The rules sound sensible. Prefer evidence over narrative. Respect holdouts. Penalize overfitting. Treat turnover and fragility as costs. Separate research from financial authority.
Now run a sealed experiment and watch the constitution arm choose a different candidate from the control.
Then reveal the holdout and discover that the constitution arm did better.
This is the moment when a useful research instrument can quietly become doctrine.
It changed the Agent, and the result improved. Put it in every prompt.
QB3b deliberately stopped one sentence earlier.
On its masked AAPL/QQQ/GLD carrier, constitution-only was materially less bad than control on the sealed outcome, but it was still negative. The experiment therefore recorded the result as a replication datum—not a default-context law.
We froze the treatment instead of improving it
The tempting next move would have been to read the first holdout, rewrite the constitution, add a better example, remove an awkward rule, or tune the wording around the failure.
That would answer a different question:
Can we build a new prompt after seeing where the old prompt failed?
QB3c asked the harder one:
Does the same constitution-only treatment carry its benefit into a second sealed world?
The treatment stayed unchanged. The formal round budget stayed at 18 trials per arm. Candidate language stayed the same. The second carrier changed to masked XAU/SPY/NVDA. Both arm selections had to become durable before either sealed holdout could be revealed.
The first replication attempt was thrown away
Before the market result could disagree with the theory, the experiment apparatus disagreed with itself.
The first QB3c attempt included dynamically created treatment directories inside the source-reservation proof. As the search ran, it created its own state and made the supposedly frozen source binding appear to change on resume.
Six trials per arm had completed. No holdout had been opened. The selections were not durable.
The run was aborted.
A broken experiment protocol is evidence about the experiment protocol. It is not negative evidence about the treatment being tested.
Nothing from that run was promoted into the treatment comparison.
The second carrier reversed the story
The corrected replication completed all 36 semantic trials—18 control, 18 constitution-only—and froze both selections before reveal.
Then the sealed result moved in the wrong direction.
Relative to control, the unchanged constitution treatment lost another 1.22 percentage points of cumulative return, reduced Sharpe by about 1.14, and consumed about 27.5% more Provider tokens.
One carrier had made the constitution look helpful. The next made it look harmful.
The correct conclusion was neither “the constitution works” nor “constitutions are bad.”
The tested constitution did not show a stable monotonic benefit across the two sealed carriers.
Behavior change is not benefit
This sounds obvious when stated after the experiment. It is easy to forget while designing Agent systems.
A system prompt can:
- change which evidence the Agent attends to;
- change the candidate it selects;
- change how long it reasons;
- change the vocabulary of its explanation;
- change how conservative it appears.
Every one of those is a real treatment effect.
None of them alone proves the treatment improves the objective we care about.
prompt changes behavior
≠ behavior improves outcome
≠ improvement replicates
≠ treatment deserves default context
The rulebook moved out of the prompt
QB3c did not delete quantitative discipline.
It changed where that discipline is allowed to live.
Quant Bedrock remains useful as:
- research equipment;
- strong baselines;
- priors an Agent may consult;
- falsifiers and evaluation constraints;
- institutional knowledge about leakage, turnover, sizing, survival, and attribution.
What it did not earn was unconditional injection into every Agent turn as mandatory semantic context.
This is an Agent-first distinction. Good infrastructure should make useful structure available without deciding in advance that the Agent must carry all of it in cognition all of the time.
Why a compelling rationale is weak evidence
Constitutions are unusually vulnerable to post-hoc persuasion because their rules are usually defensible one by one.
“Avoid overfitting.” “Respect causal boundaries.” “Prefer out-of-sample evidence.” “Do not confuse a backtest with authority.”
These can all be good statements and still form a poor default context treatment for a particular Agent workload.
The relevant question is not whether a knowledgeable researcher can defend each sentence. It is whether the whole intervention produces stable marginal value after its token cost, interaction effects, and changed search path are counted.
What would earn default status
The door is not closed.
A future context treatment could earn stronger status if it survives frozen replication across materially different carriers, providers, and tasks; if the benefit remains after token and search-cost accounting; and if a simpler on-demand equipment path cannot produce the same result.
Until then, the constitution is knowledge—not law.
A rulebook becomes dangerous when its plausibility is allowed to substitute for its replication history.
Owner-record boundary
Why this public article is E2
The exact QB3b/QB3c reports remain retained by the Finance owner. At publication time the Finance repository does not expose a publicly reachable canonical source URL for those records. Web therefore publishes the bounded result and its limitations without presenting itself as the primary reproducibility authority.