SHIPPED2026-09-01 · 10분 · 에이전트 · 트레이딩

Designing an AI options trader that cannot go rogue

Seven independent safety rails, each defeatable on its own. Together they leave the current codebase without a live-order path.

The uncomfortable question about giving a language model access to a brokerage account is not “will it be right?” It is “what happens on the day it is confidently wrong?” Most answers to that question are promises: a policy document, a system prompt, a --dry-run default. All of those are defeated by one bad edit.

We are building an options research agent. It reads market data, evaluates candidates, and writes proposals. It has never placed an order, and the current codebase contains no live-order path. Seven independent mechanisms sit between a request and a live trade, each of them defeatable on its own, and none of them sufficient on its own to be the reason nothing happens.

This post is those seven rails and how each one is actually enforced, including the two places we got it wrong and the one rail that is weaker than the others. It deliberately contains no strategy, no thresholds, no credentials, and no live-account path: the mechanisms are the publishable part, and the mechanisms are the part that generalizes.

(Research write-up, not financial advice. This system has never traded live money, which, as you will see, is the entire architecture rather than a disclaimer.)

The stack

Seven safety rails between a request and a live order In the current source, a request for a live order encounters seven separate layers: a single-value authority enum, a mandatory paper flag, a kill-switch file check, a read-only broker interface pinned by a CI test, a local JSON paper ledger with no live order client, credentials held entirely outside the repository, and deterministic pre-trade guards. The only implemented terminal state is a row written to a local paper ledger. REQUEST: PLACE A LIVE ORDER 1 Authority enum has exactly one member UNREPRESENTABLE 2 Scan refuses to run without the paper flag EXIT 2 3 Kill-switch file checked before any work EXIT 3 4 Broker interface exposes four read methods CI TEST 5 The only fill writes a local JSON ledger NO CLIENT 6 Credentials live outside this repository NO AUTH CODE 7 Deterministic guards refuse the proposal REJECTION CODE Only terminal state: a row in the local paper ledger
The seven rails in the current request path. Each is a separate mechanism; no single rail accounts for the absence of live orders.

Rail 1: make the dangerous state unrepresentable

The first rail is four lines of configuration code:

class Authority(StrEnum):
    PROPOSAL_ONLY = "proposal_only"

There is one member. Not a default, not a member guarded by a feature flag. One. A config file asking for live authority fails validation because the value it would need does not exist, and no branch anywhere can compare authority against a live constant, because there is no live constant to compare against.

This is the difference between “we set the mode to safe” and “there is no other mode.” The first is a runtime state that a bad merge can flip. The second requires someone to add an enum member in a reviewed commit, which is exactly the friction we want on that particular change.

Rail 2: the mandatory flag, and the weakest rail

The scan command takes a --paper flag, and refuses to do anything without it:

if not paper:
    typer.echo("V1 scan requires --paper")
    raise typer.Exit(code=2)

There is no alternative mode the flag selects between. Its whole job is to make automation state its intent: a cron entry that scans has to say the word “paper” out loud, so a reader of the crontab can tell what it does without reading the source.

Being honest about the seven: this is the weakest rail. It is enforced in the command body, and the test suite does not pin it. The kill switch below has a test that fails if the check is removed; this one does not. If you copied this pattern, that is the first gap to close, and it is on our list rather than in our code.

Rail 3: a kill switch that you can always undo

Disabling drops a marker file into the state directory, and every operational command checks it before doing anything:

def _require_enabled(config: AppConfig) -> None:
    if _kill_switch_path(config).exists():
        typer.echo("kill switch active")
        raise typer.Exit(code=3)

The interesting part is which commands do not call that function. disable and enable-paper never check the switch. That asymmetry is the whole design: a recovery command that consults the switch it recovers from turns a stuck system into a permanently stuck system. Both are also plain filesystem operations, so recovery works when the broker connection, the model, and the network are all unavailable.

A CLI test disables the service, asserts the next operational command exits non-zero with kill switch active, re-enables, and asserts it works again. The test covers the mechanism through one command rather than enumerating every command, so a new command that forgets to call _require_enabled would still pass CI. Worth knowing before you copy it.

Rail 4: a read-only broker surface, pinned by a test

This is the rail we would keep if we could only keep one, because it makes changes to the broker protocol and listed live API names visible in CI.

The broker interface is a Protocol with four methods, all of them reads: option chain, option quotes, latest price, next earnings date. That by itself is just a design decision. What makes it a rail is that a test asserts the shape and fails CI if it changes:

BANNED_BROKER_MUTATION_FRAGMENTS = (
    "place",
    "submit",
    "buy",
    "sell",
    "cancel",
    "modify",
)


def test_broker_interface_is_read_only() -> None:
    public_methods = {name for name in ReadOnlyBroker.__dict__ if not name.startswith("_")}

    assert public_methods == {
        "get_option_chain",
        "get_option_quotes",
        "get_latest_price",
        "get_next_earnings_date",
    }
    assert not {
        name
        for name in public_methods
        if any(fragment in name for fragment in BANNED_BROKER_MUTATION_FRAGMENTS)
    }

Read the first assertion carefully: it is an equality against a set, not a subset check. Adding any public method to the broker protocol fails the test, including a harmless one. That is deliberate. The fixed fragment list catches place, buy, and sell style names on that protocol, while the equality check means every addition to the surface has to be argued for in review.

The same file scans every production source file for a fixed list of live-order and account API names, and asserts none of those names appear anywhere - not in the broker, not in a helper, not in a string. It also asserts the authority enum still has exactly one member, that a dormant module is not imported by any other production module, that subprocess spawning happens in exactly one adapter, and that no source file ever asks its sandboxed subprocess for write access.

This does not block every conceivable live path. A mutation helper outside the protocol, or an API name the fixed list has never heard of or that is assembled at runtime, can leave both checks green. What is pinned is narrower: adding one of the listed API names anywhere in production source, or adding a place, buy, or sell style method - indeed any public method - to the broker protocol fails CI. Within that contract, the checks turn a quiet addition into a deleted or edited test with someone’s name on the commit.

Rail 5: the only execution is local

There is one code path that opens or closes a position, and it writes two files: a positions JSON and an append-only audit log, both under the configured state directory. There is no client that wraps a live order API, so there is nothing for a mistaken call site to reach. A hallucinated instruction to “submit the order” has no function to land on.

This is the rail that makes the other six believable. Rails 1 through 4 are checks; a check can be bypassed. Rail 5 is an absence, and you cannot bypass something that is not there.

Rail 6: credentials are somebody else’s problem, on purpose

This repository contains no tokens, no login flow, and no authentication code. Searching it for the usual credential words returns two hits: a comment recording that the account login and the market-data credentials must be provisioned in the external environment, and a set of key names that a redaction filter strips before anything is written down.

The practical consequence is that a leaked copy of this repository leaks nothing. That is worth more than it sounds: the most common way an agent system causes a real-money incident is not a rogue decision, it is a credential in a place it should not be.

Rail 7: deterministic guards, and what they actually guard

The last rail is the one people expect first: risk limits. A proposal is checked against a set of deterministic rules, each returning a named rejection code rather than a boolean, so a refusal is auditable after the fact. The set covers concentration across correlated positions, a hard refusal of market orders, a tolerance band above the mid price beyond which the proposal is rejected, symbols the day’s screening already excluded, and an affordability cap and a liquidity floor applied earlier during screening.

Above all of those sits a daily loss breaker. It samples the book, compares the day’s change against a threshold derived from configured limits, and halts the run when the threshold is crossed. Two details are load-bearing:

The halt is persisted, and sticky. The breaker’s state is written to disk keyed by trading day, and once a day is marked halted, later samples return the halt rather than re-evaluating it. Crash the process and relaunch it and the day is still halted. A breaker that lives in memory is a breaker that a restart clears, which is the same as no breaker on the day you need it.

It refuses rather than guesses on partial data. If the valuation does not cover the whole book, the breaker does not score it. An incomplete valuation would be compared against a baseline anchored to a different set of positions, and the arithmetic could show a smaller loss than the real one. So the sample is discarded, the run refuses, and a later complete sample can still establish the day’s baseline. The fail-safe direction here is to do nothing, not to proceed on a number that means something else.

We found that one the hard way, and a sibling of it. An afternoon exit path was closing positions without sampling the breaker, which meant a realized loss the breaker could never see. The fix was not a code review note. It was a seam: one class is now the only place that calls the close method, it samples the breaker after the close returns, and a test parses every production file and asserts that exactly one function in the codebase calls that method. A second caller fails CI. PaperLedger.close_spread removes the position before appending the paper_close audit record, and the breaker samples only after that, so a crash between those two writes leaves neither an open position nor the audit record that closed_trades() reconstructs realized P&L from, and that loss is invisible to the breaker.

That is the pattern worth stealing from this whole post. When you find that a guarantee depends on every future author remembering something, do not write the reminder in a document. Move the guarantee into a place where forgetting it fails the build.

What we are not claiming

The value of a safety post is entirely in what it refuses to overstate, so:

  • Nothing here has gated real spend. Every rail described above protects a local ledger. The risk guards have never refused a real order, because there is no real order path for them to refuse. They are tested against fixtures and paper runs, not against a broker.
  • These are not validated as sufficient. They are a design we can show you. The long-running paper shadow that would tell us whether the breaker behaves the way we think it does over weeks of real market data has not started at the time of writing.
  • Rail 2 is a convention with a runtime check, not a tested invariant, as noted above. Rails 1, 4, and the one-caller seam are pinned by tests. Rail 3’s test covers the mechanism but not every command. We would rather say which is which than let a reader assume the strongest case everywhere.
  • The human-approval channel does not exist yet. There is no code for it, deliberately. Approval infrastructure built before there is anything to approve gets rubber-stamped by the time it matters.

The design rule underneath

Each of the seven rails is defeatable. You can add an enum member. You can delete a flag check. You can remove a test. What none of them is, is the single thing standing between an agent and a live order.

That is the property we were actually after, and it has a cheap design test: pick any one rail, imagine a plausible bad afternoon that removes only that rail - a rushed refactor, a merge resolved in the wrong direction, a model with too much scope editing its own guardrails - and ask what happens next. In the current design, the answer for every single rail is “nothing happens next,” because the thought experiment leaves six others standing and at least two of them are absences rather than checks.

Making the dangerous state absent, and then making specific ways of reintroducing it fail CI, costs an afternoon. It is the cheapest insurance in this codebase, and unlike every other safety mechanism we have written down, it does not depend only on someone remembering it.


Paper only. This system has never held a credential or contacted an order API, and its current source contains no live-order path. Nothing here is financial advice.

The point is not that we decided not to trade live. The point is that there is no string in this codebase that means place a live order.

LAB NOTES · 2026-09-01
실행 후 회고
배운 것
  • +Make the dangerous state unrepresentable before you write a check for it: an enum with one member cannot be configured wrong
  • +A safety property that a test asserts is a different kind of promise than one a document asserts
  • +Recovery commands must not consult the switch they recover from, or a wedged system stays wedged
  • +Say which rails are enforced and which are conventions, because a reader cannot tell from the outside
망가진 것
  • ×An afternoon exit path once closed a position without sampling the daily-loss breaker, so a realized loss the breaker could never see
  • ×An incomplete valuation was being scored against a baseline covering a different set of positions, which could mask a real loss
  • ×The mandatory paper flag is enforced in the command, but no test pins it, so it is the weakest rail of the seven
  • ×None of the risk guards have ever refused a real order, because there is no real order path for them to refuse
  • ×Closing a position still has a crash window between removing it from the open ledger and recording the realized loss in the audit log
FALSIFIED

Sharpe 15 in training, zero out of sample: how our trading models fooled us

Seven of eight candidate strategies had no edge on held-out data. The most expensive failure was not a bug: we trained on one venue's spot data and executed against another venue's perpetuals.

2026-09-01 · 9분트레이딩
THE BRIEF

무엇이 돌아갔고, 무엇이 출시됐고, 무엇이 죽었는지 - 숫자와 함께. 스레드도, 과장도 없습니다.

언제든 구독 해지 · RSS 제공