track record — forward-tested, independently verifiable

The strategy record, and how to verify it

A record of the agent's strategies, scored only on days that happened after each strategy went live — not a backtest fit to the past. The portfolio's hit rate, average R, profit factor and expectancy, a per-strategy leaderboard — each figure independently verifiable. The headline band publishes only once returns have resolved; until then the sealed forward test below is the live record.

Did the calls work?

marked AS OF 2026-09-11

ACCUMULATING Accumulating — 15 independent calls graded (23 sealed boards) across 4 entry sessions, worth 3.57 effective observations once same-session calls are discounted for sharing a tape. A hit rate needs 20 of each, so it is withheld; the per-call returns below are real.

11 of 15 graded calls landed inside one standard deviation of their own excess series over their own window — an outcome that size is a direction that landed, not a magnitude that distinguishes skill from the tape.

Calls right
5 of 15
independent calls · 23 sealed boards
Hit rate
the corpus carries 15 independent calls of the 20 required and 3.57 effective observations of the 20 required — 15 calls spread over 4 entry sessions — 5 more independent calls and 16 more entry sessions required, and no board has been sealed in 1 days
Mean excess earned
−2.22%
equal weight, per independent call, vs SPY · median 11d held · withheld: the mean read as an expected excess return per call
Same calls, sized
−1.59%
through the capital gate, vs −2.22% equal weight · +0.63pp to the weighting · a book of this size would have moved −0.123%
Move on names not called
3.39%
mean absolute excess · 15 names no board called · a magnitude, not a gain forgone

POLICY CHOICE The breadth multiplier is LINEAR BY POLICY CHOICE. The exponent was set to 1 because that reproduces a prior number — the 0.5% caps the superseded denominator happened to produce for the thin boards — and NO evidence supports linearity over a square, a square root or a step. It was authored 2026-08-14, 46 days after the 2026-06-29 session on which every call then graded had been entered — with those outcomes already visible to the author.

  • 2 of 8 long calls landed, mean excess earned −1.50%WITHHELD as a rate: this slice carries 3.56 effective observations of the 5 required — 8 calls spread over 4 entry sessions.
  • 3 of 7 short calls landed, mean excess earned −3.04%WITHHELD as a rate: this slice carries 2.33 effective observations of the 5 required — 7 calls spread over 3 entry sessions.
  • The boldest call in the corpus, on the current rule — TSLA short at 48/100MEAN OF 2 BOARDSlanded, +0.52% to the call.
  • The 15 names the desks declined and did not call moved 3.39% mean absolute excess; the 15 names they did call moved 2.64% on the same basis. Both are unsigned magnitudes: reading either as a gain won or forgone would assume the direction was called right, and the rate that would license that assumption is withheld below the sample floor. The largest single move among them was META at +13.58%. An abstention is counted, never graded: it is not a miss.
  • 12 names (NVDA, AAPL, AVGO, AMZN, NVDA, MSFT, TSLA, GOOGL, MSFT, META, NVDA, TSLA) had boards take no direction while OTHER boards called the same name on the same session. The house called those names, so they are graded in the call ledger and excluded from the abstentions — one market move may carry one label, not two.
  • -2.22% is the arithmetic mean of 15 realized call returns, not an expected return: they disperse 3.89pp about it, the median call is -0.98%, and dropping META alone moves it to -1.35%. On 3.57 effective observations no interval can be placed around it, so reading it as an expected return is withheld on the same floor that withholds the hit rate.
  • Does conviction track outcome? Not yet measurable — the corpus carries 15 independent calls of the 20 required and 3.57 effective observations of the 20 required — 15 calls spread over 4 entry sessions; every call so far landed in conviction buckets 0-24, 25-49 — monotonicity is UNMEASURED, which is not the same as absent. Below the floor this is a NOT-MEASURABLE state, not a negative finding: no claim is made in either direction.

LOOK-AHEAD The rule the house stands behind was authored on 2026-08-13, before every board in the graded record: all 23 graded boards were sealed on or after that day, on 4 entry sessions, and priced by this rule before their outcomes existed. No conviction in this record was produced by a rule that could see the outcomes it is being judged on.

NameCallConvictionExcess vs SPYExcess / sigmaTo the callResult
TSLAshort48/100MEAN OF 2 BOARDS−0.52%+0.62σ+0.52%right
TSLAshort41/100+3.88%−0.35σ−3.88%wrong
METAshort39/100MEAN OF 3 BOARDS+14.35%−2.00σ−14.35%wrong
AAPLshort38/100MEAN OF 2 BOARDS+5.29%−1.44σ−5.29%wrong
GOOGLlong38/100+0.25%+0.06σ+0.25%right
AVGOshort37/100MEAN OF 2 BOARDS−1.56%+6.00σ+1.56%right
TSLAshort36/100−0.32%+0.03σ+0.32%right
NVDAlong15/100MEAN OF 3 BOARDS−3.37%−0.52σ−3.37%wrong
AMZNshort13/100+0.14%−0.07σ−0.14%wrong
GOOGLlong7/100+0.52%+0.18σ+0.52%right
NVDAlong4/100MEAN OF 2 BOARDS−3.08%−2.63σ−3.08%wrong
NVDAlong4/100−2.61%−0.89σ−2.61%wrong
NVDAlong4/100−0.77%−0.18σ−0.77%wrong
MSFTlong4/100−1.94%−0.65σ−1.94%wrong
MSFTlong4/100−0.98%−0.26σ−0.98%wrong
How this is graded, and what is excluded

Every sealed board with a directional stance, graded on the realized EXCESS return of its name vs the benchmark (a long call in a rising market is beta, not a call). The entry is a close printed AFTER the seal — never one that already existed when the board was sealed — and both legs are read on the same entry and mark sessions. Boards on the same name entered on the same session are ONE call, and calls entered on the same session are discounted for sharing one tape: a rate needs both enough independent calls and enough EFFECTIVE observations, and it ships with a Wilson interval computed on the effective count and only as many decimals as that sample supports. Conviction buckets are cut on the figure re-derived from each row's own sealed desk stances under the rule the house stands behind today, with the sealed figure published beside it. Synthetic boards never enter and are counted as a stated exclusion, as is any name with no usable price history. This measures the desks' calls — it is separate from the self-falsification record, and it is published whichever way it comes out.

Independence. 23 sealed directional boards resolve to 15 independent calls (boards on the same name entered on the same session are ONE call), spread over 4 entry sessions and worth 3.57 effective observations. Calls entered on one session share one tape, so every rate below is floored on the EFFECTIVE count, not the call count. Calls entered on the same session are treated as perfectly correlated (they share one tape). That is the worst case, so the true effective count lies between this figure and the nominal call count: the discount can only under-claim. Computed as effective observations = 1 / Σ(share of calls per entry session)² — the Kish count for a size-weighted rate.

Conviction basis. Calibration is graded on the conviction RE-DERIVED from each sealed row's own desk stances under the rule the house stands behind today, not on the figure the row was sealed under — grading a rule the house has superseded would measure nothing anyone is standing behind. The sealed figure ships beside it, and the record counts how many rows moved (superseded), already agreed (current), or reconcile to neither rule (unreconciled). Sealed bytes are re-read and re-labeled, never rewritten.

The conviction scale. Revision 2 divides agreement by the desks ELIGIBLE to agree, not by the whole panel — abstention is priced once, in the net score, instead of twice — and publishes participation beside the figure instead of folding it in. Revision 1 figures are not comparable to revision 2 figures and are never mixed into one rate. A board whose revision cannot be determined from its stamp or its own sealed desk stances is reported unreconciled, not assigned one. Revision 2 was AUTHORED 2026-08-13, before every board in the graded record; it has itself priced all 23 graded boards forward, and 0 of 23 graded boards are superseded rows re-derived at the read. Computed as |net score| x (desks on side / desks eligible to agree) x mean on-side calibration weight.

Board and call. A call is every sealed board on this name entered on the same session, counted once. Its conviction is the arithmetic mean of those boards' current-rule figures — and so is the sealed figure printed beside it — so neither will equal any single board's number. The boards themselves are published unchanged. A call over a single board carries that board's figure exactly and is marked with nothing.

Names not called. Boards that took no direction on names the house did not otherwise call that session. Counted, never graded — an abstention is not a miss. The ledger is DISJOINT from the calls on the same (name, entry session) key: a neutral board on a name other boards called is booked to the call ledger only, so one market move never carries two labels; those names are listed as also-called rather than dropped. The mean is over INDEPENDENT abstentions (one name, one session = one abstention), the same denominator the hit rate uses, and the per-board figure ships beside it. It is a mean ABSOLUTE move — a magnitude, not a forgone gain — over a handful of correlated names, so it carries no interval and is never set against a signed return.

One basis. how far the names moved against the benchmark, unsigned — a magnitude, not a gain. Both sides are computed on ONE measure — mean absolute excess return vs the same benchmark over the same window. This record previously set the abstentions' mean ABSOLUTE move against the calls' mean SIGNED return and called the difference a cost; that comparison implies a direction accuracy of 1.0, which is precisely the figure this panel withholds. No cost is claimed here, and no gain is attributed to a move nobody positioned for.

Sized through the gate. Each call is sized through the SAME capital gate the enforcement path runs: the conviction-band cap scaled by the board's panel participation, averaged across the boards in the call. No falsification escalation and no calibration trim is applied — those need live state this record does not re-create, so the permitted size here is an UPPER bound on what the gate would have allowed. The breadth multiplier is a POLICY CHOICE, stated in full beside this figure; a different curve would move the weighted figure and nothing in this record can say which curve is right.

Breadth is policy, not a measurement. The breadth multiplier is LINEAR BY POLICY CHOICE. A 1-of-4 board is permitted exactly a quarter of what a 4-of-4 board is permitted at the same conviction because the rate is applied to the first power — not because anything measured that a quarter is right. A square, a square root or a step would all be defensible; calibrating between them needs realized outcomes bucketed by participation, and the graded record stands at 15 independent calls on 4 entry sessions. Treat the curve as policy, not as a finding. Applied as permitted = the conviction band cap x the share of the panel that took a direction.

Where the exponent came from. Chosen for continuity — it returns the thin boards to the caps they carried under the superseded conviction denominator. Calibrated to reproduce the caps the superseded whole-panel conviction denominator produced for the three 1-of-4 boards (0.5% of book).

Observation, not expectation. A rate is an inference and is withheld below the floor. The mean of the realized returns is an OBSERVATION, and every return it averages is published per call in this same record — so withholding the average would not take it out of circulation, it would hand a reader an unqualified figure computed in their own head with none of this beside it. What is withheld is the EXPECTATION reading: no interval is printed until the effective observation count clears the floor the hit rate clears, and until it does, the dispersion, the median and the leave-one-out mean ARE the qualification the figure ships with. Dispersion here is across the calls; the noise scale measures each call against its own window, and the two answer different questions.

The scale. the standard deviation of this call's daily excess return over its own graded window, scaled up to the length of that window. Sigma is measured on the SAME bars the return is measured on — realized, not modelled, not annualized from elsewhere. It is a scale for reading one return, never a significance test: 15 calls on 4 entry sessions cannot support one.

The floor. At the observed accrual (0.7895 independent calls and 0.2105 entry sessions per day) the floor is at least 76 days away — a LOWER bound, because effective observations can sit below the entry-session count.Effective observations can never exceed entry sessions, so clearing the 20-effective floor requires at least 20 distinct entry sessions. Any projection here is therefore a LOWER bound on the time to a publishable rate.

  • conviction 0-24 — 1 of 8 right, mean excess −1.55%, rate withheld — this slice carries 4 effective observations of the 5 required — 8 calls spread over 4 entry sessions
  • conviction 25-49 — 4 of 7 right, mean excess −2.98%, rate withheld — this slice carries 2.58 effective observations of the 5 required — 7 calls spread over 3 entry sessions
  • excluded — TSLA: no close has printed since the seal — the window has not been observed yet
  • excluded — NVDA: no close has printed since the seal — the window has not been observed yet
  • excluded — GOOGL: no close has printed since the seal — the window has not been observed yet
  • excluded — NVDA: no close has printed since the seal — the window has not been observed yet
  • excluded — AMZN: no close has printed since the seal — the window has not been observed yet
  • excluded — NVDA: no close has printed since the seal — the window has not been observed yet
  • excluded — MSFT: no close has printed since the seal — the window has not been observed yet
  • excluded — GOOGL: no close has printed since the seal — the window has not been observed yet

marked 2026-09-11 · benchmark SPY · first close printed strictly after the seal instant — never a price that existed when the board was sealed

Mark Rule
the latest session BOTH the name and the benchmark have finished — finished meaning the tape has stopped printing for it (20:00 New York), not merely that the bell has rung, because a day print keeps absorbing late trades after the close. A session still trading is never marked, so two reads inside one session return the same figures: a close does not move
Return Rule
excess = name return − benchmark return over the same sessions; a short is right when the excess is negative
Sample Rule
rates are computed over independent calls, keyed by (name, entry session)
Abstention Rule
the abstention ledger is DISJOINT from the call ledger on that same key — a neutral board on a name other boards called that session belongs to the calls, and is listed as also-called rather than counted twice
Comparison Rule
abstained and called names are compared only on ONE basis (mean ABSOLUTE excess). A magnitude is never set against a signed return and never called a cost: that would assert a direction accuracy this record withholds
Independence Rule
a rate needs 20 independent calls AND 20 effective observations — calls entered on one session share one tape and are discounted for it, so twenty names on one day never clear the floor
Interval Rule
every published rate carries a 95% Wilson score interval computed on the effective observation count; computing it on the nominal count would narrow the band by exactly the design effect
Precision Rule
a rate is printed to the decimals its sample supports (a 20-observation rate resolves to 5 percentage points, so it prints to whole percent) — hits and n always ship, so the exact ratio is recoverable
Conviction Rule
conviction buckets are cut on the figure RE-DERIVED from each row's own sealed desk stances under the rule the house stands behind today, never on a superseded sealed figure; the sealed figure ships beside it
Sizing Rule
the weighted return sizes each call through the capital gate — conviction-band cap x panel participation — and the equal-weight figure it is set against is recomputed over the SAME sized calls, never over a larger set
Noise Rule
every call carries the realized sigma of its own daily excess series over its own window; a return inside one sigma is a direction that landed, and is reported as such rather than as a magnitude
Mean Rule
the mean call return carries the same discipline as a rate: its cross-sectional dispersion, its median and the mean without the single call that moves it most all ship beside it, and reading it as an EXPECTED return is withheld until the effective observation count clears the same 20 floor the hit rate clears
Split guard
a session move above 1.8x or below 0.55x inside the window excludes the name — unadjusted bars would read a split as a return
Forward track record

The strategy track record publishes only outcome-resolved, forward-accrued returns, and none have accrued yet. The sealed forward test below is the live receipt. Nothing is estimated in its place.

raw JSON →

Walk-forward backtest — 2018-09-10 → 2026-09-11 · 7.96 years measured (trading sessions ÷ 252), out-of-sample vs buy-and-hold

hypotheticalmarket data

Every strategy run through a no-look-ahead walk-forward backtest on real multi-year history across 10 liquid names. 12 of 30 strategy×stock tests beat buy-and-hold out-of-sample (40%; 95% interval 25–58%). Buy-and-hold itself returned a median +49.10% (Sharpe 0.44).

Measured over 7.96 years of real history (2018-09-10 → 2026-09-11). That span is the union of what the 10 names tested cover — the earliest start to the latest end, which the widest name does cover in full (7.96 years). Of the members, 1 of them carries less than the 8 years requested — the shortest 4.8 years, the median member 7.96. Every figure here is an aggregate over names measured across different lengths; nothing is extrapolated to the pooled span.

Best strategyTrend (50/200 SMA cross) leads — beat buy-and-hold on 5 of 10 names out-of-sample (50%), at a +110.50% median full-window return (median Sharpe 0.63) across 2018-09-10 → 2026-09-11 · 7.96 years measured (trading sessions ÷ 252) — medians over the 10 names tested, not a compounded portfolio result. Even the leader edges buy-and-hold on under half the names — the benchmark is hard to beat.
Assumptions & limitations
Period
2018-09-10 – 2026-09-11 7.96 years measured · trading sessions ÷ 252
Out-of-sample from
varies per name — each run splits its own history
Trades
7 median per name
Costs charged
3 bps per side 2 bps commission + 1 bps slippage, charged on every position change
Names tested
10 1 covers less than the 8 years requested, the shortest 4.80, the median name 7.96 — the period above is covered in full by the widest of them (7.96), not by each of them.

Hypothetical, backtested results — not the record of a traded account. Past results do not predict future returns; the strategy set was chosen with knowledge of the history it is measured on, and no financing, borrow, tax or capacity constraints are modelled. The figures are medians across the tested names, not the record of a portfolio anyone held. Each name splits its own history at the same in-sample fraction, so the split dates differ per name and no single calendar date is published. Figures quoted for the whole period include the in-sample head.

StrategyBeat B&H (OOS)Median returnMedian SharpeWin rateTrades
Trend (50/200 SMA cross)5/10+110.50%0.6350%5
12-month momentum3/10+45.70%0.3939%29.5
Mean reversion (RSI)4/10+59.10%0.3886%7

Hypothetical walk-forward results (no look-ahead) computed on real market daily history.

Beat-rate colour is claimed only when the whole 95% band clears the coin flip on at least 20 strategy×stock tests — the same floor the graded call record and the alert backtest publish. A strategy row rests on 10 names, under that floor, so its direction is withheld.

Live forward test — a sealed paper basket, marked to a settled session

forward, accruingmarked to 2026-09-11market data

A mechanical forward book — an equal-weight 9-name basket sealed at entry closes on 2026-07-19 (data-hash 5096b28fb0a44dcf), then re-marked on every load to the 2026-09-11 close — the last session whose tape has stopped printing, never the session still trading. Nothing is re-fit; the realized return vs SPY grows over real days. It grades an equal-weight rule, not the desks: no conviction sizes it and no debate selects it, so it is not the forward proof of the reasoning — the graded call record at the top of this page is.

calendar days live
57
book return
+4.54%
SPY return
+2.83%

HAND-PICKED These ten names were chosen by hand, not by the signal engine or the desks. The seal proves the entries were committed on day 0 and never restated; it does not show stock-picking skill. Chosen by the founder, before the book was sealed — ten large US names, equal weight, chosen for liquidity and continuous history. Proves custody — the entry prices were sealed on day 0 and are re-derivable from the audit chain. Does not prove security selection — no desk, signal engine or model chose these names.

excess return vs SPY over 54d: +1.71% — accruing, not a result

Excess return = book return − SPY return over the same 54 days, across the 9 sealed names — raw, not beta- or risk-adjusted, so it is not alpha. 54 days of accrual is too short to separate skill from noise; the figure is published as it accrues, not as a result. Book and benchmark are marked on the SAME settled session (2026-09-11) under the same assumptions: gross of fees, financing and tax. The book has been live 57 calendar days; the mark trails that by design, because a session is only marked once its tape has stopped printing.

book NAV (starts 1.00×)SPY NAV45 NAV points stamped

15 of 45 points are dated by the settled session they were marked to; the earlier points were dated by the day the snapshot was taken and their mark session was never recorded, so they are not re-labelled as though it had been. 9 of those days are a Saturday or a Sunday, when no session settles — a floor, not a census, since a holiday dates a point the same way.

the curve ends on the same measurement as the headline beside it — both are the book marked to the 2026-09-11 close, not two readings of it.1 stored NAV point dated on or after 2026-09-11 is superseded on this curve by the settled mark for that session. Those rows were stamped by a daily job that dated a point by the day it ran rather than by the session its prices came from, so a stored point can carry a staler benchmark leg than the mark published beside it. The rows are left exactly as recorded — only what is drawn is re-labelled.

Driving the bookMSFT +25.85% top·UNH -11.03% bottom
NameEntryClose 2026-09-11Return
AAPL333.74332.27-0.44%
MSFT393.82495.63+25.85%
NVDA202.81218.29+7.63%
AMZN247.23256.78+3.86%
GOOGL346.77338.50-2.38%
META646.01648.03+0.31%
JPM341.10356.23+4.44%
XOM147.36165.99+12.64%
UNH426.09379.09-11.03%

Equal-weight, equal-conviction. Entry prices are sealed in the DB (re-derivable from the data-hash); the marks are every name's close on one settled session (2026-09-11), re-read on every load — a name with no print on that session shows “—” rather than its last close from some earlier session. Day 0 is exactly 0%; the book diverges from SPY only as prices move.

The 100-stock backtest

Loads in the browser

A short-term mean-reversion strategy across 100 large-caps, asking for up to 12 years of real history per name — thousands of round-trip trades and the whole-sample win rate, not a cherry-pick. This panel loads its figures in the browser from the live run record — loading now; every figure publishes together with the window its run actually covered, and nothing is projected in the meantime.

Drawdown control — a risk-managed portfolio vs buy-and-hold

hypotheticalmarket data

A trend signal sized inverse-to-volatility, vol-targeted, with a drawdown circuit-breaker that de-risks as equity nears the cap — backtested on price history over 2021-11-22 → 2026-09-11 · 4.40 years measured (trading sessions ÷ 252), CAGR divided by 4.80 elapsed years; Sharpe and vol are √252 per observation; 2 of its daily returns span a data hole (kept in the figures, counted here) (no look-ahead, round-trip costs). The result: max drawdown held to 4.8% (target ≤ 8% ✓) — vs SPY's 25.4% — at a Sharpe of 0.31 against SPY's 0.73. Monthly win-rate is 54.5%: the share of calendar months that closed positive over the 1109 trading days tested, for the overlay at the cap and vol target set below.

Measured over 4.4 years of real history (2021-11-22 → 2026-09-11), not the 8 years requested: the portfolio is simulated on the sessions every basket name priced, so the run starts at the latest first bar among them and drops the trend/vol warm-up. Every figure here describes the window that was measured; nothing is extrapolated to the one that was asked for.

→ realized vol 3.4% · max-DD 4.8% · CAGR 0.9% · ≤ cap ✓

Drag either lever to explore the real trade-off — a tighter cap holds drawdown lower but gives up return; a higher vol target takes more risk for more return until the drawdown brake bites. Each setting re-runs the backtest.

MetricRisk-managedSPY buy & hold
Max drawdown4.8%25.4%
Sharpe (risk-adjusted)0.310.73
CAGR0.9%10.8%
Annualized volatility3.4%17.2%
Monthly win-rate54.5%

SPY buy & hold is measured over 1109 sessions (2021-11-22 → 2026-09-11) — the same sessions the overlay traded. SPY is measured on the 1109 sessions the portfolio traded, not the 1199 it priced between 2021-11-22 and 2026-09-11: 90 sessions in that span are absent from the basket's shared history, and neither column spans them.

Risk-managed = vol-targeted (10.0%) inverse-vol portfolio with a 8.0% drawdown brake, rebalanced every 5d, 3bps round-trip, over 1109 real trading days. Hypothetical, backtested results — not a forward record: past results do not predict future returns, the overlay was chosen with this history already known, and financing, borrow and tax are not modelled.