Walk-forward validation
A backtest optimized and reported on the same date range tells you how well a config fit that window's noise — not whether it will work tomorrow. Walk-forward validation is the honest alternative: optimize on one slice of time, measure on a later slice the optimizer never saw, and reserve a final slice nobody touches until the very end.
uv run python main.py walkforward \
--strategy demo_trend --scanner demo_volume --symbols NVDA,AAPL,MSFT \
--start 2024-01-01 --end 2025-12-31 --capital 100000 \
--mode anchored --folds 6 \
--embargo-days 5 --holdout-days 60 \
--method grid --objective sharpe_ratio --max-evals 50
Why it matters
The failure mode this guards against is overfitting — mistaking the random noise in a particular stretch of history for a real, repeatable edge.
It's easy to do by accident. Search enough parameter combinations against one date range and something will look brilliant, purely by luck — the more configs you try, the better the best one looks even if none of them has any edge. (This is the multiple-testing problem, and it's why TradeFlow also reports a deflated Sharpe that discounts your result by how many configs you tried.) Optimize and report on the same window and you've measured how well you fit that window's noise, not whether the strategy will work next month.
make demo shows the trap in one screen: a moving-average crossover posts
+16.8% in-sample, then collapses to a −0.42 median Sharpe out-of-sample and
fails every promotion gate. Same strategy, same data — the only thing that
changed is that it had to perform on bars it wasn't tuned on.
Walk-forward closes the gap by splitting time into three roles:
- In-sample (IS) — the only data the optimizer is allowed to fit.
- Out-of-sample (OOS) — later data the optimizer never saw, used to measure whether the chosen config generalizes. Repeated across folds so the verdict doesn't hinge on one lucky window.
- Holdout — a final slice carved off first and scored exactly once, at the very end. It's the closest thing to "live" you get before risking real money, so it must never influence any decision along the way.
The payoff isn't a higher return — it's an honest one. A config that clears the promotion gates here has at least been asked the right question: does this work on data it has never seen?
What it does
For each fold it optimizes parameters on an in-sample (IS) window, then scores the chosen config on the following out-of-sample (OOS) window. An embargo gap separates IS from OOS so indicator warm-up can't leak across the boundary. A holdout window is carved off the end first and scored exactly once, at the very end — it never reaches any optimizer call.
The honest performance number is the OOS aggregate: metrics recomputed over the concatenation of every fold's OOS trades, not an average of per-fold numbers.
--save-config produces a run config — the artefact backtest,
walk-forward and live all read. That page covers what is in it and how the universe is
replayed rather than re-scanned.
Key options
| Flag | Meaning |
|---|---|
--mode anchored|rolling | Expanding IS window (anchored) or fixed-width sliding IS (rolling). |
--folds N | Number of folds. Alternatively set --train-days / --test-days. |
--embargo-days N | IS→OOS gap. Defaults to the strategy's required lookback in calendar days. |
--holdout-days N | Final sacred window, scored once. |
--method grid|random|bayesian | Search method used per fold (bayesian needs the optimize extra). |
--objective | Metric to optimize in-sample (default sharpe_ratio). |
--pbo | Also estimate the Probability of Backtest Overfitting (slower). |
--monte-carlo | Block-bootstrap the OOS trades for a 5th-percentile Sharpe. |
--param-sensitivity | Perturb the chosen params ±10% and re-test robustness. |
--leakage-probe | Shift the data feed forward to detect future-data leakage. |
--bootstrap-skill | Nonparametric own p-value (stationary block bootstrap) next to the FAMILY p from White's Reality Check over every OOS return series the trial store has recorded for this strategy/universe/accounting — advisory only, not a gate. See below. |
--save-config PATH | Save the chosen config — params and run inputs — for a human to review. See Reusing a saved config. |
--results-csv PATH | Write the per-fold table to CSV. |
Reading the output
The report prints a per-fold table (IS vs OOS headline metrics, OOS trade count),
the OOS aggregate block, walk-forward efficiency, IS→OOS degradation, the holdout
block, and the promotion-gate verdict — a pass/fail per gate plus an overall
promotable. A config is only promotable if it clears every gate (median OOS
Sharpe, profit factor, efficiency, drawdown ratio, minimum OOS trades, deflated
Sharpe, and — when requested — parameter sensitivity and the leakage probe).
Saving a config never changes live behavior. It writes a JSON file to a gitignored
configs/directory; promoting it to live trading is a manual human step.
Excess return by fold
The prerequisite gates on a median, and a median is exactly where regime failure hides:
=== Excess return by fold (diagnostic) ===
+0.25% -2.02% +3.18%
median +0.25% spread 5.20pp
1 of 3 folds lost to the benchmark - the median is an average over folds that
disagree, not a typical fold.
A median of +0.25% over a five-point spread is not the same result as +0.25% in every fold, and only one of those two numbers makes anyone look closer. The spread is reported, not gated — the median already gates, and nobody yet knows what "too much disagreement" is worth failing a candidate over.
Leg beta by fold
When a strategy trades both sides, each leg's beta is reported per fold:
=== Leg beta by fold (diagnostic) ===
long +0.88 +0.91 +0.62
short -0.85 -0.90 -0.14
A book neutral on average can still be directional inside a fold.
A book that is neutral on average and directional within folds is a different proposition from one that is neutral throughout, and no aggregate separates them — the same reason the benchmark prerequisite takes a median across folds rather than a figure over the stitched curve. In the example the third fold's short leg has largely stopped hedging, which the two-year net beta would hide completely.
Requires --benchmark; diagnostic, and gates nothing. See
long/short legs for the per-leg return, volatility,
drawdown and cost breakdown on a single backtest.
Promotion prerequisites
promotable stays statistical: median OOS Sharpe, profit factor, walk-forward
efficiency, drawdown ratio, trade count, parameter sensitivity, deflated Sharpe. It
means the same thing it meant for every trial already recorded, and nothing added since
has changed it.
Two further questions come after a candidate clears those, and are reported beside it:
=== Promotion prerequisites (separate from `promotable`) ===
[PASS] cost_stress 5 vs 3
[ -- ] family_bootstrap not evaluated - needs 10 usable return-series trials to
mean anything; 2 available
Prerequisites: 1 of 3 evaluated - clear so far; 2 unknown
An unevaluated check is not a passed one - what is unknown stays unknown.
cost_stress— the edge survives at least 3x its own assumed cost, stressing the config the folds actually chose. On by default here: this is where a promotion decision is made, so cost sensitivity belongs in that story rather than in an optional follow-up.--no-cost-stressskips it; it re-runs the chosen config once per multiple.family_bootstrap— still notable once every trial the campaign tried is priced in. It does not run below 10 usable return-series trials: a family test over two series is arithmetic rather than evidence, and a striking p-value on K=2 is exactly the kind of number that should not gate anything.benchmark_excess— the median per-fold return less the benchmark's over the same steps is positive. A different question from the ratio below: a strategy can hold a good risk-adjusted number while losing to the benchmark outright, which is not something to promote on.benchmark_relative— the median per-fold information ratio against--benchmarkis positive. Per fold, then median, because every other fold statistic here is a median and a second aggregation convention in one report would differ from its neighbours most exactly when the folds disagree. They do: a real run produced per-fold IRs of[0.13, -1.25, 2.10]for a median of+0.13, and a single figure over the stitched curve would have hidden that spread entirely.
An unevaluated check is not a passed one. A cost curve nobody ran and a family too
thin to test are both unknown, and ready always travels with the count of what was
actually evaluated — a clear verdict over one of two checks is not a clearance.
Reusing a saved config
The file holds the whole run configuration, not just tuned params: strategy,
params, scanner, symbols (the universe the scanner resolved, not the candidate
list), candidate_symbols, capital, position_limits and the cost model.
position_limits is written out in full, not left to the strategy's defaults. A
config is what a paper or live run gets frozen from, and the shipped default of
max_positions: 1 is exactly the kind of inheritance nobody chose — a 61-name config
that quietly holds one position. What a run risks belongs in the file rather than being
resolved from somewhere else at start-up. A config written before this still works:
recorded keys win, unrecorded ones fall back to the strategy's own. So one file drives any run type, and can be
versioned in a private repository beside the strategies it belongs to:
tradeflow backtest --config configs/alpha.json --start 2024-01-02 --end 2024-06-01
tradeflow verdict --config configs/alpha.json --start 2024-01-02 --end 2024-06-01
tradeflow alphas --config configs/alpha.json
tradeflow risk --config configs/alpha.json
--config is accepted by backtest, live, verdict, info, alphas, horizon,
allocate and risk. Four rules make it predictable:
-
Anything you type wins. The file fills in only what the command line left unsaid, so
--symbols ZZZbeats the file's universe. Each run prints where every value came from. -
A command takes only the fields it has.
risksummarizes a universe's covariance and has no strategy, so it uses the file'ssymbolsand nothing else — and says so rather than naming a strategy it never ran. -
A contradictory
--strategyis refused. The params in the file belong to the strategy in the file; handing one strategy's tuned params to another is not something to guess at. -
The saved universe is replayed, not re-scanned. A config records the book that was validated. Re-running the scanner over it would turn that into a new book from an old recipe — a different experiment, which can move results either way without the reader knowing the universe changed.
--re-resolve-universeopts into a genuine re-scan, over the saved candidate list rather than the resolved book, because re-scanning what the scanner already picked is a second filter rather than the original decision repeated. Every run says which it did:universe=<replayed from config, 61 symbols>universe=<re-resolved from 85 saved candidates>universe=<--scanner demo_volume given; saved book re-scanned>universe=<--symbols given; --re-resolve-universe has nothing to re-resolve>
The window is never stored. A config carrying its own tuning dates would make
every later run silently re-evaluate that period, so --start/--end always come
from the run and the output says so. What the config was tuned on is recorded under
provenance.windows, for reading rather than replaying.
Nonparametric skill check
--bootstrap-skill adds a second, assumption-free verdict next to the deflated
Sharpe: an own p-value from a stationary block bootstrap of this run's OOS
returns (no assumption about the return distribution's shape), always reported
next to the family p-value from White's Reality Check over every other
trial recorded for the same strategy/universe/accounting in the
trial store — a great own p and a terrible family p is
exactly the selection-luck signature this test exists to catch. Family scoring
is advisory, not a hard gate. See
Evaluation metrics — bootstrap skill inference.
The trial store
Every walk-forward run (and every backtest/optimize/alphas run, and the
research agent) is dual-written into a queryable SQLite index over the research
journal — logs/trials.db. It's what makes campaign-wide counts answerable
without reading the whole journal:
python main.py trials query --strategy demo_trend --symbols NVDA,AAPL,META
python main.py trials status # row/journal-line counts + a drift check
python main.py trials rebuild # rebuild from the journal — safe, it's derived
trials query --strategy ... --symbols ... prints the campaign's real
n_trials — useful to check by hand, since the promotion gate above still only
counts the current run (deliberately: wiring the campaign count into the gate
automatically would make every gate strictly harder and reclassify configs
already saved as promotable — an open, evidence-backed decision, not an
oversight). See
Walk-forward validation — the trial store (engineering).
See the engineering wiki's Walk-forward validation page for the design, fold geometry, and the leakage-safety guarantees.
What makes a run a repeat
A walk-forward is memoized by its validation recipe rather than by its parameters — the window, the folds, the embargo, the objective, the search method and seed, the cost assumptions, and the book it validated at. Ask the same recipe twice and the second is answered from the first.
The book that goes into that identity is the resolved one: the strategy's own limits with any config override merged over them. It used to be only the override, which meant a run that overrode nothing recorded no book at all — while still having one. A class default moving from one position to eight changed the experiment without touching the params, the universe or the window, so two such runs looked identical to the memo and the second was served the first's answer.
Runs recorded before this change key differently and will recompute once. That is
deliberate. Recomputing is cheaper than reusing a one-position result to answer an
eight-position question, and trials rebuild is unaffected — it reads each journal
line's own recorded recipe, so historical rows keep the identity they were written with.
Running folds faster (--workers N)
--workers parallelizes each fold's in-sample candidate search across worker
processes. Folds stay sequential — the candidates are where the work is, and
per-fold progress stays readable.
It changes wall-clock only: the same seed produces the same folds, the same chosen config, and the same campaign trial count as a sequential run, because workers only execute and this process still does every journal write. See parallel candidates for the full contract and its costs.