Four allocation rules, three crisis windows, nineteen years of regime change.
Equal Weight, minimum variance, Max Sharpe, and a rule-based regime-aware allocation are run on identical data, constraints, and costs. The question is not which wins, but how each behaves when diversification assumptions stop holding.
- Universe
- 13 ETF proxies
- Rules
- 4 allocation methods
- Benchmarks
- SPY + 60/40
- Rebalance
- Monthly
- Costs
- 5 bps × turnover
- Lookbacks
- 6M / 1Y / 3Y
- Equal Weight
- GMV
- Max Sharpe
- Regime-Aware
Market data through 2026-08-19: daily adjusted closes and Cboe VIX from 2007, alongside FRED CPI (July 2026) and Fama-French factors (2026-06-30), which publish on slower calendars. Monthly rolling backtests with no look-ahead; every figure below is generated by the Python research engine.
The diversification assumption under test
Static diversification assumes asset relationships stay useful. The analysis treats that assumption as the variable under test.
Stock-bond diversification is conditional.
When equity and Treasury returns turn positively correlated, the hedge that most allocation frameworks rely on weakens at exactly the wrong time.
Inflation regimes move both legs at once.
Persistent inflation and rate shocks can push equities and nominal bonds down together, as both did through 2022.
So allocation rules are tested across regimes.
Correlation, inflation, and volatility stress are classified explicitly, and every rule is evaluated inside and outside those states using liquid ETF proxies.
Sources and provenance
- [1]
The stock-bond correlation is not structurally stable
Research on long-run U.S. data shows the stock-bond correlation flipped from mostly positive (1970s-1990s) to mostly negative (2000s-2010s), and rose again as inflation returned in 2021-2022, meaning the diversification benefit of Treasuries is regime-dependent, not guaranteed.
Source supports the broader correlation-regime relationship. The SPY-TLT 90-day rolling correlation used in this analysis is calculated from ETF return data, not taken from the source.
- [2]
Inflation regimes reshape asset relationships
U.S. CPI inflation (FRED series CPIAUCSL) exceeded 3% year-over-year for an extended stretch in 2021-2023 after a decade mostly below it. Inflation persistence is a key driver of how stocks, nominal bonds, TIPS, and real assets co-move.
Federal Reserve Bank of St. Louis, FRED, CPIAUCSL, May 2026
CPI levels are sourced from FRED; the year-over-year transformation, one-month lag, and 3% stress threshold are calculations and assumptions of this analysis.
- [3]
VIX as a market stress gauge
The CBOE Volatility Index measures 30-day expected S&P 500 volatility implied by option prices. It is widely used as a market stress indicator; it is an index, not a directly investable asset.
Cboe Global Markets, VIX Index, Jan 2026
VIX levels come from the published index; the 25 stress threshold is an assumption of this analysis. The strategies never allocate to VIX.
- [4]
Institutional allocators monitor cross-asset behavior continuously
Standard institutional references such as J.P. Morgan's quarterly Guide to the Markets track asset-class returns, valuations, and correlations across cycles, reflecting that allocation decisions are evaluated against changing market conditions, not a single historical average.
J.P. Morgan Asset Management, Guide to the Markets, Apr 2026
Used as context for why the analysis reviews returns, correlations, and risks together rather than treating allocation as a static average, a smaller-scale version of standard institutional allocation review practice.
- [5]
Factor exposures are a standard return diagnostic
The Fama-French research factors (market, size, value, profitability, investment) from the Kenneth French Data Library are the standard academic benchmark for decomposing portfolio returns into systematic exposures.
Kenneth R. French, Data Library, Dartmouth, May 2026
Factor return data comes from the library; the regressions themselves are computed by this analysis as an exposure diagnostic, not as proof of alpha.
Investable assets and signal series
Thirteen liquid ETFs serve as transparent public-market proxies. VIX, CPI, and the SPY-TLT correlation enter as signals only, never as holdings.
Equity, credit, and listed real estate exposure
Treasuries across the curve plus a T-bill cash proxy
TIPS, gold, and broad commodities
Inputs to the regime engine, never allocations
Asset role matrix
| Ticker | Asset class | Role | Investable? | Model role |
|---|---|---|---|---|
| SPY | U.S. large-cap equity | Equity benchmark | Investable | Benchmark + growth/risk sleeve |
| QQQ | U.S. growth equity | Growth exposure | Investable | Growth/risk sleeve |
| IWM | U.S. small-cap equity | Size exposure | Investable | Growth/risk sleeve |
| EFA | Developed international equity | International diversification | Investable | Growth/risk sleeve |
| EEM | Emerging markets equity | EM growth / risk exposure | Investable | Growth/risk sleeve |
| TLT | Long-duration Treasuries | Duration exposure | Investable | Duration sleeve |
| IEF | Intermediate Treasuries | Core bond exposure | Investable | 60/40 benchmark + duration sleeve |
| TIP | TIPS | Inflation-linked bond exposure | Investable | Inflation hedge sleeve |
| HYG | High-yield credit | Credit risk exposure | Investable | Credit/risk sleeve |
| GLD | Gold | Real asset / crisis hedge | Investable | Real asset hedge sleeve |
| DBC | Broad commodities | Inflation-sensitive exposure | Investable | Commodity stress basket |
| VNQ | REITs | Real estate exposure | Investable | Listed real estate / risk sleeve |
| SHV | Treasury bills / cash proxy | Defensive cash-like exposure | Investable | Cash proxy / defensive sleeve |
| VIX | Implied volatility index | Market stress indicator (signal only, never an allocation) | Signal only | Signal only |
| CPI YoY | Inflation series | Inflation regime indicator (lagged one month, signal only) | Signal only | Signal only |
| SPY-TLT corr | Derived diversification signal | Stock-bond diversification indicator (signal only) | Signal only | Signal only (calculated from ETF returns) |
Signal-only rule: VIX is used as a market stress indicator only. The strategies never allocate directly to VIX, volatility-linked products behave very differently from the spot index and are out of scope here.
Four allocation rules, one specification
Every rule runs on the same universe, rebalance calendar, constraints, and transaction-cost model. The matrix lines up their mechanics so differences in behavior can be traced to differences in inputs.
Baseline Equal Weight Naive diversification baseline with no optimization. | Optimized GMV Covariance-driven defensive allocation. | Optimized Max Sharpe Return-risk optimization, and a live demonstration of its input sensitivity. | Rule-based Regime-Aware Test how conditional defensive allocation behaves across changing regimes; the purpose is comparison, not proof of superiority. | |
|---|---|---|---|---|
| Inputs |
|
|
|
|
| Constraints |
|
|
|
|
| Expected behavior | Steady diversified exposure; risk dominated by the equity sleeve. | Defensive, bond- and cash-heavy allocations; low participation in equity rallies. | Chases whatever had the best recent risk-adjusted run; allocation shifts sharply across lookbacks. | Holds more defensive assets when multiple stress signals are active; lags fast reversals. |
| Known weakness | Treats a T-bill ETF and emerging-market equity as equally deserving of capital. | Can concentrate in whatever looked calm during the lookback window. | Expected returns estimated from short windows are noise-dominated; instability is the lesson, not a bug. | A rules-based framework tested on the same history that motivated its rules. |
| Primary diagnostic | Benchmark-relative drawdown | Concentration and weight stability | Turnover and lookback sensitivity | Crisis-window drawdown vs. benchmarks |
The regime-aware rule is one specification among four. The comparison is the output; no rule is presented as a recommendation.
Rolling backtest construction
At each rebalance date, the model uses only historical data available before that date to estimate inputs, choose weights, and hold the portfolio until the next rebalance.
Monthly rebalancing is a historical simulation rule, it does not mean the published figures update monthly.
- 01Historical data
Daily adjusted-close returns up to t−1
- 02Estimate inputs
Means, covariances, signals on the lookback
- 03Solve weights
Strategy rule or optimizer, capped long-only
- 04Apply costs
5 bps × turnover vs drifted weights
- 05Compute returns
Hold weights until next month-end
- 06Measure risk
Drawdowns, VaR/CVaR, turnover, factors
No look-ahead bias. A rebalance at month-end t sees returns and signals only through t−1. CPI is additionally lagged one month so the strategy never reacts to a print before its public release. New weights earn returns starting the next trading day.
Costs and drift. Weights drift with daily returns between rebalances. Turnover is measured against those drifted pre-trade weights, so even Equal Weight pays realistic rebalancing costs. The initial cash buy-in is excluded from costs and turnover statistics, which is why the 100% SPY benchmark shows zero cumulative cost drag.
Metric conventions. Annualized returns are geometric annualized returns from daily net returns. Annualized volatility is daily return volatility scaled by √252. Sharpe ratios use daily excess returns scaled by √252; Sortino uses downside deviation relative to a zero daily return threshold. VaR and CVaR are daily historical loss estimates at 95% confidence, reported as positive loss magnitudes.
Risk-free conventions. The Max Sharpe optimizer uses the annualized SHV return over its lookback; reported Sharpe ratios use the daily Fama-French risk-free series. SHV is used inside the optimizer as an investable cash proxy, while Fama-French RF is used for standardized performance reporting, two conventions, each fit for its purpose.
Automated research-engine checks
18 of 18 checks passed · ledger generated 2026-08-19
- data completenesspass
- return calculationpass
- weights sum and boundspass
- no lookahead biaspass
- transaction cost applicationpass
- turnover calculationpass
- var cvar conventionpass
- factor regressionspass
- regime signal availabilitypass
- cpi lagpass
- benchmark constructionpass
- data coverage honestypass
- correlation matrix boundspass
- correlation matrix symmetrypass
- effective bets boundspass
- pca share boundspass
- hedge effectiveness integritypass
- export completenesspass
warning · gfc partial data: The common ETF history begins in 2007 (SHV inception 2007-01, HYG inception 2007-04), so longer lookback windows do not cover the full 2008-2009 GFC window. For consistency, crisis comparisons use the common valid start date required by the selected strategy universe and lookback window; benchmarks are also aligned to that common sample and marked insufficient when they cannot be compared over the same full aligned window. Affected and reported as insufficient_data: Equal Weight (1Y), Global Minimum Variance (1Y), Max Sharpe / Risk-Adjusted Return (1Y), Regime-Aware Allocation (1Y), SPY (1Y), 60/40 (1Y), Equal Weight (3Y), Global Minimum Variance (3Y), Max Sharpe / Risk-Adjusted Return (3Y), Regime-Aware Allocation (3Y), SPY (3Y), 60/40 (3Y).
Full-sample performance and risk
Cumulative wealth and risk metrics over the full available sample, net of transaction costs by default. Historical results under these assumptions, not a promise about the future.
Cumulative wealth, growth of $100
1Y estimation lookback · monthly rebalancing · net of transaction costs
Benchmarks shown dashed. Series start when a full lookback window of common ETF history is available, so longer lookbacks begin later. Historical simulation, not a forecast.
Rankings within this sample, among allocation strategies (benchmarks excluded)
+225% cumulative
-11.3% peak-to-trough
0.65 vs FF risk-free
-0.29% cumulative cost drag
Full-sample risk and performance
All metrics computed from daily net returns · 1Y lookback
| Strategy | Ann. return | Ann. vol | Sharpe | Sortino | Max DD | VaR 95 | CVaR 95 | Turnover /mo | Cumulative cost drag |
|---|---|---|---|---|---|---|---|---|---|
| Equal Weight | 6.7% | 11.2% | 0.51 | 0.84 | -34.9% | 1.01% | 1.71% | 2.6% | -0.29% |
| GMV | 2.3% | 3.4% | 0.30 | 0.99 | -11.3% | 0.32% | 0.50% | 7.7% | -0.84% |
| Max Sharpe | 6.0% | 7.3% | 0.65 | 1.14 | -18.2% | 0.75% | 1.15% | 38.4% | -4.11% |
| Regime-Aware | 6.1% | 10.5% | 0.49 | 0.81 | -33.3% | 0.93% | 1.62% | 3.8% | -0.42% |
| SPYbench | 11.9% | 19.7% | 0.60 | 0.85 | -51.5% | 1.80% | 3.04% | 0.0% | 0.00% |
| 60/40bench | 8.5% | 11.1% | 0.67 | 1.09 | -31.9% | 1.05% | 1.69% | 1.8% | -0.20% |
VaR and CVaR are daily historical loss estimates at the 95% confidence level, reported as positive loss magnitudes; drawdowns are shown as negative values. Annualized returns are geometric, from daily net returns; volatility and Sharpe scale daily figures by √252; Sortino uses downside deviation relative to a zero daily return threshold. Turnover is the monthly average of total absolute weight changes vs drifted pre-trade weights (initial cash buy-in excluded). Benchmarks are shown net of the same turnover cost model where rebalancing applies, the 100% SPY benchmark has no ongoing rebalancing turnover, so its cumulative cost drag is 0.00%.
Drawdown paths and tail loss
Volatility measures dispersion, but drawdown shows the path of losses from peak to trough. VaR and CVaR summarize downside tail risk.
Drawdown from running peak
Peak-to-trough losses through time · 1Y lookback · net of costs
Values below zero are losses from the prior wealth peak. Volatility measures dispersion; drawdown shows the actual path of losses an investor would have lived through.
Downside and tail risk
Historical daily measures at 95% confidence · 1Y lookback
| Series | Equal Weight | GMV | Max Sharpe | Regime-Aware | SPY * | 60/40 * |
|---|---|---|---|---|---|---|
| Max drawdown | -34.9% | -11.3% | -18.2% | -33.3% | -51.5% | -31.9% |
| Max 10-day drawdown | -17.8% | -7.3% | -8.8% | -15.2% | -26.8% | -16.7% |
| Historical VaR 95 (daily) | 1.01% | 0.32% | 0.75% | 0.93% | 1.80% | 1.05% |
| Historical CVaR 95 (daily) | 1.71% | 0.50% | 1.15% | 1.62% | 3.04% | 1.69% |
| Worst daily return | -5.91% | -1.74% | -4.43% | -5.99% | -10.94% | -5.52% |
| Annualized volatility | 11.2% | 3.4% | 7.3% | 10.5% | 19.7% | 11.1% |
* benchmark series
VaR and CVaR are daily historical loss estimates at the 95% confidence level, reported as positive loss magnitudes; drawdowns and worst-day returns are shown as negative values. VaR is a historical threshold estimate, not a worst-case guarantee; CVaR is the average loss beyond VaR, so CVaR ≥ VaR by construction. Max 10-day drawdown is the worst rolling two-week cumulative return.
Signal classification and allocation response
The regime engine tracks stock-bond correlation, CPI inflation, and VIX stress to classify allocation conditions. The panels below share one time axis: signal changed → regime changed → weights changed.
Signal-to-Allocation Stack
Three stress signals, the resulting regime classification, and the Regime-Aware strategy's allocation response
CPI YoY is lagged one month in the backtest to avoid look-ahead bias. The regime-aware strategy does not forecast stress, it reacts to observed, lagged indicators and changes defensive exposure when multiple signals are active. The bottom panel shows final capped sleeve weights after inverse-volatility allocation, so realized weights can differ slightly from the raw regime target. When the defensive regime is active, the growth sleeve target moves toward 45%; across the full sample, realized allocation shifts gradually because regimes change through time.
Crisis-window behavior
Strategy behavior during 2008, 2020, and 2022, periods chosen because each broke a different assumption. Rankings can change across regimes; that is the finding.
Combinations without full crisis coverage are labeled insufficient_data rather than estimated. The 2008 window is only fully covered at the 6M lookback because the common ETF history begins in 2007. Benchmarks are aligned to the same common strategy sample so crisis comparisons remain like-for-like, SPY is marked insufficient at longer lookbacks for alignment, not because SPY lacks data.
Global Financial Crisis
Jan 2008 - Dec 2009
- Stress mechanism
- A credit and banking crisis: equity, credit, real estate, and commodities sold off together while Treasuries rallied hard.
- Why it matters
- The classic test of whether duration actually diversifies an equity shock, and in this window, it did.
- Main lesson
- In this window, defensive allocations built on Treasuries held up; diversification across risk assets alone did not.
Cumulative wealth through the crisis
Rebased to $100 at crisis start · 6M lookback · net of costs
Only series with full coverage of the crisis window are shown. Strategy rankings during one crisis are historical observations, not predictions.
max drawdown -8.7%
max drawdown -34.9%
SPY crisis return-20.1%
60/40 crisis return-7.7%
Crisis risk metrics
2008 GFC · 6M lookback · daily net returns within the window
| Strategy | Crisis return | Volatility | Max DD | CVaR 95 | Worst day | vs SPY |
|---|---|---|---|---|---|---|
| Equal Weight | -2.0% | 20.5% | -34.9% | 3.03% | -5.29% | +18.1% |
| GMV | +4.3% | 4.9% | -8.7% | 0.65% | -1.58% | +24.4% |
| Max Sharpe | -4.1% | 13.9% | -26.4% | 1.94% | -6.20% | +16.1% |
| Regime-Aware | -3.0% | 19.4% | -33.5% | 2.99% | -5.17% | +17.2% |
| SPYbench | -20.1% | 34.8% | -51.9% | 5.27% | -9.84% | 0.0% |
| 60/40bench | -7.7% | 18.8% | -31.9% | 2.80% | -5.25% | +12.4% |
Correlation Breakdown & Crisis Hedge Lab
Diversification is conditional. This module measures when cross-asset correlations rise, how many independent bets remain under stress, whether portfolio risk collapses into one dominant factor, and which hedge sleeves actually reduce crisis drawdowns.
A portfolio can hold many tickers and still behave like one concentrated trade. Correlation estimates are sample-sensitive, crisis windows are historical observations, and hedge behavior is regime-dependent — this lab is a diagnostic, not a forecast or recommendation. VIX remains signal-only and is never used as an investable hedge.
Stress rising — correlations climbing
Defensive minus normal correlation, per asset pair
Each block is one asset-pair relationship; rows and columns are ETFs. Taller, warmer blocks mean the pair moved more together in the selected view, and the uplift view isolates where diversification weakened most under stress. Hover a block for the pair's normal, defensive, and uplift correlation — the exact values also appear in the heatmap and tables below.
Average pairwise correlation across the full book rose from 0.23 to 0.30 from normal to defensive regimes.
13 assets, but the correlation structure leaves only a few independent bets at the depth of this crisis.
Share of cross-asset variance explained by a single dominant factor in this crisis window.
Largest crisis-window drawdown reduction in 2022 Inflation, matching the dossier below.
Regime correlation matrices
Average pairwise correlation, normal vs defensive regimes · 1Y lookback
The difference matrix shows where diversification changed most under stress. Positive (warm) values mean the pair moved more together in defensive regimes than in normal ones; negative (blue) values mean the pair diversified better under stress. The strongest warm band sits where Treasuries meet equities — duration's hedge weakened — while broad commodities (DBC) cool, decoupling from the rest of the book.
Effective independent bets
Full universe · 1 / Σ pᵢ² from correlation eigenvalues · by lookback
The portfolio holds 13 assets, but the correlation structure can collapse the number of independent bets to only a few during stress.
Common-factor concentration
Stacked PC1 / PC2 / PC3 variance share (top of stack = top-3 cumulative) · full universe · 1Y lookback
When the PC1 band widens, cross-asset returns are being driven by a more concentrated common factor.
2022 Inflation / Rate-Hike Drawdown
Jan 2022 – Dec 2022 · 1Y lookback
- Stress mechanism
- An inflation and rate shock. Stocks and bonds fell together as the stock-bond correlation turned positive — the exact regime that breaks static 60/40 diversification.
- Main lesson
- Duration added to losses instead of offsetting them; inflation-sensitive real assets and cash were the sleeves that actually defended.
- Avg correlation
- 0.41
- Growth/risk corr
- 0.79
- Effective bets
- 3.3 / 13
- PC1 share
- 49%
Hedge effectiveness
90/10 crisis-window overlays · 2022 Inflation · 1Y lookback · sorted by drawdown reduction
| Base | Hedge | Return Δ | Drawdown reduction | CVaR reduction | Full-sample impact | Status |
|---|---|---|---|---|---|---|
| 60/40 | DBC | +3.3% | +3.5% | +0.13% | -0.74% | ok |
| Regime-Aware | DBC | +2.8% | +3.1% | +0.06% | -0.51% | ok |
| Equal Weight | DBC | +3.1% | +2.5% | -0.00% | -0.58% | ok |
| 60/40 | SHV | +1.7% | +2.0% | +0.21% | -0.69% | ok |
| Equal Weight | SHV | +1.5% | +1.8% | +0.18% | -0.49% | ok |
| Regime-Aware | SHV | +1.2% | +1.5% | +0.16% | -0.44% | ok |
| 60/40 | GLD | +1.6% | +1.2% | +0.16% | +0.20% | ok |
| Equal Weight | GLD | +1.3% | +1.0% | +0.06% | +0.35% | ok |
| 60/40 | TIP | +0.5% | +0.8% | +0.18% | -0.54% | ok |
| Regime-Aware | GLD | +1.1% | +0.7% | +0.09% | +0.43% | ok |
| Equal Weight | TIP | +0.2% | +0.6% | +0.11% | -0.35% | ok |
| 60/40 | IEF | +0.2% | +0.5% | +0.17% | -0.54% | ok |
| Regime-Aware | TIP | -0.1% | +0.3% | +0.12% | -0.29% | ok |
| Equal Weight | IEF | -0.1% | +0.3% | +0.12% | -0.34% | ok |
| Regime-Aware | IEF | -0.4% | -0.0% | +0.13% | -0.29% | ok |
| 60/40 | TLT | -1.5% | -1.0% | +0.15% | -0.51% | ok |
| Equal Weight | TLT | -1.8% | -1.3% | +0.08% | -0.31% | ok |
| Regime-Aware | TLT | -2.2% | -1.6% | +0.12% | -0.24% | ok |
A 10% sleeve is overlaid on a 90% base portfolio, rebalanced monthly and charged the same 5 bps turnover cost. Cash-like and duration hedges tend to help in equity-led crises; an inflation and rate shock favors different sleeves. Full-sample impact is the overlay's effect on full-period annualized return — positive when the sleeve adds return over the whole sample, negative when it costs return. A hedge is only useful when its crisis-window downside improvement justifies that full-sample impact.
Systematic factor exposure
Factor regressions help diagnose whether strategy returns reflect exposure to common risk factors. They do not prove investment skill by themselves.
Factor loadings by strategy
OLS coefficients on daily Fama-French factors · 1Y lookback strategies
The zero line separates positive from negative loadings. R² is shown in the table as model fit, not strategy quality. Benchmarks are omitted from the chart for readability.
How to read this. Market beta tells you how much of each strategy is simply equity exposure: the defensive strategies hold betas well below the equity benchmarks, which explains most of their drawdown behavior before any other factor matters.
The Fama-French model is used as an equity-factor exposure diagnostic. Because the strategies include bonds, TIPS, gold, commodities, credit, REITs, and cash proxies, the model does not fully explain all sources of multi-asset return and risk, lower R² for bond- and commodity-heavy strategies is expected, and is itself informative. A fuller multi-asset attribution model would also include term, credit, inflation, commodity, and real-asset factors.
The intercept is treated as residual return unexplained by this limited factor model, not as proof of investment skill. Factor regressions diagnose exposures; they do not prove skill or future alpha.
Factor exposure table
All series · 1Y lookback · daily Fama-French 5-factor regression · Intercept: diagnostic residual return, not proof of skill
| Strategy | Intercept (ann.) | Mkt β | SMB | HML | RMW | CMA | R² | Obs |
|---|---|---|---|---|---|---|---|---|
| Equal Weight | -0.2%t=-0.2 | 0.49 | 0.06 | 0.01 | -0.04 | 0.02 | 0.84 | 4569 |
| GMV | 0.1%t=0.1 | 0.08 | 0.00 | -0.03 | -0.01 | 0.01 | 0.20 | 4569 |
| Max Sharpe | 3.0%t=1.9 | 0.14 | 0.00 | -0.11 | -0.08 | 0.06 | 0.17 | 4569 |
| Regime-Aware | -0.8%t=-1.2 | 0.49 | 0.05 | 0.03 | -0.01 | 0.00 | 0.93 | 4569 |
| SPYbench | -0.4%t=-0.8 | 1.00 | -0.11 | 0.01 | 0.06 | 0.05 | 0.99 | 4569 |
| 60/40bench | 0.8%t=1.1 | 0.55 | -0.05 | -0.03 | 0.04 | 0.02 | 0.93 | 4569 |
The intercept is annualized from the daily regression intercept and treated as residual return unexplained by this limited factor model, not as proof of alpha or skill. In particular, the Max Sharpe intercept is not interpreted as tradable alpha: the regression is equity-factor-only, the strategy holds multi-asset exposures, and the figure is sensitive to lookback construction and transaction-cost assumptions. t-statistics are shown for the intercept only.
Turnover, concentration, and feasibility
Backtest-attractive is not the same as implementable. Weights through time, turnover, and concentration reveal whether a strategy could actually be run.
Allocation heatmap, Max Sharpe
Target weights at each monthly rebalance · 1Y lookback · rows are assets, columns are time
Brighter teal = higher weight (scale capped at the 35% constraint). Stable strategies show long horizontal bands; unstable ones show vertical striping as the optimizer jumps between assets.
Monthly turnover
Total absolute weight change at each rebalance · 1Y lookback
Turnover is measured against drifted pre-trade weights; the initial cash buy-in is excluded. Each unit of turnover costs 5 bps.
Concentration & implementability
Averages across all rebalances · 1Y lookback
| Strategy | Avg effective positions | Max single weight | Avg turnover /mo | Max turnover | Annualized turnover | Cumulative cost drag |
|---|---|---|---|---|---|---|
| Equal Weight | 13.0 | 7.7% | 2.6% | 9.5% | 31% | -0.29% |
| GMV | 3.6 | 35.0% | 7.7% | 42.7% | 92% | -0.84% |
| Max Sharpe | 3.6 | 35.0% | 38.4% | 137.7% | 461% | -4.11% |
| Regime-Aware | 5.6 | 35.0% | 3.8% | 19.4% | 46% | -0.42% |
Effective positions = 1 / HHI of weights. Optimized portfolios can look attractive in return metrics yet become unrealistic if they demand excessive turnover or concentration.
Transferable findings
This is a comparison, not a recommendation. The analysis studies how allocation rules behave under stated historical assumptions; the transferable lessons are about process.
Diversification is conditional
Stock-bond relationships change with the inflation and rate regime. A mix that hedged well for two decades added to losses in 2022. Diversification should be evaluated across regimes, not assumed from a long-run average.
Drawdowns matter more than averages
Two strategies with similar annualized returns can differ enormously in peak-to-trough losses. For most investors the path, and whether they can stay invested through it, matters as much as the endpoint.
Constraints make models usable
Unconstrained optimizers concentrate and churn. Weight caps, long-only rules, and transaction-cost awareness are not afterthoughts; they are what separates a paper portfolio from an implementable one.
Real-world individual, HNI, and institutional portfolios often consider diversifiers beyond public equities and bonds, cash, real assets, private credit, hedge funds, commodities, structured products. This analysis deliberately uses liquid ETFs as transparent public-market proxies rather than modeling illiquid private assets, whose data, valuation, and liquidity characteristics require different methods. The ETF proxies are imperfect representations of those broader exposures.
Boundary conditions
These results are historical and assumption-dependent. The analysis is only as good as its assumptions, so they are stated here rather than buried.
Backtest assumptions
- Historical ETF adjusted-close data (Yahoo Finance)
- Monthly rebalancing at month-end trading days
- Long-only portfolios, 35% per-asset weight cap
- 5 bps transaction cost per unit of turnover, measured against drifted pre-trade weights; initial cash buy-in excluded
- Geometric annualized returns; volatility and Sharpe scaled by √252 from daily figures
- VaR/CVaR reported as positive daily loss magnitudes at 95% confidence; drawdowns reported as negative values
- SHV as the optimizer's investable cash proxy; Fama-French RF for standardized performance reporting
- CPI signal lagged one month to avoid look-ahead
- VIX used as a signal only, never an allocation
- Crisis windows selected manually and documented
Data limitations
- Common ETF history begins 2007 (SHV, HYG inceptions), limiting GFC coverage for longer lookbacks
- ETF proxies imperfectly represent asset classes
- CPI data is released with delay and may be revised
- Yahoo Finance adjusted prices can be restated
Model limitations
- Expected-return estimates are noisy; sample means are used deliberately to expose this
- Covariance estimates are unstable across windows
- Regime thresholds are judgment calls that may overfit past crises
- The Fama-French model is an equity-factor diagnostic: it does not fully explain strategies holding bonds, TIPS, gold, commodities, credit, REITs, and cash; a fuller attribution would add term, credit, inflation, commodity, and real-asset factors
- Regression intercepts are residual returns under this limited model, not proof of alpha or skill
- Results are sensitive to start date, end date, and lookback choice
Implementation limitations
- Linear cost model ignores spreads, market impact, and capacity
- Taxes are ignored
- Liquidity is assumed sufficient at all rebalances
- No shorting, leverage, or derivatives are modeled
Automated research-engine checks
18 of 18 checks passed · ledger generated 2026-08-19
- data completenesspass
- return calculationpass
- weights sum and boundspass
- no lookahead biaspass
- transaction cost applicationpass
- turnover calculationpass
- var cvar conventionpass
- factor regressionspass
- regime signal availabilitypass
- cpi lagpass
- benchmark constructionpass
- data coverage honestypass
- correlation matrix boundspass
- correlation matrix symmetrypass
- effective bets boundspass
- pca share boundspass
- hedge effectiveness integritypass
- export completenesspass
warning · gfc partial data: The common ETF history begins in 2007 (SHV inception 2007-01, HYG inception 2007-04), so longer lookback windows do not cover the full 2008-2009 GFC window. For consistency, crisis comparisons use the common valid start date required by the selected strategy universe and lookback window; benchmarks are also aligned to that common sample and marked insufficient when they cannot be compared over the same full aligned window. Affected and reported as insufficient_data: Equal Weight (1Y), Global Minimum Variance (1Y), Max Sharpe / Risk-Adjusted Return (1Y), Regime-Aware Allocation (1Y), SPY (1Y), 60/40 (1Y), Equal Weight (3Y), Global Minimum Variance (3Y), Max Sharpe / Risk-Adjusted Return (3Y), Regime-Aware Allocation (3Y), SPY (3Y), 60/40 (3Y).