I want to start with the number, because it's the whole reason I'm writing this. LLY's profit factor read 1,156. Not 1.156. Eleven hundred and fifty-six. If you've spent any time backtesting you already know what that number means: not "found an incredible edge," but "found a bug or a coincidence and haven't caught it yet."
I hadn't caught it yet. It sat in an overrides file I trusted for months before I ran a proper re-validation pass — 2-year time-stability, quarter by quarter, instead of one aggregate number calculated over the whole window. LLY's PF didn't just come down. It came down to 1.92, and the quarterly breakdown showed exactly why: 13 in the first quarter, decaying to 0.7 by the fourth. The aggregate number was an average dragged upward by one freak early quarter on a small sample. By the time anyone would have traded it live, the edge was already dead.
LLY wasn't the only one
Once I stopped trusting single aggregate numbers, the rest of the shortlist didn't hold up much better. ZRO went from a profit factor of 6.55 to 0.88 under the same validation — and 0.88 means net-negative. Not "smaller edge." No edge, once measured honestly. MRVL, LAB, GOOGL and AVAAI all failed the same way: fake edges traced back to small-sample cherry-picks, the kind of result you get when a backtest window happens to contain a handful of trades that all went the same way by chance.
Seven candidates went into that validation pass. Three came out the other side.
| Symbol | PF | WR | Trades | Verdict |
|---|---|---|---|---|
| MEGAUSDT | 3.31 | 98.0% | 2,000 | Stable, holdout > train |
| EWJUSDT | 1.41 | 56.0% | 124 | Real, too small to matter much |
| NAORIS | 1.17 | 87.0% | 2,000 | Positive but decaying |
| LLY | 1156 → 1.92 | — | — | Fake, edge dead by q4 |
| ZRO | 6.55 → 0.88 | — | — | Fake, net −$177 |
What a real edge looks like instead
MEGAUSDT is the one I trust, and the reason isn't just that its profit factor (3.31) is high. Plenty of fake edges have high profit factors too — LLY's was 300+ times higher. What makes MEGA different is the direction it moved when I split it out of sample: trained PF was 1.66, holdout PF came in at 3.38. It got better on data it hadn't seen. That's the opposite of what an overfit result does, and across 2,000 trades and 5 stable templates it held up the same way in every regime I checked except one mixed BEAR/LOW-vol slice. It's the only pair in this batch I've marked trust_level HIGH.
EWJ is the honest small case. A profit factor of 1.41 isn't exciting, but it's stable across all four quarters on real trades — it just doesn't have the sample size or the liquidity to be more than a footnote. NAORIS is the one that worries me a little: PF 1.17 sounds fine until you look at the quarterly trend — 1.14, then 1.33, then 1.25, then 0.98. Still on the right side of breakeven, but the line points at zero, and I'm not adding size to something that's actively decaying just because this quarter's number is still positive.
What I check now before I trust a profit factor
Not the aggregate number. The quarter-by-quarter breakdown, first. A two-year backtest with a single PF calculated across the whole window can hide an edge that died eighteen months ago and is only still showing positive because one early quarter was strong enough to drag the average up. LLY is the extreme version of that — PF=1156 was never "an edge," it was one lucky quarter wearing a costume.
The second check is holdout versus training, and which direction it moves. MEGA got better out of sample. Everything that got worse out of sample is off the list, no matter how good the headline number looked going in. I run every profit-factor claim, mine or anyone else's, through the profit factor calculator with the actual trade counts before I let a single aggregate number talk me into anything, and I check what a string of losses against that "edge" would actually do to an account with the risk of ruin calculator.
What to check before you trust a backtest's profit factor
- Split it into quarters before you look at the aggregate. An average PF over two years can hide an edge that's been dead for eighteen months.
- Check which direction holdout moves versus training. A real edge should hold or improve out of sample. LLY and ZRO both got dramatically worse.
- Distrust absurd numbers on principle. A PF over roughly 5-10 on real market data is a reason to dig for a small-sample artifact, not a reason to celebrate.
- Weigh sample size against liquidity. EWJ is real at 124 trades but too illiquid to matter. NAORIS has 2,000 trades but a visibly decaying trend. Neither number alone tells the full story.
FAQ
How can a backtest show a profit factor of 1,156?
Almost always a small sample. A handful of lucky trades in a narrow window can produce a profit factor that looks astronomical because there are barely any losing trades to divide by. LLY showed PF=1156 on an early pass, and the number wasn't a bug, it was just measuring a tiny, cherry-picked slice of history that happened to go one direction. The fix isn't a better formula, it's more data across more time.
What does 2-year time-stability validation actually check?
It splits two years of results into quarters and recomputes profit factor in each one separately, instead of trusting a single number calculated across the whole period. A real edge holds up quarter to quarter. LLY's quarterly PF went 13 in the first quarter down to 0.7 by the fourth — the aggregate PF=1156 was an average dominated by one freak early quarter, and the edge was already dead by the time anyone would have traded on it live.
Which pairs actually survived validation, and why?
Three, out of a shortlist that included LLY and ZRO among others. MEGAUSDT passed cleanly: PF 3.31 on 2,000 trades, stable across 5 templates, and its holdout-period PF (3.38) actually came in higher than its training PF (1.66) — the opposite direction a fake edge moves. EWJUSDT is real but tiny: PF 1.41, stable across all 4 quarters, but only 124 trades and very low liquidity, so it barely moves the needle. NAORIS is the one I'm watching, not trusting: PF 1.17 with quarterly PF sliding 1.14 to 1.33 to 1.25 to 0.98 — still positive, but the trend line points at zero.
What do I check before trusting a backtest's profit factor now?
The quarter-by-quarter breakdown before the aggregate number. A single profit factor calculated over two years can hide an edge that died 18 months ago and is being propped up by one strong early quarter. I split every candidate into quarters, check the direction it's moving (not just whether it's currently positive), and separately re-run holdout-versus-training to make sure the edge got stronger out of sample, not weaker.