The evening shift
How it learns
Most market sites grade themselves on the one stock they told you about. We grade the whole exchange. Every evening, once the closing bell has rung, the day's move is recorded for every listed stock we track — the ones we picked, the ones we ranked and passed over, and the thousands we never mentioned. Then the morning's scoring is marked against it.
Sessions graded
1
Stock-days measured
6,773
Average rank correlation
-0.153
Why measure stocks we did not pick
A screen that quietly throws away a whole family of winners leaves no trace of that mistake in the stocks it kept. The evidence sits entirely in what it discarded. So the only way to find that kind of blind spot is to keep score on the discards too, which is what the evening run does.
That turns one graded call a day into a hundred graded rankings and thousands of measured stock-days. It is the difference between an opinion about how the day went and a dataset.
What each number means
- Rank correlation
- Did the stocks we scored highest actually outrun the ones we scored low? Measured against the market's own move that day, so a session where everything rose does not flatter us. Zero is a coin toss. Above zero, the ranking carries information.
- Decile spread
- The same idea in percentage points: our best-scored tenth of the day minus our worst-scored tenth.
- Pick vs. best candidate
- What the published stock did, next to the best stock we had already shortlisted. The gap is regret — value that was in our hands and not chosen.
- Top-20 recall
- Of the twenty biggest relative movers that were actually tradable that day, how many had made it into our shortlist at all. This is the one that grades the filter rather than the scoring. "Tradable" matters: the largest movers on any given session are usually two-dollar micro caps spiking on an announcement, which our screen rules out on price and liquidity by design. Judging it against those would be scoring it on a race it deliberately does not enter. The column beside it shows what a randomly drawn shortlist of the same size would have caught, because a recall without its own chance level is not a result, only a number.
Session by session
| Session | Stocks measured | Rank corr. | Decile spread | Pick | Best candidate | Top-20 recall | Chance would give |
|---|---|---|---|---|---|---|---|
| 2026-08-20 | 6,773 | -0.153 | -3.29% | +1.52% | +17.73% | 6/20 | 0.95/20 |
Take the data and check it yourself
Nothing above is worth much if you have to take our word for the arithmetic. So the raw rows are downloadable: one file with every closed session and its graded metrics, one with every candidate we scored that morning and what it went on to do. Load either into a spreadsheet or into pandas and you can recompute every figure on this site — including the ones that do not flatter us.
Sessions
One row per closed session: the published pick, its open-to-close return against SPY, the algorithm version that produced it, the grading metrics and the sealed fingerprint.
What we are not claiming
The measuring runs from day one; the retuning does not. Coefficients are only adjusted once there are enough graded sessions for a change to mean something, and when they are, it will be on the entire accumulated history rather than on yesterday — so the newest session counts for one in N, and the scoring converges instead of lurching about after every bad afternoon.
Until then the honest statement is the narrow one: the record is being built, in public, and every session that passes adds to it. The counter at the top of this page is the whole claim. It is also, deliberately, the easiest thing on this site to check.
Each morning's call is sealed and timestamped into the Bitcoin blockchain before the market opens, which is what stops any of this from being tidied up after the fact.