A due-diligence layer that will not grade itself is not one. This page keeps two ledgers, both computed live from the same data the verdicts use, and both published whether they flatter the model or not: what a copier actually realized after each score, and whether the calls — copyable, avoid, flagged — held up.
The honest headline first: three consecutive 30-day reads of the return table have found no rank correlation between the CopyGrade Score and a copier's realized return (Spearman ρ ≈ 0: −0.001 on 2026-07-01, −0.064 on 2026-08-01, −0.081 on 2026-09-01, the last over 47,102 matured outcomes). The score remains a documented v1 heuristic, not a proven predictor — we said so when the first read landed, and the model's weights are frozen until a signal appears. Why we published a null result · How the score is built
What the metric is: for each snapshot, a simulated copier mirrors only the positions the wallet opens after that moment (buy-time windowed), and the outcome is the realized return of the round-trips that close within the horizon, net of modeled fees. What it isn't: the realized figure counts only round-trips that close inside the horizon, which understates slow, patient strategies. Beside it we now show a mark-to-market estimate that also marks the positions still open at the horizon's end to current price — it exists for recently-matured cohorts and reads “—” for older windows (marking a long-closed window to today's prices would be a fiction, not a measurement). Dormant windows — where the wallet did nothing for a copier to mirror — are excluded, so the table is not diluted by no-observation rows. Only windows the sync observed end-to-end count, and a band is published only once it holds at least 30 outcomes.
| Score band at snapshot | n | Median return | Mean return | Median MTM | Mean MTM | % positive |
|---|---|---|---|---|---|---|
| 75–100 | 1,257 | 0.00% | +2.69% | 0.00% | −4.66% | 31% |
| 50–74 | 15,474 | 0.00% | +1.62% | −0.35% | −11.05% | 21% |
| 30–49 | 39,351 | 0.00% | +1.39% | −1.04% | −9.36% | 23% |
| 0–29 | 1,333 | 0.00% | −0.92% | −2.32% | −16.44% | 7% |
| Score band at snapshot | n | Median return | Mean return | Median MTM | Mean MTM | % positive |
|---|---|---|---|---|---|---|
| 75–100 | 937 | 0.00% | −0.17% | −2.66% | −20.17% | 36% |
| 50–74 | 13,664 | 0.00% | +1.31% | −0.30% | −11.32% | 18% |
| 30–49 | 35,194 | 0.00% | +1.70% | −1.73% | −11.60% | 29% |
| 0–29 | 767 | 0.00% | −2.03% | −15.90% | −26.01% | 12% |
| Score band at snapshot | n | Median return | Mean return | Median MTM | Mean MTM | % positive |
|---|---|---|---|---|---|---|
| 75–100 | 84 | 0.00% | −5.75% | −18.46% | −29.51% | 38% |
| 50–74 | 749 | 0.00% | +0.60% | −1.84% | −15.76% | 33% |
| 30–49 | 1,057 | 0.00% | +2.24% | −1.37% | −13.79% | 29% |
Until these tables show a persistent positive relationship, the Copy Score remains a documented v1 heuristic — our assessment methodology, not a proven predictor. Aggregates only; no individual wallet's outcome is published. Last computed Sep 4, 2026 · refreshed daily.
Why liveness matters more than it sounds: Polymarket's own study of the ten most-copied wallets (COPYCAT, April 2026) found at least two of them inactive since March. A grade on a wallet that then goes quiet is a grade no copier can use, whatever its number. “Still trading” means as our sync records it: a wallet our index no longer covers reads as quiet even if it trades elsewhere. What this is not: a return. A wallet still trading in the same band has not necessarily made its copiers money; the table above is where that question is answered.
| Score band then | n | Still trading in our index | Still in that band |
|---|---|---|---|
| 75–100 | 32 | 56.3% | 75.0% of 32 still tracked |
| 50–74 | 1,097 | 45.5% | 94.1% of 1,097 still tracked |
| 30–49 | 1,938 | 57.9% | 93.7% of 1,938 still tracked |
| 0–29 | 65 | 43.1% | 83.1% of 65 still tracked |
| Score band then | n | Still trading in our index | Still in that band |
|---|---|---|---|
| 50–74 | 101 | 34.7% | 84.2% of 101 still tracked |
| 30–49 | 185 | 46.5% | 83.2% of 185 still tracked |
Computed Sep 4, 2026 23:05 UTC · refreshed hourly · risk-assessment language throughout: a flag is our model's read of public trade patterns, never an accusation of conduct.