CopyGrade
The receipts

Our track record, published either way.

A due-diligence layer that will not grade itself is not one. This page keeps two ledgers, both computed live from the same data the verdicts use, and both published whether they flatter the model or not: what a copier actually realized after each score, and whether the calls — copyable, avoid, flagged — held up.

The honest headline first: three consecutive 30-day reads of the return table have found no rank correlation between the CopyGrade Score and a copier's realized return (Spearman ρ ≈ 0: −0.001 on 2026-07-01, −0.064 on 2026-08-01, −0.081 on 2026-09-01, the last over 47,102 matured outcomes). The score remains a documented v1 heuristic, not a proven predictor — we said so when the first read landed, and the model's weights are frozen until a signal appears. Why we published a null result · How the score is built

The receipts

Track record: scores vs. realized outcomes

Every score snapshot is graded later against what a simulated copier would actually have realized over the following weeks. These tables are computed live from that grading — they are the model's report card, published whichever way they point.

What the metric is: for each snapshot, a simulated copier mirrors only the positions the wallet opens after that moment (buy-time windowed), and the outcome is the realized return of the round-trips that close within the horizon, net of modeled fees. What it isn't: the realized figure counts only round-trips that close inside the horizon, which understates slow, patient strategies. Beside it we now show a mark-to-market estimate that also marks the positions still open at the horizon's end to current price — it exists for recently-matured cohorts and reads “—” for older windows (marking a long-closed window to today's prices would be a fiction, not a measurement). Dormant windows — where the wallet did nothing for a copier to mirror — are excluded, so the table is not diluted by no-observation rows. Only windows the sync observed end-to-end count, and a band is published only once it holds at least 30 outcomes.

7-day horizon
n = 57,415
Score band at snapshotnMedian returnMean returnMedian MTMMean MTM% positive
75–1001,2570.00%+2.69%0.00%−4.66%31%
50–7415,4740.00%+1.62%−0.35%−11.05%21%
30–4939,3510.00%+1.39%−1.04%−9.36%23%
0–291,3330.00%−0.92%−2.32%−16.44%7%
30-day horizon
n = 50,562
Score band at snapshotnMedian returnMean returnMedian MTMMean MTM% positive
75–1009370.00%−0.17%−2.66%−20.17%36%
50–7413,6640.00%+1.31%−0.30%−11.32%18%
30–4935,1940.00%+1.70%−1.73%−11.60%29%
0–297670.00%−2.03%−15.90%−26.01%12%
90-day horizon
n = 1,914
Score band at snapshotnMedian returnMean returnMedian MTMMean MTM% positive
75–100840.00%−5.75%−18.46%−29.51%38%
50–747490.00%+0.60%−1.84%−15.76%33%
30–491,0570.00%+2.24%−1.37%−13.79%29%

Until these tables show a persistent positive relationship, the Copy Score remains a documented v1 heuristic — our assessment methodology, not a proven predictor. Aggregates only; no individual wallet's outcome is published. Last computed Sep 4, 2026 · refreshed daily.

Do the calls hold up?

Liveness and persistence, by the band we assigned

Of the wallets we graded into each band on a day 30 and 90 days ago: how many are still trading now (a fill within the last 30 days), and how many still sit in that band. Below it, whether our farming-risk flags held. Aggregates only, over wallets we still track; a band publishes only past 30 wallets.

Why liveness matters more than it sounds: Polymarket's own study of the ten most-copied wallets (COPYCAT, April 2026) found at least two of them inactive since March. A grade on a wallet that then goes quiet is a grade no copier can use, whatever its number. “Still trading” means as our sync records it: a wallet our index no longer covers reads as quiet even if it trades elsewhere. What this is not: a return. A wallet still trading in the same band has not necessarily made its copiers money; the table above is where that question is answered.

Graded 30 days ago
sampled Aug 5, 2026 · n = 3,132
Score band thennStill trading in our indexStill in that band
75–1003256.3%75.0% of 32 still tracked
50–741,09745.5%94.1% of 1,097 still tracked
30–491,93857.9%93.7% of 1,938 still tracked
0–296543.1%83.1% of 65 still tracked
Farming-risk flags
Of the 1,434 wallets our model flagged that day, 98.5% of the 1,434 still tracked are flagged today. Of the 1,698 it read as clean, 2.2% of the 1,698 still tracked have since been flagged (a “—” means fewer than 30 were still tracked, so the share is withheld). A flag that holds is a call; one that flickers is noise — both are ours to own.
Graded 90 days ago
sampled Jun 6, 2026 · n = 311
Score band thennStill trading in our indexStill in that band
50–7410134.7%84.2% of 101 still tracked
30–4918546.5%83.2% of 185 still tracked
Farming-risk flags
Of the 141 wallets our model flagged that day, 88.7% of the 141 still tracked are flagged today. Of the 170 it read as clean, 5.3% of the 170 still tracked have since been flagged (a “—” means fewer than 30 were still tracked, so the share is withheld). A flag that holds is a call; one that flickers is noise — both are ours to own.

Computed Sep 4, 2026 23:05 UTC · refreshed hourly · risk-assessment language throughout: a flag is our model's read of public trade patterns, never an accusation of conduct.

How to read this page

Three rules we hold ourselves to

  • Pre-committed framing. The return table's definition and its publication floor shipped before the first cohort matured, so the numbers could not be shaped after the fact.
  • Floors, not cherry-picks. A band appears only past 30 wallets; a small cell reads as signal it is not.
  • Misses stay up. Every change to the model is a dated entry on the changelog, and a null result is a result.