CopyGrade
← Guides
Strategy

Copying a season vs copying a bracket: why sample size decides how much to trust a sports wallet

Updated August 28, 2026 · By Simon Lee — CopyGrade Research

A season gives you a sample; a bracket gives you a story. A wallet that trades the daily MLB slate can put hundreds of comparable, independently resolved markets behind its record in one summer, while a wallet that "called" a knockout tournament has at most a handful of correlated bets behind the same-looking profit line — and no number of them adds up to evidence of a repeatable edge. That difference decides how much a copier should trust a record, and therefore how much money to put behind it. Across the nine sports boards we cut on 2026-08-28, 15.4% of the wallets that traded the MLB board have 20 or more settled MLB markets; on the World Cup board the figure is 5.9%, and the typical World Cup wallet settled exactly one.

Why does a season produce evidence a bracket can't?

Because evidence is counted in resolutions, and the two formats resolve on different scales. An MLB team plays 162 regular-season games and the league settles 2,430 of them; an NBA team plays 82 (1,230 league-wide); even the NFL, the shortest major season, gives each team 17 games across an 18-week slate and the market 272 game lines a year. A 48-team World Cup settles 104 matches in total, and the champion plays eight of them. A wallet that follows one team through a bracket cannot generate more than eight game results, however sharp it is — the format caps the sample before the trader's skill gets a say.

The second difference is independence. Game lines across a season are close to repeated draws from the same distribution: same market type, same resolution rule, a fresh result every day. Bracket markets stack on each other. "Team X to advance," "Team X to win the title," and "Team X to win Sunday" are one opinion sold three times, and they resolve together — a copier holding all three has one position, not three. The World Cup final made the point the expensive way: the 90-minute match markets went to zero while the advance market on the same team paid in full.

Here is what that looks like in the graded universe. For each board, the table counts every wallet in our universe with at least one fill on the board's markets and asks how many distinct markets that wallet has settled there (as of 2026-08-28, boards are tag-derived, a market counts as settled once it is past its expiry, and we hold up to the 10,000 most recent fills per wallet).

BoardWallets with a fillMedian settled markets90th percentile20+ settled50+ settled100+ settled
Tennis2,26126519.8%11.7%7.4%
MLB2,47124015.4%8.9%5.3%
NBA4,1393177.9%3.2%1.1%
World Cup1,2941125.9%1.9%0.6%
Champions League2,6443105.3%1.8%0.5%
Golf721174.6%1.8%1.0%
Premier League2,703294.4%1.1%0.3%
NHL1,792294.1%1.2%0.2%
NFL2,437292.8%0.7%0.2%
Share of a board's wallets with 20+ settled markets on that board — 2026-08-28
Each bar is the share of wallets with at least one fill on the board that have settled 20 or more distinct markets there, as of 2026-08-28. Tennis 19.8%, MLB 15.4%, NBA 7.9%, World Cup 5.9%, Champions League 5.3%, Golf 4.6%, Premier League 4.4%, NHL 4.1%, NFL 2.8%. Amber marks the boards whose format is a knockout or weekly tournament (World Cup, Champions League, golf); blue marks league seasons and the daily tennis slate. Boards are tag-derived; a market is settled once past its expiry. Source: CopyGrade graded universe.

Two things to read off it. The daily-slate sports — tennis, with a match somewhere almost every day of the year, and baseball — are the only two of the nine boards where more than one wallet in ten has a 20-market settled record, and the only two where a meaningful cohort (7.4% and 5.3%) has settled a hundred. The tournament boards sit where the format puts them: the 90th-percentile World Cup wallet settled 12 markets, the 90th-percentile golf wallet 7. And the NFL row is a reminder that the calendar matters as much as the sport — twelve days before kickoff the board is futures and props, so even its best-sampled wallets have single-digit settled game markets until the season is under way. Board figures recompute with every sync; treat these as the late-August snapshot they are.

How many resolved markets does it take to tell skill from luck?

More than almost any bracket can supply. The arithmetic is the plain binomial one: the uncertainty around a hit rate measured on n resolved bets is about ±1.96 × √(p(1−p)/n) at 95% confidence, which for a coin-flip-shaped record is:

Resolved bets95% error bar on the hit rate
8 (one team's tournament run)±35 points
20±22 points
50±14 points
100±10 points
400±5 points
1,000±3 points

A record built on eight resolved bets cannot distinguish a 55% trader from a 45% one. A 5-point edge — a large one in a fee-paying market — does not clear its own error bar until roughly 400 resolved bets, and a properly powered test wants closer to 800. Prediction markets are not fair coins, so the right measurement is edge against the price paid rather than a raw hit rate, but the law is the same: the noise on the average shrinks with the square root of the sample, and nothing about a trader's confidence changes the exponent. Reading a wallet like a quant walks the edge-versus-price version.

Against that bar, the graded universe is thin. Of the 8,417 wallets we had graded on 2026-08-28, the median had 25 closed round-trips; 24.1% had 100 or more and 5.4% had 400 or more. Our own confidence band reflects this: a CopyGrade Score does not reach High confidence below 120 fills, 40 closed round-trips and 30 days of watched history — enough to read realized returns, and deliberately not presented as enough to certify a 5-point edge. The methodology explains how the score weighs a thin record.

What does sample size do to position sizing?

It shrinks it, and it should shrink it before conviction gets a vote. Every sizing rule that isn't a guess — Kelly and its fractions included — scales the stake with the estimated edge, and the estimate comes with the error bar from the table above. On 20 resolved bets that bar spans more than 40 points, so the interval around any plausible edge contains zero; the honest sizing answer is small or nothing, whatever the point estimate says. On 400 the interval is narrow enough that a fraction of the estimated edge is a defensible stake. The rule of thumb is to size a copy allocation on the lower bound of what the record can support, not on the record's midpoint — and to let the allocation grow as the sample does, rather than sizing up because a run looks good.

Two practical consequences follow. First, a bracket run should never raise a wallet's allocation, because it adds correlated outcomes rather than independent ones — the position count went up, the evidence barely moved. Second, the drawdown-based rule in position sizing for copy trading needs the same haircut: a maximum drawdown measured over a 30-day tournament window is a floor on how bad it gets, not an estimate. The Copy Simulator replays a candidate's fills under your own capital and latency, and the sample it replays is the same sample the score saw — if the replay covers a fortnight, that is what you know.

Why is a tournament record the weakest evidence in sports?

Because four things go wrong at once, and each would be enough on its own.

  1. The format caps the sample. A team's run is at most eight matches in a World Cup, seven games in a best-of-seven series, three in a Wild Card round. Nobody's October series record is a sample, and no amount of conviction adds resolutions the schedule did not play.
  2. You see the survivors. After a tournament, the wallets that are visible are the ones whose bracket hit. 1,294 wallets traded the 2026 World Cup board and the median settled a single market there; the handful with a spectacular profit line are exactly what a field that size would produce by chance, and the leaderboard shows you the top of that distribution, not the distribution.
  3. The bets are correlated. Advance, champion, and match markets on the same side resolve together. A wallet's ten winning tournament positions may be two opinions, and the World Cup final showed the same opinion resolving in opposite directions across market types.
  4. Attention attracts manufactured records. Event windows are where farmed records get built because that is where copiers are looking. When we graded the World Cup board on 2026-07-07, 94% of the top 50 wallets by tournament profit carried a farming-risk flag — our model's algorithmic risk assessment from public trade history, not a claim about anyone's conduct, and a dated snapshot rather than a trend. The farming check is the first veto for a reason.

None of this says a sharp trader cannot trade a bracket well. It says a bracket cannot show you that they did.

How should a copier use a season and a bracket differently?

  1. Vet on the season, act during the bracket. Grade a candidate on their resolved game-market record — the Wallet Scout board filters and a verdict's track record show how many markets stand behind the score — and treat everything that happens after the bracket starts as out-of-sample in both directions. A 6-2 October is not evidence, and neither is a 2-6.
  2. Separate the market types before you read the profit. The category-edge breakdown on a verdict shows whether a wallet's return came from the daily slate or from one futures hit. The first can survive copying; the second is variance wearing a costume.
  3. Size on the sample, not the story. Use the error-bar table above as the ceiling on how much of an allocation a record can justify, and prefer the drawdown-based rule measured over the longest window the wallet has.
  4. Watch the between-seasons gap. A record ages. A wallet whose whole sample is last season and whose bracket bets are this month's is a wallet you are copying on stale evidence; when to stop copying covers the decay signals.
  5. Let a change in behaviour be the alert. A wallet that triples its per-fill size on "obvious" bracket spots is tilting, not sharpening. Alerts on a vetted candidate tell you the moment its fills stop looking like the season that earned it the grade.

The base rate is the honest frame for all of it: 1.1% of actively-traded wallets in our July 2026 snapshot passed every test a copier should apply, and a liquid tournament does not change the denominator — it adds attention and variance around the same small number of genuinely sharp wallets. A season lets you find them. A bracket mostly lets you find the lucky.

Read the methodology← All guides