Last Updated: August 2026
Disclaimer: This article is for informational purposes only and is not financial advice. Crypto trading involves significant risk of loss. Never trade with money you cannot afford to lose. Always do your own research (DYOR).
GhostCopy scanned 23,495 traders across Hyperliquid, OKX, and Polymarket public leaderboards. Of 150 candidates deep-judged after a forensic pre-filter, 45 cleared the naive probabilistic skill test — including one whose score came in at a perfect 1.0, meaning 100% certainty of skill by that standard. Applied to the same 45 candidates, the Deflated Sharpe Ratio — which penalises for the 23,495-person search that produced them — returned a maximum score of 0.0692 against a bronze bar of 0.60. Zero candidates cleared it. This article explains what that gap means, what it doesn't mean, and how to run the same test yourself.
Why "45 Traders Look Skilled" and "Zero Are Statistically Skilled" Can Both Be True
Before getting into mechanics, it is worth naming the paradox that every leaderboard contains: the number at the top of a ranking is a function of two things — the actual quality of the traders in the field, and the size of the field itself. Most people think about only the first. The Deflated Sharpe Ratio forces you to think about both.
Here is the problem restated without jargon. If you flip a fair coin one time, you expect to get some heads and some tails. If you flip it one million times and photograph only the single longest streak, the resulting photograph looks like a very lucky coin — but the coin hasn't changed. You just searched a very large space and kept the tail. The "luck" in the image belongs to the search, not the object photographed.
Public leaderboards work the same way. Every exchange sorts traders by realized performance and presents the top of the list. Nothing about that process screens for whether performance reflects skill or whether it reflects what pure randomness inevitably produces across a large enough population. The person you see at rank one on Bybit's copy-trading board may be the most skilled active trader in crypto. They may also be the luckiest member of a very large random walk. The raw leaderboard cannot tell you which.
The standard tool for answering this question is the Probabilistic Sharpe Ratio — and for a single candidate, evaluated in isolation, it is a good tool. It adjusts for track-record length and for the fact that real trading returns are not normally distributed. What it does not do is adjust for how many candidates you searched through before you landed on this one. That is the single correction the Deflated Sharpe Ratio adds, and as the scan results show, it changes the answer completely.
Forty-five candidates clearing the naive 95% bar and zero clearing a 60% bar after deflation is not a contradiction. It is the expected outcome of applying a corrected test to the winners of a 23,495-person search. The naive test saw 45 skilled-looking traders. The deflated test saw 45 candidates whose records are consistent with being the luckiest subset of a very large crowd.
Free: Crypto Trading Platform Cheat Sheet
Side-by-side fee comparison, ratings, and quick-pick recommendations for every major exchange and trading bot. Save hours of research.
No spam. Instant download on the next page.
PSR and DSR: The Same Metric, Decades Apart in Honesty
Both tests trace back to a Sharpe ratio — return divided by volatility — and both ask a version of the same question: "Is this trader's observed performance significantly better than chance?" They diverge on what they count as the baseline chance you are comparing against.
The Probabilistic Sharpe Ratio (PSR) was introduced by Bailey and López de Prado in 2012. Its key innovation over the raw Sharpe ratio is that it accounts for the shape of the return distribution. The ordinary Sharpe ratio secretly assumes returns are normally distributed. Real trading returns rarely are: they are typically negatively skewed (steady small gains, infrequent large losses) and fat-tailed. Both of those distortions make an observed Sharpe ratio look more impressive than it really is. PSR adjusts for skewness and kurtosis, and it adjusts for track-record length — a 30-day Sharpe is much noisier than a 3-year Sharpe and should be weighted accordingly. The result is a probability: PSR = 0.97 means "there is a 97% chance this track record reflects a non-zero true Sharpe."
What PSR does not adjust for is the number of candidates you evaluated before choosing this one. It treats the trader in front of it as if they were the only candidate ever considered.
The Deflated Sharpe Ratio (DSR) was introduced by the same authors in 2014, and it makes exactly that correction. Instead of testing whether the observed Sharpe is greater than zero, it tests whether the observed Sharpe is greater than the expected maximum Sharpe you would see from pure luck given how many candidates were searched. The more candidates you scan, the higher the luck baseline rises, because the best of a large random field will always outperform the best of a small one. DSR places the trader's result against that inflated baseline rather than against zero.
The gap between the two tests is not a statistical curiosity. It is the gap between asking "is this trader better than nothing?" and asking "is this trader better than the luckiest member of the crowd we searched?" When the crowd is 23,495 people, those are very different questions.
| What the test measures | PSR | DSR |
|---|---|---|
| Adjusts for return skewness | Yes | Yes |
| Adjusts for return kurtosis | Yes | Yes |
| Adjusts for track-record length | Yes | Yes |
| Adjusts for the number of candidates searched | No | Yes |
| Baseline compared against | Zero Sharpe | Expected max Sharpe given search size |
| DSR is always ≤ PSR | — | Always |
| Appropriate when evaluating one candidate | Good | Equivalent to PSR |
| Appropriate when evaluating from a leaderboard of thousands | Overconfident | Correctly calibrated |
The practical implication of the bottom two rows: PSR and DSR give similar answers when you are evaluating a single, referred candidate. They diverge massively — as the scan shows — when you are working with the output of a large ranked search.
The Scan: 23,495 Traders, Three Venues, One Question
GhostCopy's judgment system as of 2026-08-01 covers three public leaderboard venues: Hyperliquid, OKX, and Polymarket. Across those three platforms, the active population of ranked traders with sufficient public data for evaluation was 23,495.
The single question the scan asked was: does any of these traders show statistically significant skill after accounting for the fact that 23,495 candidates were searched?
The process ran in two stages. The first was a forensic pre-filter: a set of gates designed to eliminate candidates whose records are structurally unreliable before expensive statistical computation begins. These gates examine things like minimum trade count, track-record length, and return patterns that are inconsistent with genuine trading. Of the 23,495 total, 150 candidates passed the pre-filter and were deep-judged. Of those 150, 33 were eliminated by a forensic gate during deep judging, leaving 117 structurally clean records to evaluate statistically.
Those 117 then received both a PSR computation and a DSR computation. The PSR used a standard benchmark Sharpe of zero with a 95% confidence threshold — the most common setup in academic work and the closest equivalent to how a retail investor thinks about "is this person skilled?" The DSR used N = 23,495 to set the expected-maximum-Sharpe baseline, with a bronze clearance bar of 0.60.
The results, in summary form:
| Stage | Count |
|---|---|
| Total traders scanned | 23,495 |
| Deep-judged after pre-filter | 150 |
| Clean after forensic gates | 117 |
| Killed by a forensic gate | 33 |
| Clearing PSR > 95% (naive test) | 45 |
| Clearing DSR ≥ 0.60 (deflated test) | 0 |
| Best DSR in population | 0.0692 |
| Best PSR in population | 1.0 |
The Forensic Gates: Why 33 Candidates Were Eliminated Before the Statistics Started
The 33 candidates killed by a forensic gate are worth pausing on, because they represent a class of failure that is distinct from statistical luck. These are track records with structural problems that invalidate the statistical computation before it begins — not traders who looked skilled and failed the DSR, but traders whose records were not valid inputs for either test.
Common examples of the type of issue forensic gates target include track records that are too short to produce a meaningful estimate (a few dozen trades does not give enough observations to measure skewness or kurtosis reliably), return series that show suspicious patterns of serial correlation inconsistent with genuine market exposure, or records where the displayed statistics don't cohere with the underlying trade log.
The result of removing these 33 is that the 117 clean candidates represent the highest-confidence subset of the deep-judged population — the traders whose records are structurally sound enough to take the statistical question seriously. The fact that none of them cleared the DSR bar is therefore not an artifact of garbage input data. The 45 who passed the PSR test had clean, genuine track records. The deflated test found them insufficient regardless.
The Results: The Cliff Between PSR and DSR
Forty-five candidates cleared a PSR above 0.95. That means 45 traders, on a naive reading, have records showing a greater-than-95% probability of genuine skill. One of those candidates has a PSR of exactly 1.0 — a number that, interpreted naively, means 100% certainty of skill. These are not borderline results. These are records that, evaluated in isolation, would clear the bar you would use if you were evaluating a single referred candidate.
Applied to the same 45 candidates, using N = 23,495 to set the DSR baseline, the highest score returned was 0.0692.
That is not a near-miss. The bronze bar is 0.60. The best candidate in the full population scored 0.0692. The gap is not 5 points or 10 points — it is a factor of eight. The candidate who came closest to the DSR bar is, by that test, structurally indistinguishable from the luckiest member of a large random field. Their PSR of 1.0 became a DSR of 0.0692 because the test properly asked: "Is this trader better than what pure chance would produce across 23,495 people?" The answer, for all 45 candidates, was no.
This is not a failure of the traders. This is the expected mathematical output of testing the winners of a large population search. The expected maximum Sharpe ratio from pure luck across 23,495 traders is high enough that the observed records cannot clear it. The leaderboards, by construction, are presenting the extreme right tail of a massive distribution — and that tail, in the absence of very strong persistent skill, is largely manufactured by variance.
What 0.0692 Actually Means
It helps to be precise about what the DSR score means numerically. A DSR of 0.0692 is a probability of approximately 7% — specifically, the probability that the best trader in the scanned population has a genuine positive true Sharpe after accounting for the size of the search.
To be clear about what that does and does not say: a 7% probability is not zero. It is not proof that this trader lacks skill. It is a statement that, given the evidence available (track record length, return shape, and the fact that 23,495 candidates were searched to find this one), the record is consistent with luck roughly 93% of the time and consistent with skill roughly 7% of the time. A decision-maker trying to allocate capital responsibly would need that number to be significantly higher before acting on it.
The intuition for why the score is so low even for a PSR = 1.0 candidate: when the search space is large, the expected best-of-the-field Sharpe under the null hypothesis of pure luck is very high. The candidate's observed Sharpe must not merely exceed zero — it must exceed this inflated luck baseline by enough to be statistically significant. In a field of 23,495, the luck baseline is high enough that even a record which looks overwhelmingly impressive in isolation cannot clear it with the data available.
That is not a broken test. That is the correct answer. The appropriate response when a test returns "insufficient evidence" is not to lower the bar. It is to gather more data — longer track records, more observations per period — until the test can resolve the question.
This Does Not Prove That No One Can Trade
It is important to say clearly: a result of zero candidates clearing the DSR bar does not mean that none of the 23,495 traders in the population has genuine skill.
It means that the available track records, evaluated against the size of the search that produced them, do not provide statistically sufficient evidence to distinguish skilled traders from the lucky tail of a large random field. Those are meaningfully different claims.
Several things are consistent with this result. A genuinely skilled trader with a shorter track record will score low on both PSR and DSR, because there are simply not enough observations to resolve the signal from the noise yet. A skilled trader on a venue like Polymarket, where prediction markets have different statistical properties than perpetual futures, may require a different test parameterization to surface. A trader with real edge who also runs high position concentration will show inflated volatility that compresses their observed Sharpe below what their true skill warrants.
The DSR is a test designed around a specific null hypothesis and a specific kind of evidence. It cannot speak to forms of skill the observed data does not contain. A trader who made a single enormous informed bet, got it right, and then stopped — accumulating a very short track record — may be the most skilled person on any of these leaderboards. The DSR would score them poorly, not because they lack skill, but because one trade is not enough observations to make the statistical case.
What the result does say, and says firmly, is that the records publicly visible on these leaderboards — taken as they are, evaluated in the context of how many traders were searched to surface them — do not clear the bar for a conservative allocation decision. If someone is copying traders based on leaderboard rank alone, the deflated evidence for that decision is a maximum 7% confidence. That should cause a serious reconsideration of how the selection is being made.
How to Apply the Same Test Yourself
The full DSR computation requires access to a trader's periodic return series and a choice of N — the number of candidates searched to surface them. Neither piece of information is straightforwardly available inside a copy-trading app. But the DSR's logic can be applied by hand in a way that captures most of the value.
Step 1: Establish the field size. Before evaluating any individual trader, ask how many traders the platform ranks. This is often findable in the platform's documentation or by counting the total listed on the leaderboard. Bybit's copy-trading board (explore Bybit copy trading →), Binance's lead trader system (see Binance lead traders →), and Bitget's copy-trading tab (browse Bitget copy trading →) each surface thousands to tens of thousands of traders. That field size is the N in your mental DSR.
Step 2: Adjust your required evidence accordingly. In a field of 1,000 traders, a 12-month track record with a 1.5 observed Sharpe is modest evidence of skill. In a field of 25,000 traders, you should need either a much longer track record, a much higher observed Sharpe, or both. The bar rises with N.
Step 3: Require long, unfavorable-period track records. The PSR adjustment for track-record length is the part you can apply directly without computation. A 30-day or 90-day record is nearly useless for this purpose. Demand at least 12 months of trade history, and specifically look for history that spans a market downturn or a regime change — not just the most recent bull leg. Any platform that only shows the last 30 days is hiding exactly the information you need.
Step 4: Treat negative skew as a red flag. A return series that is dominated by small consistent gains interrupted by large occasional losses is structurally the pattern the PSR and DSR penalize most heavily. This is also exactly the pattern produced by martingale grid strategies, options-selling approaches, and high-leverage "trend until the drawdown kills you" methods — the strategies most common at the top of short-window leaderboards. If the account shows a very smooth equity curve interrupted by sudden large drawdowns, the Sharpe ratio is probably being inflated by a structural feature of the strategy, not by genuine skill.
Step 5: Compute a rough PSR as a sanity check. Even without the DSR deflation, the PSR gives you a more honest read than raw Sharpe. A PSR computation needs the annualized Sharpe, the number of monthly or daily return observations, the skewness, and the kurtosis. For a rough mental version: a Sharpe ratio computed over fewer than 24 months of monthly data should be treated as nearly meaningless for small or moderate Sharpe values. The standard error on a 12-month Sharpe estimate is large enough that a "1.0 Sharpe" could reflect anything from a -0.5 to a +2.5 true edge.
The deeper lesson: the DSR is not a tool for rejecting traders — it is a tool for calibrating how much evidence you actually have. A low DSR does not mean "don't ever copy this person." It means "you do not yet have sufficient statistical evidence to make this decision confidently." The correct response is to wait for more data, not to lower your bar.
The Honest Limits of This Data
Any analysis of this kind has limits, and being explicit about them is part of what makes the findings trustworthy rather than a pitch.
The 150 deep-judged candidates were not randomly sampled. The pre-filter that selected them from 23,495 was designed to surface the most promising candidates — those with the longest records, sufficient trade counts, and initial-screen Sharpe ratios above a threshold. The 150 represent the strongest candidates available, not a representative cross-section. This means the true rate of DSR-clearance across all 23,495 traders is almost certainly lower than the zero found in the deep-judged pool — the best candidates were evaluated, and none passed.
The choice of N in the DSR is a modelling decision. Using N = 23,495 (the full scanned population) is the conservative and defensible choice. A researcher could argue for using only the 150 deep-judged candidates, which would produce a more lenient test. Under that alternative, the results would still show zero candidates clearing 0.60, because the best DSR of 0.0692 is too far from the bar for the deflation adjustment to matter — but it is worth noting that the choice of N changes the scores numerically.
The scan covers three venues as of one date. Hyperliquid, OKX, and Polymarket are major venues, but they do not represent all of crypto trading. A genuinely skilled trader who operates primarily on Bybit, Binance, dYdX, or smaller decentralized venues would not appear in this scan. The result is specific to the population on these three platforms as of 2026-08-01.
The DSR bar of 0.60 is a choice, not a universal standard. The bronze clearance threshold of 60% probability of genuine skill is a conservative but reasonable starting point for a copy allocation decision. Some practitioners use 90% or 95% as their bar. Under any of those alternatives, the result here is the same — zero candidates clear — but the bar itself reflects a design choice about how much uncertainty you are willing to accept.
A zero result is consistent with both "no one here is skilled" and "the available data is insufficient to confirm skilled people who exist." The scan cannot distinguish between these two interpretations. It can only say that, given the current evidence, the statistical case for any individual candidate does not meet the threshold. That is a meaningful statement for a capital allocation decision, even if it is not a statement about the underlying truth of whether skill exists in the population.
FAQ
What is the difference between PSR and DSR, in plain terms?
PSR asks: "Given this track record, how confident are we that this trader is better than a coin-flipper?" DSR asks: "Given this track record and the fact that we searched 23,495 people to find them, how confident are we that they are better than the luckiest member of that crowd?" PSR treats the candidate as if they are the only person ever evaluated. DSR accounts for the search. When the search is large, the correction is enormous — as the jump from PSR = 1.0 to DSR = 0.0692 illustrates.
Does zero candidates passing mean no one on these leaderboards can trade?
No. It means the available public track records, evaluated against the size of the search, do not provide sufficient statistical evidence to make that determination. A skilled trader with a short public history will fail the DSR not because they lack skill but because the data is not yet long enough to clear a conservative threshold. The correct interpretation is "insufficient evidence," not "no skill exists."
Why does the best trader have a PSR of 1.0 and a DSR of 0.0692?
PSR = 1.0 means the observed Sharpe is so far above zero that it is essentially certain to be positive on a naive test. DSR = 0.0692 means that same Sharpe, once compared against the expected best-luck score from a 23,495-person field, falls well short of what pure chance would produce. The candidate looks certain to be skilled when evaluated alone. They look likely to be the luckiest of a large crowd when evaluated in context. The context is the correct frame.
What DSR score should I require before copying a trader?
The 0.60 bar used here is a reasonable starting point — it means 60% confidence of genuine skill after deflating for the search. A more conservative allocation decision might use 0.80 or 0.90. The key point is that any meaningful bar requires a DSR, not just a PSR, whenever you are choosing from a leaderboard rather than evaluating a referred candidate. A PSR of 0.95 in a field of 25,000 is weak evidence. A DSR of 0.70 in a field of 25,000 is meaningful evidence.
Can I run this test myself without a quant background?
Partially. The full computation requires return data at the trade or daily level, plus an estimate of field size. If a platform does not expose the return series, you are limited to approximations. The most accessible version of the DSR instinct: require very long track records (24 months minimum), avoid strategies with obvious negative skew, and mentally inflate the required Sharpe threshold based on the size of the board you are selecting from. The longer you wait for more data, the more the PSR and DSR converge — because time is the variable that resolves statistical ambiguity.
Disclaimer: This article is for informational purposes only and is not financial advice. Crypto trading involves significant risk of loss. Never trade with money you cannot afford to lose. Always do your own research (DYOR).
Affiliate Disclosure: This article contains affiliate links to Bybit, Binance, and Bitget. If you sign up through one of our links, we may earn a commission at no extra cost to you. Our editorial analysis — including the statistical methodology described here — is not influenced by affiliate relationships.