A P&L screenshot is a story. The wallet is the receipt.

Most of the time the story wins, because the receipt takes work to read. Crypto Twitter runs on screenshots — a vertical chart, a fat green arrow, "+387% in 30 days," ten thousand likes, and three reshares before lunch. By Monday the trade is folklore, and nobody asked a single question. The reason this works is also the reason it should bother you: our brains anchor on the most concrete thing in front of them, which is exactly what a JPEG looks like. It's not immutable. It's thirty seconds and a charting tool.

This piece is the buyer-side counterweight. Five checks, in order of how much they expose. Apply them to any track record — bot, agent, fund manager, Discord guru — and you'll cut the marketing fog down to what actually matters. I'll cite BullSpot once as a working example because the company publishes a verifiable on-chain record, but the framework isn't vendor-specific. If a system can't survive this audit, you don't need it.

Why Smart Money Still Gets Fooled

Three biases do most of the damage.

Confirmation bias is the obvious one. If you want to believe a vendor is good, you'll skim the trade list for wins and stop reading. The brain treats disconfirming evidence as friction. So the audit isn't about willpower — it's about structuring the questions so the brain can't skip them.

Survivorship bias is the quieter killer. You're only seeing the products that survived long enough to be marketed. The 38 trading bots launched last quarter and killed by their first drawdown? Gone. The two funds that blew up in February? Quietly renamed. Survivorship doesn't just apply to strategies — it applies to vendors, traders, and even the screenshots themselves, because losers don't get screenshotted.

The narrative fallacy finishes the job. We want a clean story: "the agent saw the funding flip and went long." Stories are memorable. Trade logs are not. So the seller always leads with the story. The buyer's job is to refuse it.

The fix is mechanical. You run the same five checks every time, in the same order, the way an accountant runs a balance sheet. Here they are.

Check 1 — Sample Size and Time Under Market

A trade log with 18 entries is a sample, not a record. Statistically, 18 trades from a discretionary strategy has roughly the same predictive power as 18 coin flips if outcomes are even remotely noisy — which they are in crypto. The signal-to-noise ratio on Hyperliquid perpetuals over a month is brutally low. You need to see hundreds of closes across multiple regimes to even start estimating expectancy.

What you're actually looking for:

  • Closed trade count. Not "signals generated." Not "trades placed and currently open." Closed and resolved.
  • Span across at least one full regime cycle. A vendor whose entire track record is from a one-way bull market hasn't traded. They've been long. The current tape — BTC chopping in the $83,500–$84,200 range while the 1D EMA ribbon stays bullish and 4H SuperTrend stays bearish — is exactly the kind of grinding, conflicting-signal environment that punishes one-direction strategies. If the record doesn't include at least one tape like this, you don't have a record. You have a screenshot.
  • Outliers called out. If one trade made 40% of the total P&L, that's not a track record. That's a lucky flip. The system should still be interesting — but only after you mentally delete that trade and see what's left.

A workable floor: 100+ closed trades across at least three distinct market regimes (trending up, trending down, range-bound) before the word "edge" earns its place in your notes. Anything less is a marketing artifact.

Check 2 — The Full Drawdown Story

This is where the screenshot economy loses most of its camouflage. Vendors will show you the equity curve. They will not show you the underwater curve. The difference matters more than anything else on the page.

Two track records can end at the same number and be radically different businesses. One path: −4%, −3%, +1%, +2%, +6%, +9%. Another: +18%, −22%, +14%, −19%, +11%, +2%. Same destination, completely different experience. The second one is a yo-yo that any sane allocator would have killed after the first drawdown. The equity curve doesn't tell you which you're buying.

What to demand:

  • Maximum drawdown, in percent and in duration. "−32% over 47 days" is information. "We had some drawdowns" is not.
  • Recovery time. How long from peak to new equity high after the worst stretch? A strategy that takes 9 months to recover from a 20% drawdown is not the same business as one that recovers in 6 weeks, even if the final number matches.
  • The shape of the distribution tells you whether the vendor got unlucky once or whether the strategy bleeds regularly.

Flat equity curves are themselves suspicious. Real strategies have rough patches. A line that goes up and to the right without interruption is usually one of three things: a survivorship-filtered window, a leveraged bet that hasn't met its reckoning yet, or a fabrication.

Check 3 — Survivorship and Selection

Even an honest vendor can hand you a misleading record by accident. The most common way is selection: showing only the trades that survived in their system, after they've already filtered or restructured.

The structural tells:

  • Closed positions vs. open positions in the log. If the public log shows 100% closed winners and 0% losers, somebody is editing. A real book has both, in roughly the proportions the strategy would generate.
  • The "we stopped trading that one" tell. When a vendor explains away a bad stretch by saying "we paused the system during that period," ask whether pausing was a discretionary choice made after knowing the period was going to be bad. If yes, you're looking at a backtest-shaped record.
  • Strategy-version drift. "The current version is v4.2 — earlier versions had a rough Q1." Translation: the early version lost money and the marketing uses the new one. Fine, but you want the v4.2 record specifically, including its early losses, not its mature self.

The deeper form of survivorship is vendor survivorship. You won't see the graveyard. You won't see the bot that blew up in March, the AI agent that got shut down after a single −60% week, the fund that quietly returned capital and pivoted to NFTs. They are not competing for your dollar right now, so they are not on your screen. Assume any category of trading product you encounter is the survivor of an iceberg. Price the iceberg in.

Check 4 — On-Chain Verifiability

A screenshot can be Photoshopped in 30 seconds. A wallet cannot. On Hyperliquid, every fill, every funding payment, every liquidation is signed and visible on-chain. This is the single most important property of the venue for buyer protection, and almost nobody uses it.

What counts as verifiable:

  • A wallet address you can audit yourself. You should be able to paste it into the explorer and see the trade history, the deposits, the withdrawals, the funding receipts. Not a curated summary. The actual chain.
  • Read-only access where possible. Vendors that let you connect a read-only API key or view permissions on the wallet itself are taking custody risk off the table. They cannot move your funds, and you can prove that.
  • Closed positions visible, not just current state. A vendor showing only the open position is hiding the closed trades — the actual edge or lack of one.

What does not count, no matter how it's framed:

  • A dashboard screenshot with no underlying wallet.
  • A "live" page on the vendor's site that you can't independently verify against the chain.
  • A PNL CSV you have to take on faith.
  • A third-party "verifier" that itself can't be verified.

BullSpot is a useful example here: the company publishes a verifiable on-chain record. Whether or not you end up a customer, that's the bar. If a system can't show you the chain, the chain is telling you something.

Check 5 — Methodology on the Record

P&L is an outcome. Methodology is the process. You can have a great outcome from a bad process (luck) or a bad outcome from a great process (variance). You cannot tell the difference from the equity curve alone. You need the reasoning.

What to look for:

  • Reasoning logs per trade. Not just entry and exit — the why. What was the setup, what was the opposing case, what was the invalidation. The reasoning can be terse. It cannot be absent.
  • Vetoes and skipped trades. A system that took every signal it generated is a rule bot, not an agent. The interesting question is what it didn't do. Real agents pass on setups. Real agents override their own rules when the tape disagrees.
  • Consistency under stress. Compare the reasoning quality in winning trades versus losers. If the winners are explained in detail and the losers are summarized in a sentence, the vendor is curating again.

A useful analogy: two surgeons can have identical success rates over a year. One keeps meticulous notes, changes gloves between patients, and double-checks imaging. The other has good hands and vibes. After a thousand operations, you might not be able to tell their outcomes apart. But the day something goes sideways, you'd rather be on the table of the one who documented everything.

The Market Context Test

One last filter, applied to the vendor's record and your own priors: does the track record include a tape like the one you're about to trade?

Right now the broader structure is constructive — BullSpot's market brief on Saturday framed the $84K reclaim as initiative buying after a clean $83,508 bear-trap sweep, with two bullish volume displacements (2.2x and 2.3x) doing the actual work. But the internal signals are conflicted: 1D EMA ribbon bullish with RSI at 63.7, 4H/12H SuperTrend bearish, 1H ribbon bearish. Funding is flat. Positioning is balanced. This is a grinding, mixed-signal tape — exactly the kind that exposes strategies built only for clean trends.

If a vendor's record is 80% one-direction in trending markets and 20% in chop, and you size up during chop, you've already lost. The audit isn't only "did this work somewhere." It's "did it work somewhere that looks like this."

The Checklist

Five questions, in order. If any answer is missing, vague, or footnoted with "trust me," walk.

  1. How many closed trades, across how many regimes? Floor: 100 closed, three regimes. Lower than that and you're reading a sample, not a record.
  2. What was the worst drawdown, in percent and in time? Plus recovery duration, plus the shape of the drawdown distribution.
  3. Are losers visible in the public log? Same granularity as winners, same time stamps, no "we paused during that period" hand-waving.
  4. Can I see the wallet on-chain? Independently, with a read-only view, including closed positions — not a dashboard screenshot.
  5. Is the reasoning on the record? Per-trade, including the trades it didn't take, with comparable detail on winners and losers.

Anything that survives all five is worth a deeper conversation. Anything that fails two or more is marketing. The middle — passing four of five — is where most real products live. Treat those failures as the conversation you'd need to have before clicking allocate, not as automatic disqualifications.

Track records are stories until they aren't. The chain, the closed-trade log, and the reasoning trace are the receipts. Build the habit of asking for all three, every time, and the curated-screenshot economy loses most of its power over you.


Source context: BullSpot report from 2026-09-26T04:18:31.423Z (Fresh report: generated this cycle).