This InsightMeter guide teaches a calm way to read any public forecast track record—including InsightMeter-style scorecards and third-party leaderboards—without turning a percentage into a personality cult. Your reader job is to demand a fixed horizon, an explicit benchmark, a countable sample, and visible failure modes before you grant a scoreboard any research weight. Filing literacy and research-workflow articles elsewhere on this site already cover EDGAR mechanics and thesis falsification. Here the object is the scorecard itself: how to evaluate it as education about measurement, not as a sales pitch.
Open the scorecard. Do not ask “is this person a genius?” first. Ask: “Can I restate the scoring rule in one sentence a skeptic could audit?” If you cannot, you are looking at branding with numbers attached. Branding can be honest marketing; it is still not an auditable track record until the rule is locked.
Write four fields before you quote the hit rate: horizon (when does a call resolve?), benchmark (against what is success measured?), hit definition (what counts as hit vs miss?), and sample (which calls are included, and which are excluded by rule rather than by embarrassment?). Empty fields mean the percentage is not ready for transfer into your research notes.
Soft product mention once: InsightMeter publishes explicit scoring rules on methodology.php so readers can see how horizons and benchmarks are supposed to be stated. Use that page as a template for reading any scoreboard—ours or elsewhere—not as a claim of guaranteed returns.
Horizon. “Up next week,” “outperform over 20 trading days,” and “beat estimates next earnings” are different games. A scorecard that mixes horizons inside one percentage without labeling strata is averaging unlike objects. If the marketed story implies multi-month skill while the engine scores five-day direction, the transfer is illegitimate.
Benchmark. Absolute-up hits (“stock finished higher”) are not the same as excess-return hits (“stock beat sector ETF / index / stated hurdle”). A forecaster can look brilliant on absolute hits in a roaring bull tape and ordinary on excess returns—or the reverse in a bear tape. If the pitch says “alpha” while the scoreboard uses absolute direction, the words and the math disagree.
Sample. Sample size belongs beside every rate. 8/10 and 80/100 are not interchangeable stories. Tiny samples swing wildly; large samples can still be biased. Ask whether the sample is all calls under the locked rule in a dated window, or a curated subset. Curation after outcomes is survivorship wearing a lab coat.
Survivorship and silent edits. Dropping misses because they were “not real setups,” deleting accounts that failed, or rescoring with a friendlier horizon after browsing results inflates rates. Educational scorecards show misses, restatements, and rule changes in the open. If you cannot find failure modes, assume they may be offstage.
Minutes 0–3: copy the stated rule (or note that it is missing). Minutes 3–7: identify horizon, benchmark, hit definition, and sample window from the page—not from a social caption. Minutes 7–12: spot-check two hits and two misses (or the nearest available outcomes) against the rule; see whether you would score them the same way. Minutes 12–15: write a one-line verdict: auditable under rule R / incomplete / mismatched to the claim being marketed.
Do not spend the fifteen minutes hunting for the highest percentage on the site. Marketers put the highest percentage where eyes go. Your job is definition lock and failure visibility.
If the scorecard shows backtested results, label them backtests in your note. Backtests can be useful educational simulations; they are not live proof. Pair this habit with the backtest-vs-live scoring guide so “historical study” never silently becomes “live verified edge.”
If the scorecard is live-scored, still ask about revisions: when prices are corrected, when calls are amended, when timestamps disagree across platforms. Live scoring with quiet restatements can recreate survivorship in slow motion.
Fiction A — Clean but modest. Scorecard states: 20-session excess return vs fictional sector ETF SECT; hit if excess > 0; sample = all River calls on MAPL 2022–2024 (n=40); hit rate 55%; misses listed. Verdict: auditable description. Educational weight: limited, sample moderate, not a tip.
Fiction B — Cherry-picked window. Same engine, but banner shows “92% accurate in Q2 2024” with n=12 and no full-year panel. Verdict: incomplete for general skill claims; window may be honest as a slice if labeled—overclaim if used as the only number in a pitch.
Fiction C — Benchmark swap. Page scores absolute 5-day ups at 70%; sales copy says “market-beating alpha.” Verdict: definition mismatch falsifies the alpha wording until excess-return results are shown under a locked hurdle.
Fiction D — Survivorship. Leaderboard removes retired forecasters who finished below 45%. Remaining cohort averages 61%. Verdict: cohort average is conditional on survival; not comparable to an all-forecaster pool without the dropouts restored or explicitly modeled.
Overclaim: “Proven 80% accuracy—follow every long.” Quieter: “Under rule R, 80 of 100 resolved calls in window W were hits; past rates are not future promises; here are the misses.”
Overclaim: “Backtest confirms alpha.” Quieter: “Simulation under assumptions A/B/C produced positive excess return in-sample; out-of-sample/live results are separate; no guaranteed edge.”
Overclaim: “Everyone missing this scoreboard is leaving free money on the table.” Quieter: “A transparent scoreboard helps you compare definitions; it does not create free money.”
Ban in research notes: guaranteed returns, proven alpha, “can’t lose,” and SEO-style superlatives that treat a percentage as destiny. Filings and forecasts support modest verbs: consistent with, inconsistent with, insufficient sample, definition mismatch.
Use this template on any public forecast track record before you cite it.
Checklists are educational process tools. They do not create profitable trades.
Most scorecard errors are social—people want a hero percentage.
Correcting them improves research hygiene even if you never publish a forecast.
A fully auditable scorecard can still describe luck, regime fit, or a game that will not continue. Measurement literacy reduces overclaiming; it does not manufacture foresight.
Different honest rules produce different rankings. Two careful scorekeepers can disagree on ties, corporate actions, and timestamp conventions. Document the rule you used; do not pretend rankings are metaphysics.
Public scorecards omit private information, position sizing, costs, and constraints. A directional hit can still be a poor economic outcome after friction. Hit rates are not portfolio returns unless the methodology explicitly scores returns that way.
Nothing in this guide is a recommendation to buy or sell any security, to follow any forecaster, or to expect future performance from past percentages.
Give juniors Fiction B and Fiction C and ask for a checklist verdict. If they return “looks strong—92%” without naming horizon and sample, the assignment fails.
Require every cited rate to travel with rule R, window W, and n. Adjectives without those fields are incomplete.
Pair with personal forecast journals so learners feel the pain of scoring their own misses. Empathy for misses beats abstract lectures about survivorship.
Prefer: “Under 20-day SECT-excess rule R, River’s MAPL sample n=40 hit rate was 55% in 2022–2024; misses included; not a forecast of future results.” Avoid: “River is 55% genius smart money.”
When you share a screenshot, crop in the rule text—or admit the rule was not on-screen. Caption-only percentages are how overclaiming spreads.
If someone asks “so should I buy?”, answer with process: the scorecard does not answer that question. Education ends at measurement; advice begins somewhere this guide refuses to go.
Use forecast track records and scoring/backtests guides for deeper metric design. Use backtest-vs-live to keep simulation labels honest. Use a personal journal to practice the same rules on your own calls.
Falsification habits from research-workflow guides apply: a marketed percentage with mismatched horizon is already falsified as support for the marketed claim.
Methodology and glossary keep shared vocabulary—hit rate, benchmark, horizon, survivorship—so arguments stay about evidence.
Evaluating Forecast Track Records Fairly · Backtest vs live forecast scoring pitfalls · Forecast scoring & backtests · Building a personal forecast journal
Lock horizon, benchmark, hit definition, and sample before you trust a percentage; look for misses and survivorship; label backtests vs live; and refuse guaranteed-return language. A scorecard you can audit is education—a scorecard you can only admire is advertising. Continue with related guides, methodology, and glossary. Nothing here recommends buying or selling any security.