Accuracy audit · not a betting claim · est. 2026
Elon tweet-count model accuracy, audited.
An honest look at the record — over-confident bands and all.
This page audits Elon tweet-count model accuracy: point and interval accuracy for forecasts made 24–72 hours before a market closes, and nothing else. Across 211 forecasts on 90 markets, our nominal 90% prediction band actually covered the outcome about 67% of the time — so the model is over-confident, not calibrated. There's no profit, ROI, or tradeable edge here; the sample is small, versions are entangled with calendar time, and every trend is suggestive, not established.
Computed from our logged forecast record — past accuracy does not predict future accuracy
What the record actually shows
FIG. 01 — CALIBRATIONActual coverage
vs 90% nominal
Close date vs error
r, all months
Same r
June removed
Market size vs
error, r
Our 90% prediction band is over-confident: across 211 forecasts on 90 markets it actually covered the outcome about 67% of the time, not 90%. There's a weak hint that later forecasts are more accurate — the correlation between close date and error is r = −0.22 — but that entire signal comes from one month. Remove June and it collapses to r = −0.02: no trend. Bigger markets do forecast a little better (market size vs error r = −0.38), which is the one relationship that survives dropping June. None of this is a claim about the final-16-hour window, any other account, or profit.
June was anomalous: mean absolute error was 11.9% that month against a March–May baseline of 39.5%. We can't explain it as a model improvement rather than an easy month, so we don't count it as one. Read this as an audit of Elon Musk tweet-count markets forecasts, not a signal to trade them.
Does a 67% band mean the model can't trade? Not on its own — and that's why we measure calibration separately from anything else. A Polymarket tweet market is a set of discrete buckets, and trading one only needs the single most-likely bucket to be right often enough and priced too cheaply by the book. You don’t need a tight, high-confidence interval for that — a wide, over-confident band can still name the right bucket. But that is a different measurement — modal-bucket hit rate versus the market's own price — and we haven't validated it here. These markets price efficiently, so this page makes no profit or edge claim either way; it grades interval calibration on purpose, because it's the harder test to pass.
Month by month
FIG. 02 — MONTHLY BREAKDOWN| Month | Forecasts | Markets | Mean abs. error | Coverage | Note |
|---|---|---|---|---|---|
| Mar 2026 | 46 | 20 | 46.8% | 48% | Baseline |
| Apr 2026 | 47 | 20 | 36.1% | 53% | Baseline |
| May 2026 | 50 | 22 | 36% | 70% | Baseline |
| Jun 2026 | 56 | 23 | 11.9% | 96% | Anomaly — not representative |
| Jul 2026 | 12 | 5 | 41.5% | 50% | Partial month |
June (mean abs. error 11.9%, coverage 96%) is the single month that makes the record look like it's improving — and it's the month we least trust. July's partial data (12 forecasts, 41.5% error) already reads like the March–May baseline again. Drop June and the improvement disappears. That's the honest bottom line: we can't yet prove the model is getting better.
This page is an accuracy audit, not a betting claim. It measures point and interval accuracy for Elon Musk tweet-count forecasts made 24–72 hours before close, and nothing else — no profit, ROI, or tradeable edge; nothing about the final-16-hour window; nothing about any other person. The sample is small and model versions are entangled with calendar time, so every trend here is suggestive, not established. June was anomalous and is not representative. Numbers are computed from our logged forecast record; past accuracy does not predict future accuracy.