How this works, and how we stop ourselves cheating
Anyone can claim their forecasts are good. The hard part is building something that would catch you out if they weren't. That's most of what this page describes.
Every day, a cast of AI characters argues about each market on our list. Three independent readers then turn that argument into actual odds — "58% higher, 22% lower, 20% little changed" — plus a guess at the size of the move. Before any of it is published, the forecast is locked with a digital fingerprint and registered with independent outside services, so we can never quietly rewrite it later. When the market closes, code compares what we said to what happened.
- Small moves don't count as wins. A prediction only counts as right if the market moves more than its own normal daily wobble. Anything smaller is "little changed". That threshold is set from the market's recent volatility and locked in before we publish — so we can't move the goalposts after seeing the result.
- Being confidently wrong costs more than being unsure. We're scored on the odds we gave, not on a right/wrong tick. Saying "90% higher" and being wrong hurts far more than saying "40% higher" and being wrong. This makes honest uncertainty the profitable strategy.
- Weak inputs produce softer odds. If one of the three independent readers is unavailable or gives inconsistent answers, code pulls the final probabilities farther back toward the market's long-run frequencies. Once enough genuinely independent calls have resolved, a point-in-time calibration layer can soften or sharpen confidence — never change the chosen direction — using only results known before the new forecast.
- We always call a direction — and always show what it's worth. Every debate commits to higher or lower; we never sit one out. But a call is only meaningful if it beats what the market does anyway. If gold rises 40% of days, forecasting "40% up" is a lookup, not a prediction — so every call carries how far it departs from that history, and the scoreboard reports strong calls separately from near-coin-flips. A record built on the latter reads as exactly that.
- “Which way” and “how big” are two different questions. Every forecast produces three numbers — higher, lower, and barely moved — and only two of them are about direction. So the headline answers the direction question on its own terms: of the times this market actually moves, which way does it go? A day with odds of 31% higher, 31% barely moved and 38% lower is published as lower, 55% against 45%, with the 31% chance of a quiet day stated separately. Reading “lower, 38%” off the raw three would suggest we expect the opposite — we don't, and we shouldn't print something that implies it. Grading is unaffected: all three numbers are scored, untouched.
- Grey means quiet, not undecided. It is the probability that the closing move stays inside the volatility-based quiet range. Inside that range the outcome is neither “higher” nor “lower” for grading, so assigning it a direction would be false precision. The analysts still state which way they think a break would go.
- Odds and move size share one distribution. Code reweights the instrument's historical one-day or one-week returns so the higher, lower and quiet buckets carry the final judged probabilities. The displayed average move and 80% range are then read from that same distribution, preventing a separately generated size estimate from contradicting the odds.
- We must beat the dumb alternatives. Every forecast is also scored against simple benchmarks: always guess the historical average, always guess yesterday repeats, a plain coin flip, and a mechanical model with no AI in it. Beating a coin flip is not the bar — beating the best of those is.
- The characters never grade themselves. Every score, every verdict, every statistic is computed by code from real prices. The AI characters can explain and argue; they cannot decide whether they were right.
A required-sample calculation publishes here once computed.
Two things make that number bigger than it looks. Gold and silver called on the same morning usually move together, so two such calls are really closer to one piece of evidence — and we count them that way. And with 27 analysts, somebody will look brilliant by pure chance, the way someone always wins a raffle; we correct for that too.
We check every price against a second source — the US Federal Reserve, the European Central Bank, government energy data, or a second exchange. If the two disagree, the forecast goes on hold publicly rather than being graded on a number we don't trust. If it can't be resolved, the call is cancelled — and the cancellation stays on the record permanently, because a call that quietly disappears is a call that was hidden.
Accuracy figures on this site describe past forecasts only, are usually not statistically meaningful, and tell you nothing reliable about the future. The analysts are fictional AI characters, not people, and not qualified to advise anyone. This is a research and curiosity project. It is not investment advice or a recommendation to buy or sell anything.