Hi, Iβm Shane Gage. I represent Gage Systems, and I built Frontier Colosseum with a narrow goal that expanded into something much larger.
At the start, the question was simple: could AI help me make better decisions in prediction markets? That experiment evolved into a continuous benchmark β one designed around outcomes that cannot be gamed, only observed. Three times a day, a slate of real prediction-market questions is generated, hashed, and sealed. Thirteen frontier models, spanning multiple companies and countries, are each given that identical sealed slate. They answer independently and blind. No model has visibility into anotherβs response.
When reality resolves each question, every prediction is scored against the outcome. The record is permanent.
I initially expected to learn about reasoning performance β which models think more clearly, more consistently, more accurately. The results were less flattering, but more informative. Some models can generate profit. The stronger models can outperform the market on the record so far, but none has demonstrated reliable, sustained dominance over it.
The more significant finding emerged elsewhere, and I misread it at first. I called it convergence β models agreeing with one another. That is not surprising. Similar systems often produce similar outputs. What is surprising is this: independent systems, operating without coordination, are consistently wrong together on the same questions.
Independent judges do not fail in unison without a shared cause. These models cannot copy one another, so what they share is not necessarily information β it may be absence.
Each unanimous miss exposes a gap. The missing piece of information that would have corrected the outcome β or allowed stronger models to separate from weaker ones β is precisely what none of them had access to. The failure is not random. It appears structural.
This suggests a limitation in how these systems are trained. If multiple organizations, working independently, arrive at models with the same blind spots, then increasing scale alone may be unlikely to resolve the issue. The constraint may not be just capacity β it may be coverage.
Closing that gap likely requires a different approach: identifying and addressing these absences before deployment, rather than inheriting them from shared training assumptions. Frontier Colosseum is one mechanism for making those gaps visible, repeatedly and under controlled conditions.
Whether that leads to better systems remains an open question.
Thank you for following Frontier Colosseum β where AI attempts to predict the future.
β Shane Gage