OpenAI 2026 hackathon

ForecastProof

An evidence-first agent that validates forecasting papers, challenges them statistically, and runs one governed next experiment.

Solo project by SUN XING · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #4,190 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

ForecastProof is a self-reported research-support system designed to validate financial forecasting papers using deterministic checks and LLM-assisted synthesis. The author states it runs an end-to-end workflow that includes analyzing, verifying, challenging, and generating audited decisions on forecasting models.

What changed

The project was built for the OpenAI 2026 hackathon. It is described as a Python/Streamlit application using GPT-5.6 and various ML libraries to create an evidence-first agent for validating forecasting papers.

Single most important open question

Is there any evidence of real-world usage, adoption or traction beyond this single hackathon submission?

Back to contents

What The Product Actually Is

The description states that ForecastProof is:

  • A Python and Streamlit application
  • Built with codex, gpt-5.6, numpy, openai, pandas, python, pytorch, responses-api, scikit-learn, streamlit, structured-outputs
  • An evidence-first agent that validates forecasting papers
  • Designed to run one governed next experiment after validation

The author claims it takes an evidence-backed forecasting paper through:

  1. Analyze the paper
  2. Verify on a common task
  3. Challenge the result
  4. Generate an audited decision memo
  5. Run one governed iteration

It is described as a "research-support system" not investment advice, and never authorizes trading.

Confidence: Low — This is entirely self-reported with no evidence of actual product usage or performance beyond the demo.

Back to contents

Positioning & Claim Evolution

The description states:

  • ForecastProof turns questions about financial forecasting papers into an executable, auditable workflow
  • It replaces asking an AI to declare whether a paper "works" with deterministic evidence gates
  • GPT-5.6 handles language reasoning while deterministic code owns decision authority
  • The system refuses to turn uncertainty into a confident-looking recommendation

The author claims the product:

  • Makes "evidence-bound proposals may change only allow-listed parameters"
  • Uses "deterministic code—not the model—owns every pass/fail gate, promotion decision, and deployment hold"
  • Is designed to be "more trustworthy when its authority is deliberately narrow"

Confidence: Low — These are claims about positioning and intent, not verified outcomes or traction.

Back to contents

Target Customer & ICP

The description states:

  • ForecastProof is a research-support system for financial forecasting papers
  • It targets researchers working with forecasting models
  • The demo contains three S&P-related papers and three genuinely executed model families
  • It runs on frozen SPY next-day-direction dataset with identical 12-lag features, 41 purged walk-forward folds, etc.

No explicit customer segment or ICP is defined beyond "researchers" or "papers." No mention of institutional users, finance teams, or commercial customers.

Confidence: Very low — No evidence of target customer identification or commercial use cases beyond the hackathon demo.

Back to contents

Business Model & Pricing Evidence

The description states:

  • The default experience is a verified replay, so judges can run the complete golden path without an API key or model cost
  • Live mode is opt-in, capped, uses low reasoning, sets store=false, and records usage telemetry without storing the key
  • It's described as a "research-support system" not investment advice

There is no mention of pricing, monetization, revenue streams, or commercial licensing.

Confidence: Very low — No evidence of any business model or pricing structure.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with Python, Streamlit, pandas, NumPy, scikit-learn, PyTorch
  • Uses GPT-5.6 Luna through the OpenAI Responses API
  • Implements deterministic checks and structured outputs
  • Has 131 automated tests including real training smoke tests for all three model families
  • Includes read-only tools like get_evidence_brief and get_verification_result

It is described as:

  • A frozen research-suite artifact keeping all paper cases on one comparable protocol
  • Reproduction tiers distinguishing strict reproduction from paper-inspired adaptation
  • Using a "bounded iteration loop" that changes parameters, trains, evaluates reserved folds, and retains the parent when guardrails fail

Confidence: Medium — Some technical details are provided, but no evidence of production deployment or scalability.

Back to contents

Traction & Maturity Signals

The description states:

  • Demo contains three S&P-related papers and three genuinely executed model families
  • All cases share exactly 656 aligned out-of-fold targets and pass four comparison-integrity gates
  • The UI separates “adaptation succeeded” from “forecast skill is established”
  • 131 automated tests pass, including real training smoke tests for all three model families
  • A live GPT-5.6 Luna acceptance run called both tools and passed all eight memo audits at 100/100 while preserving deployment HOLD

However:

  • No mention of actual users or customers
  • No revenue data or adoption metrics
  • No evidence of real-world usage beyond the hackathon demo
  • No mention of any product in production or commercial use

Confidence: Very low — The only "traction" is from a hackathon demo.

Back to contents

Competitive Context

The description does not provide:

  • Any information about competitors
  • Market positioning relative to existing tools
  • Mention of similar products or platforms

No evidence of competitive landscape analysis or differentiation strategy.

Confidence: Very low — No competitive context provided.

Back to contents

Key Risks & Red Flags

Key claims in the description raise concerns:

  1. Unproven commercial viability: The system is described as a research tool, not a product for sale
  2. No revenue or customer evidence: No signs of monetization or real-world adoption
  3. Limited scope: Only three papers and three model families tested in demo
  4. Hackathon origin: Built for a single hackathon event with no indication of further development
  5. Self-reported maturity: No evidence of production-grade reliability or scalability

Confidence: Medium to High — These are inferred risks from the lack of evidence.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual intended commercial use case for ForecastProof?
  2. Has there been any real-world testing beyond this hackathon demo?
  3. Are there plans to monetize or scale this beyond a research prototype?
  4. How does the system handle edge cases or failures in the verification process?
  5. What are the limitations of the current implementation that would prevent production use?
  6. Is there any plan for integrating with existing financial research platforms or workflows?

Back to contents

Investment/Partnership Verdict

The description states that ForecastProof is a "research-support system, not investment advice, and never authorizes trading." It was built for a hackathon.

There is no evidence of:

  • Revenue
  • Customers
  • Product-market fit
  • Commercial traction
  • Any business model beyond the demo

Confidence: Very low — This appears to be an experimental prototype with no demonstrated commercial potential or market readiness. The author states it is not investment advice, and there is no indication that it has moved beyond a proof-of-concept stage.

Verdict Not ready for investment or partnership consideration based on the provided evidence.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.