OpenAI 2026 hackathon

100 Winters

A long-horizon benchmark that reveals whether an AI agent's decisions remain survivable after 100 winters of compounding consequences.

Solo project by aLexzzz430 Huang · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,274 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

The project described as "100 Winters" is a self-reported agent benchmarking tool designed to evaluate AI agents over long time horizons (400 decisions / 100 winters) in a simulated civilization world. It includes a causal simulation engine, a browser-based workbench for replay and audit, and an API or CLI adapter for external agents.

What changed

The project was submitted to the OpenAI 2026 hackathon. The author states that the original "Agent Survival Benchmark" predates Build Week and is frozen at a pre-Build-Week baseline. During Build Week, they added a 100-winter protocol, four world profiles with private shared shocks, browser workbench, product design, GPT-5.6 evidence, responsive UI, documentation, and publishing workflow.

Single most important open question

Is there any evidence of real-world adoption or usage by external agents beyond the demo and internal testing?

Back to contents

What The Product Actually Is

The description states that 100 Winters is a benchmarking tool for AI agents, designed to evaluate their long-term decision-making capabilities in a simulated civilization world. It runs agent policies through 400 seasonal decisions over 100 winters.

  • The system simulates a "causal civilization world" with variables like food, water, population, health, ecology, knowledge, institutions, technology, infrastructure, risk, shocks, collapse, and scoring.
  • It uses a WorldCore engine that is imported into a React/Vite workbench for visualization and replay.
  • The system supports external agents via a restricted API or Codex CLI adapter, which invokes GPT-5.6 Sol with a strict JSON action schema.
  • Observations are canonicalized and SHA-256 hashed; public replays are reconstructed from accepted events, not private state.
  • It includes tools for:
    • Replay
    • Comparison of score divergence
    • Inspection of terminal failures
    • Audit of exact public observations

Inference The product appears to be a simulation-based benchmarking platform, not a commercial SaaS offering. It is built for developers and researchers evaluating AI agents over time.

Back to contents

Positioning & Claim Evolution

The description states that the project was inspired by the idea that AI agents are often evaluated one task at a time, which can miss long-term consequences like ecological debt or institutional failure.

  • The product positions itself as a long-horizon benchmark where "the easy first decision is not the test—the accumulated consequence is."
  • It claims to reveal whether an AI agent's decisions remain survivable after 100 winters of compounding consequences.
  • The project evolved from an "Agent Survival Benchmark" to a more structured, productized version during Build Week.

Inference The positioning reflects a shift from a basic idea (agent evaluation) to a more formalized tool (benchmarking with simulation and replay). It is not clear if this was a pre-existing idea or a new direction taken during the hackathon.

Back to contents

Target Customer & ICP

The description states that external agents connect through a restricted API or Codex CLI adapter. The system is designed for developers and researchers who want to evaluate AI agents in long-term simulations.

  • It supports developers using tools like GPT-5.6 Sol.
  • It includes a browser-native workbench for replay, audit, and comparison.
  • It is built for those interested in causal evaluation of agent behavior over time, not for end-users or general consumers.

Inference The ICP appears to be researchers, developers, and AI engineers working on agent evaluation, particularly those focused on long-term consequences and causal modeling. No evidence of commercial customers or end-user adoption is provided.

Back to contents

Business Model & Pricing Evidence

The description does not state anything about a business model or pricing.

  • It mentions that the system supports external agents via API or CLI adapter.
  • It includes a live demo, public source, and setup instructions.
  • No mention of monetization, subscriptions, or fees is present.

Inference There is no evidence of a commercial business model. The project appears to be a research or hackathon product, not a revenue-generating service.

Back to contents

Technical & Delivery Signals

The description provides technical details:

  • Built with:
    • React/Vite
    • Node.js
    • GPT-5.6 Sol
    • Codex CLI adapter
    • JavaScript, Lucide icons, SHA-256 hashing
    • GitHub for version control
  • Uses a WorldCore engine simulating multiple variables.
  • Observations are canonicalized and hashed using SHA-256.
  • Supports public-only replays, fail-closed validation, and JSON schema enforcement.
  • Includes tools for:
    • Arena (visual projection)
    • Compare (score curves)
    • Audit (observation/hash/action boundary)
    • Protocol (reproducibility documentation)

Inference The technical stack is consistent with a developer tool or simulation platform, built in a modern web stack. It shows attention to reproducibility, security, and auditability.

Back to contents

Traction & Maturity Signals

The description states:

  • It was submitted to the OpenAI 2026 hackathon.
  • It includes a live demo, public source, and setup instructions.
  • It has twenty passing protocol, API, export, security, and Codex tests.
  • A live GPT-5.6 Sol run is included with four accepted structured decisions and archived observation hashes.
  • The project was built during Build Week, with a changelog showing before/after.

Inference There is no evidence of customer traction or adoption beyond the demo and internal testing. It is a prototype or hackathon product, not a mature commercial offering.

Back to contents

Competitive Context

The description does not mention any competitors or direct market context.

  • The project is positioned as a benchmarking tool for AI agents.
  • It is inspired by the need to evaluate long-term consequences of agent decisions.
  • No mention of existing tools, platforms, or benchmarks in this space is provided.

Inference There is no evidence of competitive analysis or awareness of similar products. The project appears to be independent, possibly novel in its approach, but unproven in the market.

Back to contents

Key Risks & Red Flags

  • No commercial traction or adoption: The product is described as a hackathon submission with no evidence of real-world usage.
  • Unverified claims: The description states that it evaluates agents over 100 winters and includes GPT-5.6 Sol, but there is no independent verification of these claims.
  • Limited scope: It appears to be a research or prototype tool, not a scalable commercial product.
  • No pricing or monetization model: No indication of how the project would generate revenue.
  • Self-reported maturity: The project is described as built during Build Week, suggesting it is in an early stage.

Inference The risk is high that this is a non-commercial prototype, not a viable product for investment or partnership.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended use case for 100 Winters beyond the demo and internal testing?
  2. Are there any external agents currently using the API or CLI adapter?
  3. How does the project plan to scale beyond the current prototype?
  4. Is there a roadmap for monetization or commercialization?
  5. What are the limitations of the current simulation engine in terms of realism or scalability?
  6. How is the system validated for fairness and reproducibility?

Back to contents

Investment/Partnership Verdict

The description states that 100 Winters was submitted to the OpenAI 2026 hackathon and built during Build Week. It includes a live demo, public source, and internal tests but no evidence of commercial traction or adoption.

  • The project is described as a benchmarking tool for AI agents, not a commercial product.
  • There is no evidence of revenue, customers, or market demand.
  • It appears to be a research or prototype effort, not a scalable business.

Inference Based on the self-reported description only, there is no basis for investment or partnership. The project lacks commercial viability and traction. It may have potential as a research tool but is not ready for commercial use or funding.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.