OpenAI 2026 hackathon

CriteriaForge

Turn human product intent into a ratified, executable Constitution that Codex can apply—but never silently rewrite.

Solo project by 勇樹 浦田 · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,576 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

CriteriaForge is a product that enables non-technical product owners to create a human-ratified, executable "Product Constitution" from unstructured source material (e.g., documents, spreadsheets, images, video metadata). It uses GPT-5.6 to propose and apply this constitution to Codex-built work, but prevents AI from silently rewriting the intent.

What changed

The project is a self-contained tool built for the OpenAI 2026 hackathon. It includes both a public demo and a local macOS version that binds only to localhost. It uses GPT-5.6 in a structured way to propose testable criteria, identify conflicts, and apply deterministic evaluation rules.

Single most important open question

Does CriteriaForge actually solve the stated problem of preventing AI from silently rewriting human intent, or is it a conceptual framework that has not yet been proven at scale?

Back to contents

What The Product Actually Is

The description states that CriteriaForge:

  • Helps non-technical product owners turn messy source material into a "human-ratified, executable Product Constitution"
  • Uses GPT-5.6 to propose an eight-section Product Constitution
  • Keeps original documents, spreadsheets, images, video metadata, and Git evidence on the user’s Mac
  • Evaluates Codex-built work against this constitution using deterministic safeguards
  • Allows for bounded Codex repair in a disposable Git worktree

It is described as a local macOS app with a web interface built using Next.js, React, TypeScript, Tailwind CSS, and shadcn/ui. It also includes a public Vercel build that replays three recorded GPT-5.6-sol runs on fictional data.

Inference The product appears to be a hybrid tool combining local storage, AI-driven structuring of human intent, and deterministic evaluation of Codex outputs. It is not a SaaS offering but rather a self-contained desktop application with a public demo.

Back to contents

Positioning & Claim Evolution

The author states:

  • The goal is to address the gap between Codex’s speed and the need for human authority over product decisions
  • It introduces a “different authority model” where humans ratify the constitution, and GPT-5.6 may propose and apply it at scale but cannot silently redefine it
  • The tool aims to prevent AI from turning explicit promises or exclusions into indistinguishable assumptions

Inference CriteriaForge positions itself as a solution for product teams who want to use AI tools like Codex but retain control over the underlying intent and decision-making process. It is not about replacing human judgment, but about creating a structured way to ensure that AI outputs align with ratified decisions.

Back to contents

Target Customer & ICP

The description states:

  • The primary user is a "non-technical product owner"
  • It targets those who want to use Codex but need to maintain control over the output
  • The tool supports both local and public versions, suggesting it may appeal to developers or teams working in constrained environments

Inference The ICP likely includes product owners or project managers who are not technical but must oversee AI-assisted development workflows. It is not a general-purpose AI tool but one tailored for those seeking structured control over AI outputs.

Back to contents

Business Model & Pricing Evidence

The description states:

  • There is no mention of pricing, subscriptions, or monetization
  • The public demo is available without sign-in
  • A local macOS version can be run with Node.js, Git, and an authenticated Codex CLI
  • No revenue model or customer acquisition strategy is described

Not evidenced No indication of a business model, pricing structure, or monetization approach.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with Next.js, React, TypeScript, Tailwind CSS, shadcn/ui, SQLite, and Vercel
  • Uses GPT-5.6 via the official Codex CLI
  • Local edition binds only to 127.0.0.1 port and uses secure local storage
  • AI contracts use JSON Schema 2020-12 with undeclared properties rejected
  • AI citations are checked locally against approved source IDs, segment IDs, typed locators, and SHA-256
  • The public Vercel build omits local API routes and checks for forbidden markers

Inference The tool is built with a focus on security, privacy, and deterministic behavior. It uses local storage and avoids exposing sensitive data in the public version.

Back to contents

Traction & Maturity Signals

The description states:

  • The project was submitted to the OpenAI 2026 hackathon
  • It includes a public demo replay using fictional data
  • A local macOS edition is available for testing
  • The current automated suite contains 65 unit/integration checks plus four browser executions
  • It has zero vulnerabilities in dependency audit and no console errors or axe violations

Not evidenced No real-world usage, customer base, or adoption metrics are provided. The project is described as a hackathon submission with no indication of traction beyond the demo.

Back to contents

Competitive Context

The description does not mention any competitors or direct comparisons to existing tools in the market.

Not evidenced No competitive analysis or positioning relative to other AI-assisted development or product management tools.

Back to contents

Key Risks & Red Flags

  • The tool is described as a hackathon submission with no commercial traction
  • It is unclear whether the solution scales beyond the demo or local use case
  • The public demo uses fictional data, not live GPT-5.6 endpoints
  • There is no evidence of real-world adoption or feedback from users
  • The project is self-contained and does not appear to be a scalable SaaS offering

Inference The tool may be a proof-of-concept rather than a mature product. It lacks commercial viability indicators, and its utility in real-world settings remains unproven.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual use case for this tool beyond the hackathon demo?
  2. How does it handle edge cases where human intent is ambiguous or contradictory?
  3. Is there any evidence of user feedback or testing outside of the demo environment?
  4. What are the plans for scaling beyond a local macOS app and public demo?
  5. Are there any real-world integrations with Codex or other AI tools that have been validated?

Back to contents

Investment/Partnership Verdict

Not evidenced No financials, revenue, or customer data are available to assess commercial viability.

Inference Based on the self-reported description, CriteriaForge is a hackathon project with limited evidence of traction or scalability. It may be an interesting concept for product teams seeking control over AI outputs, but it lacks the maturity and commercial indicators needed for investment or partnership consideration at this stage.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.