OpenAI 2026 hackathon

PromiseProof

GPT-5.6 investigates, Codex repairs, a human approves and an unchanged deterministic verifier decides PASS.

Solo project by Alexandre Paiva · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #1,726 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

PromiseProof is a self-reported developer tool built by one person (Alexandre Paiva) during a 1-week hackathon. It is described as a system that investigates, repairs and verifies user-facing promises in software — specifically, ensuring that when a feature like personalization is turned off, no user-identifying data reaches recommendations. The system uses AI for diagnosis and repair but excludes AI from the final verification step.

What changed

The project was built entirely during a single week (Build Week) as a solo effort. It includes a reference application, deterministic evidence collection, an AI-assisted investigation workflow using GPT-5.6, a Codex-based repair process, human patch approval, and a deterministic verifier that excludes AI from declaring PASS.

The single most important open question

Is there any evidence of real-world adoption or traction beyond the author’s own demonstration? The description states no revenue, customers, or usage data exist outside of the self-contained demo environment.

Back to contents

What The Product Actually Is

The description states:

  • PromiseProof is a system that takes one concrete user-facing promise and turns it into a deterministic check.
  • It uses Playwright journeys and network captures to observe broken promises.
  • GPT-5.6 investigates, Codex repairs, and a human approves the patch.
  • A deterministic verifier (Playwright + evaluator) decides PASS or FAIL without AI involvement.
  • The system supports three developer surfaces: browser, CLI, and GitHub Action.

Inference The product is a proof-of-concept tool for integrity verification in software systems, focused on ensuring that user-facing features behave correctly under specific conditions.

Back to contents

Positioning & Claim Evolution

The description states:

  • The project was submitted to the OpenAI 2026 hackathon.
  • It aims to solve two problems: hidden system defects and AI's inability to be trusted as a judge of its own repairs.
  • The author positions it as a way to let AI do diagnosis and repair while structurally preventing AI from declaring success.

Inference The positioning is that PromiseProof is a tool for developers to validate software behavior with integrity, using AI in a controlled way without granting it authority over final outcomes.

Back to contents

Target Customer & ICP

The description states:

  • The system supports three developer surfaces: browser, CLI, and GitHub Action.
  • It targets developers working on distributed systems where user-facing promises can be broken in subtle ways.
  • The tool is built for use in CI/CD pipelines via a GitHub Action.

Inference The primary customer is likely software engineers or DevOps teams working with complex, distributed applications where integrity of user-facing features is critical.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

Explanation

There is no mention of pricing, monetization, or business model in the description. The project is described as a hackathon submission and not a commercial product.

Back to contents

Technical & Delivery Signals

The description states:

  • Built using Node.js, TypeScript, Express.js, Playwright, GitHub Actions, Cloudflare Workers, OpenAI API, Codex SDK, Zod, esbuild, Vite.
  • Uses deterministic testing, git worktrees, and a pinned evaluator fingerprint.
  • Supports Windows, Ubuntu, macOS via CLI and GitHub Action.
  • The verifier path excludes all model calls and uses unchanged Playwright + evaluator.

Inference The tool is built with modern developer tooling and emphasizes determinism and security in verification.

Back to contents

Traction & Maturity Signals

Not evidenced.

Explanation

There is no evidence of revenue, customers, usage metrics, or product adoption beyond the author’s own demonstration. The project is described as a solo-built hackathon submission.

Back to contents

Competitive Context

Not evidenced.

Explanation

No mention of competitors or market context in the description. The tool is presented as a novel approach but not positioned against existing tools.

Back to contents

Key Risks & Red Flags

  • Solo developer: The entire project was built by one person, which raises questions about scalability and long-term maintenance.
  • Limited scope: Only one synthetic reference application and one contract family are supported.
  • No real-world validation: The system is described as a demo with seeded failures; no evidence of use in production or third-party validation.
  • No external evidence collection attestation: While binding proves reports match evidence, it does not prove the evidence was collected honestly.

Back to contents

Diligence Questions To Ask The Founders

  1. What are the actual use cases for this tool beyond the demo?
  2. Has there been any independent testing or validation of the system?
  3. How would you scale this to support more than one contract family?
  4. Are there plans to integrate with real-world applications or platforms?
  5. What is the long-term roadmap for developer adoption and product maturity?

Back to contents

Investment/Partnership Verdict

Not evidenced.

Explanation

There is no evidence of funding, valuation, or investment interest in this project. The description makes no claims about commercial traction or investor interest. It is a solo-built hackathon submission with no indication of commercial viability or market readiness.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.