OpenAI 2026 hackathon

GovernDiff

Test business policy changes before they ship. GovernDiff turns draft rules into deterministic impact previews, showing who is affected, what changed, and why.

Solo project by Kouichi Namiki · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #4,367 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

GovernDiff is a local web application that processes written business policy into structured rules using GPT-5.6 and evaluates changes against synthetic records to show impact. It allows users to test draft policy versions before publishing, showing who is affected, what changed, and why.

What changed

The author describes a shift from traditional document-based policy review to a system that uses AI for language parsing and deterministic code for impact calculation. The tool separates "language boundary" (AI) from "decision boundary" (deterministic logic).

Single most important open question

Does the author's self-reported technical approach actually work as described, or is this a demonstration with limited real-world applicability?

Analysis basis

This report is based entirely on the self-reported project description supplied by the caller. All claims are unverified and should be treated as stated by the author only.

Back to contents

What The Product Actually Is

The description states that GovernDiff is:

  • A local web application
  • Designed to turn written policy into structured, reviewable rules
  • Capable of testing policy versions against synthetic records
  • A tool for reviewing not only how policy text changed, but what that change would do
  • Not designed to approve, reject, reimburse, pay, or submit expenses

The system uses:

  • GPT-5.6 for language processing (structured outputs)
  • Deterministic TypeScript code for impact calculations
  • Synthetic expense records (48 in demo)
  • A two-boundary architecture: language boundary and decision boundary

Evidence The author's own write-up describes the product as a local web application that processes policy into rules and evaluates changes against synthetic data.

Back to contents

Positioning & Claim Evolution

The description states:

  • The tool was inspired by software teams' use of tests and diffs before deploying code
  • It aims to create a similar review process for business policy
  • The name combines "governance" and "diff"
  • The author positions it as a way to test business policy changes before they ship

The claim evolution shows:

  1. Initial inspiration: comparing policy reviews to software deployment processes
  2. Core positioning: a tool for reviewing policy impact, not executing decisions
  3. Technical approach: AI for language parsing, deterministic code for results

Evidence The author's own write-up describes the inspiration and positioning as being about applying software development practices to business policy review.

Back to contents

Target Customer & ICP

The description states:

  • The tool is designed for people reviewing business policies
  • It's intended for "judges and reviewers" who want to test complete product without credentials
  • The demo focuses on expense policy review
  • It's a local web application, suggesting individual or small team use

Evidence The author describes the target as "judges and reviewers" and focuses on expense policy scenarios.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

Evidence No information provided about pricing, monetization, or business model in the description.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with: codex, css, csv-parse, git, gpt-5.6, next.js, node.js, openai, pnpm, react, responses, structured, typescript, vitest, zod
  • Uses two boundaries: language boundary (GPT-5.6) and decision boundary (deterministic code)
  • Implements strict schema validation of model outputs
  • Uses AbortController for handling asynchronous requests safely
  • Has 180 passing automated tests
  • Uses a 13-phase state machine and 12-event reducer
  • Implements immutable release snapshots
  • Separates Sample Mode from Live Mode

Evidence The author's own write-up describes the technical architecture, implementation details, and testing approach.

Back to contents

Traction & Maturity Signals

Not evidenced.

Evidence No information provided about revenue, customers, adoption, or traction beyond the demo scenario.

Back to contents

Competitive Context

Not evidenced.

Evidence No mention of competitors or market positioning in the description.

Back to contents

Key Risks & Red Flags

Inferences based on self-reported information:

  1. The tool is described as a local web application with no clear path to enterprise deployment
  2. It's built for one person (team size: 1) and may lack team collaboration features
  3. The demo uses synthetic records only, not real operational data
  4. The approach requires human resolution of ambiguities, which may limit scalability
  5. The tool is described as a hackathon submission with no indication of production readiness

Evidence The author's own write-up describes the tool as a demonstration project with limited scope and no clear path to commercialization.

Back to contents

Diligence Questions To Ask The Founders

  1. How does the current synthetic-only approach translate to real-world policy scenarios?
  2. What are the actual limitations of using GPT-5.6 for policy language parsing at scale?
  3. How would this tool integrate with existing enterprise systems or workflows?
  4. What is the path from this demo to a production-ready solution?
  5. How does the human-in-the-loop approach scale beyond the current demonstration?
  6. What are the actual use cases beyond expense policy that this system could handle?

Inference These questions arise from the self-reported description's limitations and lack of commercialization details.

Back to contents

Investment/Partnership Verdict

Not evidenced.

Evidence No information provided about investment interest, partnership opportunities, or commercial viability in the description.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.