OpenAI 2026 hackathon

EvidenceGate

Code evidence, source evidence, and a release decision you can inspect.

Solo project by Harsh Choksi · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,998 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

EvidenceGate is a system for evaluating code changes against explicit acceptance criteria using two independent evidence domains: internal (repository-based) and external (authoritative sources). It was built as a hackathon project by one person, Harsh Choksi, using TypeScript and GPT-5.6.

What changed

The author states they built a system that evaluates code changes based on both internal implementation evidence and external authoritative evidence, with a focus on preventing model-generated confidence scores from being mistaken for final release decisions. The system uses a deterministic policy layer to compute final results, while models are used only for bounded research and classification.

The single most important open question

Is there any evidence of real-world usage or adoption beyond the author's own demonstration? The description states no revenue, customers or traction data exist beyond what was self-reported.

Back to contents

What The Product Actually Is

The description states that EvidenceGate evaluates code changes against explicit acceptance criteria using two independent evidence domains: internal and external. Internal evidence comes from the repository itself, including source analysis, diffs, commands, and test results. External evidence comes from approved authoritative sources, with provenance preserved and citations tied directly to the claims they support.

A required claim passes only when both evidence domains support it. The system uses two distinct roles for GPT-5.6: Stage A gathers cited narrative and returned-source metadata from approved domains, while Stage B maps acceptance criteria and evidence IDs into structured assessments.

The system includes local validators that reject fabricated evidence IDs, unsafe URLs, malformed citation ranges, and forged gate outcomes. A versioned policy layer computes the final result deterministically.

Back to contents

Positioning & Claim Evolution

The description states the author's inspiration was to build something more useful than another model-generated confidence score: a release decision reviewers could examine, verify, and challenge. The system is positioned as a tool for ensuring that code changes meet explicit acceptance criteria through independent verification of both internal implementation and external authoritative sources.

The claim evolution shows a progression from dissatisfaction with existing AI confidence scores to building a system where "the final release policy should remain deterministic, versioned, and open to inspection." The author explicitly states this is not about replacing human judgment but about creating a verifiable evaluation framework.

Back to contents

Target Customer & ICP

Not evidenced. The description does not identify specific target customers or define an ideal customer profile (ICP). It only describes the system's functionality and the author's own use case in a hackathon context.

Back to contents

Business Model & Pricing Evidence

Not evidenced. The description provides no information about pricing, revenue models, or monetization strategies. It only describes the technical implementation of a system for code evaluation.

Back to contents

Technical & Delivery Signals

The description states that EvidenceGate was built using TypeScript and GPT-5.6. The author used Codex as their primary development collaborator to decompose requirements, design architecture, implement system, write tests, create adversarial fixtures, debug issues, prepare documentation, and organize the release.

The system includes 173 offline tests covering evidence schemas, policy enforcement, spoofed domains, fabricated citations, prompt injection, stale or conflicting sources, output limits, report escaping, and publication safety. It produces self-contained HTML reports and canonical JSON bundles.

Back to contents

Traction & Maturity Signals

Not evidenced. The description states this is a hackathon project submitted to the OpenAI 2026 hackathon on Devpost. No revenue, customer or traction data is available beyond what was self-reported by the author. The system has not been demonstrated in production environments or with real users.

Back to contents

Competitive Context

Not evidenced. The description does not mention any competitors or competitive landscape. It only describes the author's own approach to code evaluation and does not reference existing tools or systems in this space.

Back to contents

Key Risks & Red Flags

Risk 1

Single-person development. The system was built by one person (Harsh Choksi) with no evidence of team expansion or organizational support beyond the hackathon context.

Risk 2

Limited real-world validation. The description states this is a hackathon project with no evidence of real-world usage, customers, or adoption beyond the author's own demonstration.

Risk 3

Model dependency without clear operationalization. While the system uses GPT-5.6 for bounded research and classification, there's no indication of how this would scale or be maintained in production environments.

Risk 4

No evidence of business sustainability. The description shows no evidence of revenue streams, pricing models, or path to monetization beyond the hackathon submission.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific acceptance criteria are you envisioning for real-world use cases?
  2. How would this system scale beyond a single-person development context?
  3. What are your plans for operationalizing the GPT-5.6 components in production?
  4. Have you identified any potential customers or use cases outside of the hackathon context?
  5. What would be the key metrics for success in a real-world deployment?

Back to contents

Investment/Partnership Verdict

Not evidenced. The description provides no information about funding rounds, valuations, headcount, or partnership opportunities. It only describes a single-person hackathon project with no evidence of commercial traction or investment interest beyond the author's own submission.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.