Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,998 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
EvidenceGate is a system for evaluating code changes against explicit acceptance criteria using two independent evidence domains: internal (repository-based) and external (authoritative sources). It was built as a hackathon project by one person, Harsh Choksi, using TypeScript and GPT-5.6.
What changed
The author states they built a system that evaluates code changes based on both internal implementation evidence and external authoritative evidence, with a focus on preventing model-generated confidence scores from being mistaken for final release decisions. The system uses a deterministic policy layer to compute final results, while models are used only for bounded research and classification.
The single most important open question
Is there any evidence of real-world usage or adoption beyond the author's own demonstration? The description states no revenue, customers or traction data exist beyond what was self-reported.
What The Product Actually Is
The description states that EvidenceGate evaluates code changes against explicit acceptance criteria using two independent evidence domains: internal and external. Internal evidence comes from the repository itself, including source analysis, diffs, commands, and test results. External evidence comes from approved authoritative sources, with provenance preserved and citations tied directly to the claims they support.
A required claim passes only when both evidence domains support it. The system uses two distinct roles for GPT-5.6: Stage A gathers cited narrative and returned-source metadata from approved domains, while Stage B maps acceptance criteria and evidence IDs into structured assessments.
The system includes local validators that reject fabricated evidence IDs, unsafe URLs, malformed citation ranges, and forged gate outcomes. A versioned policy layer computes the final result deterministically.
Positioning & Claim Evolution
The description states the author's inspiration was to build something more useful than another model-generated confidence score: a release decision reviewers could examine, verify, and challenge. The system is positioned as a tool for ensuring that code changes meet explicit acceptance criteria through independent verification of both internal implementation and external authoritative sources.
The claim evolution shows a progression from dissatisfaction with existing AI confidence scores to building a system where "the final release policy should remain deterministic, versioned, and open to inspection." The author explicitly states this is not about replacing human judgment but about creating a verifiable evaluation framework.
Target Customer & ICP
Not evidenced. The description does not identify specific target customers or define an ideal customer profile (ICP). It only describes the system's functionality and the author's own use case in a hackathon context.
Business Model & Pricing Evidence
Not evidenced. The description provides no information about pricing, revenue models, or monetization strategies. It only describes the technical implementation of a system for code evaluation.
Technical & Delivery Signals
The description states that EvidenceGate was built using TypeScript and GPT-5.6. The author used Codex as their primary development collaborator to decompose requirements, design architecture, implement system, write tests, create adversarial fixtures, debug issues, prepare documentation, and organize the release.
The system includes 173 offline tests covering evidence schemas, policy enforcement, spoofed domains, fabricated citations, prompt injection, stale or conflicting sources, output limits, report escaping, and publication safety. It produces self-contained HTML reports and canonical JSON bundles.
Traction & Maturity Signals
Not evidenced. The description states this is a hackathon project submitted to the OpenAI 2026 hackathon on Devpost. No revenue, customer or traction data is available beyond what was self-reported by the author. The system has not been demonstrated in production environments or with real users.
Competitive Context
Not evidenced. The description does not mention any competitors or competitive landscape. It only describes the author's own approach to code evaluation and does not reference existing tools or systems in this space.
Key Risks & Red Flags
Risk 1
Single-person development. The system was built by one person (Harsh Choksi) with no evidence of team expansion or organizational support beyond the hackathon context.
Risk 2
Limited real-world validation. The description states this is a hackathon project with no evidence of real-world usage, customers, or adoption beyond the author's own demonstration.
Risk 3
Model dependency without clear operationalization. While the system uses GPT-5.6 for bounded research and classification, there's no indication of how this would scale or be maintained in production environments.
Risk 4
No evidence of business sustainability. The description shows no evidence of revenue streams, pricing models, or path to monetization beyond the hackathon submission.
Diligence Questions To Ask The Founders
- What specific acceptance criteria are you envisioning for real-world use cases?
- How would this system scale beyond a single-person development context?
- What are your plans for operationalizing the GPT-5.6 components in production?
- Have you identified any potential customers or use cases outside of the hackathon context?
- What would be the key metrics for success in a real-world deployment?
Investment/Partnership Verdict
Not evidenced. The description provides no information about funding rounds, valuations, headcount, or partnership opportunities. It only describes a single-person hackathon project with no evidence of commercial traction or investment interest beyond the author's own submission.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.

