OpenAI 2026 hackathon

Evidence Chain

Claim assurance for AI-generated professional reports

Solo project by Ding Wang · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,987 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be: Evidence Chain is a single-person project that claims to offer a methodology-aware assurance workflow for professional reports. The author states it uses AI (specifically GPT-5.6 Luna) and structured workflows to test whether claims in reports meet human-approved professional methodologies, not just whether text can be retrieved.

What changed: This appears to be an early-stage prototype submitted to the OpenAI 2026 hackathon. It includes five synthetic cases demonstrating how it evaluates claims against evidence using five explicit tests (provenance, construct validity, attribution, contradiction, and coverage). The system is described as deterministic in its status assignment, with human approval boundaries for both methodology and final wording.

The single most important open question: Is there any evidence that this product has been used or tested by professionals beyond the author’s own synthetic cases? The description states no revenue, customers, or traction data exist beyond what the author reports.

Back to contents

What The Product Actually Is

The description states that Evidence Chain is a methodology-aware assurance workflow for professional reports. It:

  • Turns framework PDFs into structured evidence contracts.
  • Requires a professional to review and lock that contract before it governs an audit.
  • Examines registered PDFs, images, and timestamped transcripts.
  • Retrieves supporting, contradictory, attribution, and coverage evidence.
  • Runs five explicit tests covering provenance, construct validity, attribution, contradiction, and coverage.
  • Validates model-supplied sources and requirement references.
  • Applies deterministic application rules to assign Supported, Contested, Unsupported, or Insufficient evidence.
  • Proposes narrower, evidence-bounded wording for human approval, then retests it.

It is described as a single-page application built with JavaScript, Vite, and Firebase Hosting, with server-side operations running in Firebase Cloud Functions using the OpenAI Responses API. The author states that GPT-5.6 Luna is used through the OpenAI Responses API to interpret source materials, compile methodology requirements, classify evidence, adjudicate tests, and propose wording.

Not evidenced: No actual deployed product, no real-world use cases, no customer data, or live integrations are described.

Back to contents

Positioning & Claim Evolution

The author states that AI-assisted reports can quote sources accurately but still reach invalid conclusions. The project positions itself as addressing the gap between accurate citation and professional validity.

It claims to test whether each consequential claim satisfies a human-approved professional methodology, not just whether related text can be retrieved.

The positioning implies a role in professional assurance, particularly for reports that must meet specific standards or frameworks, such as compliance, leadership development, or organisational assessments.

Inference: The project is positioned as a tool to prevent false claims from appearing valid due to AI-generated content that sounds plausible but lacks proper evidence. This is a response to the growing use of AI in professional writing and reporting.

Back to contents

Target Customer & ICP

The description states that Evidence Chain is designed for professionals who create or review reports that must meet specific methodologies or standards, such as:

  • Organisational assessment
  • Operational review
  • Project benefits
  • Leadership development
  • Compliance assurance

It is described as targeting a professional audience, not general users. The system requires human approval of methodology and final wording, suggesting it’s built for domain experts or auditors.

Not evidenced: No specific customer personas, use cases beyond synthetic examples, or target industries are provided.

Back to contents

Business Model & Pricing Evidence

The description does not state any business model or pricing structure. It is a single-person hackathon project with no mention of monetisation, licensing, or subscription models.

Not evidenced: No revenue streams, pricing tiers, or commercial arrangements are described.

Back to contents

Technical & Delivery Signals

The system is built as a single-page application (SPA) using:

  • JavaScript
  • Vite
  • Firebase Hosting
  • Firebase Cloud Functions (second-generation)
  • GPT-5.6 Luna through OpenAI Responses API

It uses strict structured outputs and disables provider-side storage. The system validates returned structures and references before deterministic code assigns a status.

The author states that the interface includes:

  • Three visible chapters: Understand, Test, and Resolve
  • Seven guarded internal phases
  • An Evidence Graph showing how raw sources become validated evidence, tests, and decisions

Not evidenced: No information on scalability, performance metrics, or deployment infrastructure beyond Firebase.

Back to contents

Traction & Maturity Signals

The project is described as a single-person hackathon submission, with no mention of:

  • Customers
  • Revenue
  • Users
  • Product adoption
  • Market traction

It includes:

  • Five complete synthetic cases
  • 248 automated tests
  • A deployed walkthrough that judges can complete in about one minute

Not evidenced: No real-world usage, user feedback, or product maturity beyond the prototype stage.

Back to contents

Competitive Context

The description does not mention any competitors. It is a self-contained project submitted to a hackathon and does not reference existing tools for:

  • AI-generated content verification
  • Professional report auditing
  • Evidence-based claim validation

Not evidenced: No competitive analysis, no comparison to existing platforms or methodologies.

Back to contents

Key Risks & Red Flags

  • Single-person project: The entire system is built by one person (Ding Wang), raising questions about scalability and long-term maintenance.
  • No real-world use: All cases are synthetic; there is no evidence of actual professional adoption or testing.
  • Unverified claims: The product is described as a prototype, not a production-ready tool.
  • Limited scope: It only works with PDFs, images, and transcripts, and is limited to five specific test types.
  • No commercial viability: No pricing, monetisation, or business model is described.

Back to contents

Diligence Questions To Ask The Founders

  1. What professional domains or methodologies have you tested this system against beyond the synthetic cases?
  2. How do you plan to scale beyond a single-person development effort?
  3. Have you conducted any user studies or trials with domain experts?
  4. What are your plans for integrating real-world data sources and workflows?
  5. How do you intend to monetize or commercialize this product?

Back to contents

Investment/Partnership Verdict

Not evidenced: No financials, traction, or investor interest are described.

This is a single-person hackathon prototype, not a product with demonstrated market demand or commercial viability. The author describes it as a proof-of-concept for a methodology-aware assurance system, but there is no evidence of real-world adoption, revenue, or customer feedback.

Inference: If this project were to evolve into a product, it would need significant development to move beyond synthetic cases and gain traction with professionals in reporting and auditing roles. The idea has potential, but the current version is not ready for investment or partnership consideration.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.