OpenAI 2026 hackathon

Six of Two Hundred

An evidence-linked exception queue that checks AI-drafted customer replies against order records and policies, then sends only contradictions, missing evidence, or coverage gaps to a human.

Solo project by Dried Sandwich · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #1,934 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Six of Two Hundred is a self-reported AI workflow tool designed to validate AI-drafted customer replies by checking claims against local order records and policy evidence. It operates as an "evidence-linked exception queue" that routes only contradictions, missing evidence, or coverage gaps to human review.

What changed

The author states they built this in the context of a hackathon (OpenAI 2026), using synthetic data and a deterministic system to test integration between GPT-5.6 and local commerce records. The project includes a structured output from the model, a deterministic re-checking layer, and an operator-facing UI.

The single most important open question

Is there any evidence that this tool has been used in production or tested with real customer data, or whether it can scale beyond its current synthetic scope?

Back to contents

What The Product Actually Is

The description states:

  • Six of Two Hundred is described as an “evidence-linked exception queue for AI-drafted customer replies.”
  • It checks material claims against local order, transaction, and policy evidence.
  • It uses four outcomes: EXCEPTION, NO_EXCEPTION_DETECTED_BY_CURRENT_CHECKS, INCONCLUSIVE_EVIDENCE, and NOT_CHECKABLE.
  • The main interface is an “Outbox” rather than a dashboard — showing only the small set requiring attention.
  • A separate “Self-Test” view isolates three evidence scopes: live GPT integration, deterministic replay, and development fixtures.

Inference The tool appears to be a hybrid system combining LLM-based claim extraction with deterministic validation logic that ensures claims are grounded in actual data before routing to humans.

Back to contents

Positioning & Claim Evolution

The author states:

  • The inspiration was the problem of AI drafting full reply queues but failing at the last mile — where a person must skim all replies and becomes a rubber stamp.
  • The tool aims to avoid random sampling by checking every material claim against relevant records.
  • It is positioned as a way to reduce human effort by filtering out safe claims, not by automating the full process.

Inference The positioning evolved from a general AI workflow challenge (AI drafting customer replies) into a specific solution focused on quality control and evidence-based routing.

Back to contents

Target Customer & ICP

The description states:

  • The tool is built for use in commerce environments where AI drafts customer replies.
  • It is designed to reduce human workload by filtering out safe claims and only sending exceptions to review.
  • The author mentions “synthetic commerce domain” and “local order store,” implying internal or enterprise use cases.

Inference The likely ICP includes teams or departments managing high-volume customer service workflows, especially those using AI for drafting responses and needing quality assurance before sending replies.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

Explanation

There is no mention of pricing models, monetization strategies, or business model assumptions in the description. The project is presented as a hackathon submission with no indication of commercial intent beyond its demonstration.

Back to contents

Technical & Delivery Signals

The description states:

  • Built using Codex, Python, SQLite, Markdown policies, HTML/JavaScript UI.
  • Uses GPT-5.6 for bounded semantic claim extraction and assessment.
  • Includes deterministic code that re-fetches cited locators to check relevance and consistency.
  • Implements a guarded tool loop, claim/evidence registry, canary and replay harnesses, and budget/approval gates.
  • The author reports results from three scopes: live integration (A2R2), deterministic replay, and development fixtures.

Inference The technical approach shows a clear separation between LLM inference and deterministic validation logic, suggesting an architecture designed for safety and traceability.

Back to contents

Traction & Maturity Signals

Not evidenced.

Explanation

There is no evidence of revenue, customers, usage metrics, or product adoption beyond the author’s own testing. The project is described as a hackathon submission with synthetic data and no real-world deployment.

Back to contents

Competitive Context

Not evidenced.

Explanation

The description does not mention competitors or similar tools in the market. No reference to existing solutions for AI-generated customer reply validation or exception queues was provided.

Back to contents

Key Risks & Red Flags

  • The tool is described as a hackathon project with synthetic data and no real-world testing.
  • There is no evidence of production use, scalability, or integration with actual systems.
  • The author notes that the final evaluation was not completed due to budget and time constraints.
  • No mention of privacy, security, or operational readiness for handling real customer data.

Inference The project lacks commercial viability indicators and may not be ready for enterprise deployment without significant development and testing.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual scope of your synthetic test environment? Can it be scaled to real-world use?
  2. Have you tested this with any real customer data or human operators?
  3. How would you handle edge cases like ambiguous evidence, missing records, or model hallucinations in a production setting?
  4. Is there a plan to integrate with existing CRM or helpdesk systems?
  5. What are the key assumptions about user behavior and operator workflow that underpin this design?

Back to contents

Investment/Partnership Verdict

Not evidenced.

Explanation

There is no evidence of funding, traction, or commercial readiness to support an investment or partnership decision. The project is described as a hackathon submission with limited scope and no indication of product-market fit or scalability beyond its current synthetic demonstration.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.