OpenAI 2026 hackathon

SupportTrace Evidence Gate

A keyless RAG support console that attaches source evidence, scores citation confidence, and routes risky or weak drafts to human review.

Solo project by baiyuxi Bai · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #7,062 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

SupportTrace Evidence Gate is a self-reported keyless RAG (Retrieval-Augmented Generation) support console that retrieves knowledge, composes deterministic drafts, attaches source excerpts, scores citation confidence, and routes risky or weak drafts to human review. It is described as a TypeScript-based application with a Fastify API and React/Vite dashboard.

What changed

The project was built during the OpenAI 2026 hackathon, with an emphasis on demonstrating a deterministic, evidence-driven approach to support draft generation and review. The author states it was developed using Codex with GPT-5.6 for design, implementation, testing, and documentation.

Single most important open question

Is there any evidence of real-world usage or integration beyond the hackathon demo? The description does not indicate whether this system has been deployed in production or tested on live support traffic.

Back to contents

What The Product Actually Is

The description states that SupportTrace Evidence Gate is a keyless RAG support console. It retrieves approved knowledge, composes deterministic drafts, and attaches exact source excerpts. It calculates a transparent citation-relevance heuristic, then emits one of three explicit decisions: ready_to_send, human_review_required, or blocked. It also blocks prompt injection before release.

It is described as a TypeScript application with:

  • Fastify API
  • React/Vite console
  • Deterministic local embeddings
  • In-process vector store
  • Vitest coverage
  • Docker files and CI

The system is exposed through a POST endpoint /api/support/draft and includes a judge-visible dashboard.

Inference The product appears to be a proof-of-concept or demo-level tool, not a production-ready SaaS offering. It was built for a hackathon and is described as deterministic, keyless, and local.

Back to contents

Positioning & Claim Evolution

The author states that the system addresses a problem: “Support drafts can sound confident even when retrieval is weak.” The product aims to give operators more than fluent prose — it provides exact evidence, confidence scoring, and explicit routing logic.

It positions itself as a tool for evidence-based support workflows, where human review is not a fallback but an explicit product state. It also claims to block prompt injection and separate evidence quality from policy risk.

Inference The positioning implies a shift from generic AI-generated support to a controlled, traceable, and accountable workflow — likely targeting enterprises or teams managing sensitive customer interactions.

Back to contents

Target Customer & ICP

The description does not explicitly name target customers. However, it is implied that the system is designed for:

  • Teams managing customer support workflows
  • Organizations needing evidence-based decision-making in support
  • Users who want to retain control over sensitive or high-risk cases

It is described as a support console, suggesting a B2B SaaS or internal tooling use case.

Inference The ICP likely includes enterprise support teams, compliance-focused organizations, or customer service platforms that require auditability and risk control in AI-generated responses.

Back to contents

Business Model & Pricing Evidence

The description does not contain any information about pricing, monetization, or business model. It is a self-reported demo project built for a hackathon.

Inference There is no evidence of a commercial model, revenue, or pricing structure beyond the author’s own account.

Back to contents

Technical & Delivery Signals

  • Built with TypeScript
  • Uses Fastify API
  • Includes React/Vite console
  • Implements deterministic local embeddings
  • Employs an in-process vector store
  • Has Vitest coverage and Docker files
  • CI/CD pipeline is included
  • The system is described as keyless, local, and deterministic

The author states that Codex with GPT-5.6 was used for design, implementation, testing, documentation, and privacy verification.

Inference The technical stack suggests a developer-focused tool, likely built for internal use or demonstration. It is not described as cloud-hosted or scalable beyond the demo environment.

Back to contents

Traction & Maturity Signals

The project is described as a hackathon submission, built during Build Week, and includes:

  • Eight passing tests
  • A keyless local judge path
  • Public-safe demo
  • Reproducible documentation

There is no evidence of:

  • Customers
  • Revenue
  • Product adoption
  • Live usage
  • Production deployment

Inference The system is at a demo or prototype stage, with no indication of traction or real-world maturity.

Back to contents

Competitive Context

The description does not mention competitors. However, it implies a role in the RAG and AI support workflow space, where tools like:

  • LangChain
  • LlamaIndex
  • Rasa
  • Zendesk AI
  • Salesforce Einstein

are present. The product’s focus on evidence attachment, confidence scoring, and explicit routing may differentiate it from general-purpose RAG systems.

Inference It is positioned to compete in a niche where traceability, auditability, and risk control are critical — but no specific competitors are named or compared.

Back to contents

Key Risks & Red Flags

  • The system is described as deterministic, keyless, and local, which may limit scalability or integration.
  • It was built for a hackathon, with no evidence of production use or real-world testing.
  • There is no mention of data privacy, compliance, or enterprise readiness.
  • The threshold score (0.40) is described as a demo policy, not calibrated.
  • No evidence of customer feedback, error rates, or performance metrics beyond tests.

Inference The product may be too early-stage for commercial deployment and lacks any indication of real-world validation or enterprise-readiness.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual use case or customer problem you are solving?
  2. Have you tested this system on real support traffic or with real users?
  3. How does the 0.40 confidence threshold translate into business impact or error costs?
  4. Is there a plan to move beyond the demo environment (e.g., cloud deployment, integration)?
  5. What is the long-term vision for this product — is it intended as an internal tool or a SaaS offering?
  6. How do you plan to calibrate the confidence scoring with real-world data?

Back to contents

Investment/Partnership Verdict

The project is described as a hackathon demo and lacks evidence of traction, revenue, customers, or commercial viability.

Inference At this stage, it is not suitable for investment or partnership unless there is a clear path to product-market fit and real-world testing. It may be a pre-product idea, not a product in development.

Confidence Level Low — based on self-reported, unverified evidence only.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.