OpenAI 2026 hackathon

AI Hallucination Prevention & Agent Drift Guard

Stops unsafe AI continuations before execution by detecting hallucination and agent-drift risk at token level, then producing safe rewrites and auditable replay traces.

Solo project by ÖZTÜRK TOKER · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,484 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

The description states that the project is a safety tool for AI agents, designed to detect hallucination and agent-drift risk at token level, and intervene before unsafe continuations execute. It presents a demonstration of three agent-safety scenarios with replay dashboards, audit JSON documents, and trace files. The author claims this approach allows early detection and safe rewrites, with auditable replay traces.

The project is presented as a prototype or demo submitted to an OpenAI hackathon, built using Python, JavaScript, Node.js, and Playwright, with GPT-5.6 and Codex used during Build Week for validation and documentation enhancements. The system separates prior proprietary work from new public-facing components.

Key open question

What is the actual commercial viability of this safety intervention approach? The description does not indicate any revenue, customers, or adoption beyond a hackathon submission.

Back to contents

What The Product Actually Is

The description states that the product is an AI safety tool that:

  • Detects hallucination and agent-drift risk at token level
  • Stops unsafe continuations before execution
  • Produces safe rewrites
  • Provides auditable replay traces

It includes:

  • A public-safe demo presenting three agent-safety scenarios
  • A standalone replay dashboard
  • An audit JSON document
  • A frame-by-frame JSON Lines trace
  • A readable Markdown report
  • Trigger token, stop status, intervention type, and policy fields

The system is described as demonstrating prevention-oriented approach: expose risk while an unsafe continuation is forming, stop it before execution, provide a safe rewrite, and preserve an auditable record of the intervention.

Back to contents

Positioning & Claim Evolution

The description states that this project was inspired by a safety goal: "make emerging hallucination and agent-drift risk visible early enough to stop an unsafe continuation before the demonstrated action executes."

It positions itself as:

  • A prevention-oriented approach to AI safety
  • A tool that stops unsafe continuations before execution
  • A system that provides safe rewrites and auditable replay traces

The project claims to demonstrate a "stop-before-execution" interface pattern, where the system makes interventions visible before unsafe actions advance.

Back to contents

Target Customer & ICP

Not evidenced. The description does not identify specific target customers or ideal customer profiles beyond general AI agent safety concerns.

Back to contents

Business Model & Pricing Evidence

Not evidenced. There is no information provided about pricing, business model, monetization strategy, or revenue streams.

Back to contents

Technical & Delivery Signals

The description states that the system was built using:

  • Python for artifact generation and validation
  • HTML, CSS, JavaScript for replay dashboards
  • JSON and JSON Lines for audit records
  • Node.js with Playwright for deterministic browser recording

It was built with tools including: codex, css3, github, gpt-5.6, html5, javascript, json, node.js, openai, playwright, python, schema.

The system includes:

  • An audit JSON Schema
  • Structural audit validation
  • Cross-file artifact-consistency verification
  • Portable repository paths
  • Disclosure and testing documentation
  • A clear prior-work versus Build Week record
  • An isolated Playwright-based automated demo recorder

Back to contents

Traction & Maturity Signals

Not evidenced. The description does not provide any information about traction, customers, revenue, or adoption beyond the hackathon submission.

Back to contents

Competitive Context

Not evidenced. The description does not mention competitors, market positioning, or competitive landscape.

Back to contents

Key Risks & Red Flags

  • The project is presented as a hackathon submission with no evidence of commercial traction
  • The description states that proprietary research and core implementation details were not exposed
  • The system only demonstrates proxy safety behavior, not actual runtime safeguards
  • No evidence of revenue, customers, or market adoption
  • The demonstration does not claim universal hallucination prevention
  • The project is described as a "public-safe demo" with limited scope

Back to contents

Diligence Questions To Ask The Founders

  1. What specific AI agent use cases are you targeting with this solution?
  2. How does your approach scale to production environments with multiple concurrent agents?
  3. What are the technical limitations of your current demonstration that would need to be addressed for commercial deployment?
  4. Can you describe the proprietary research and implementation details that were not included in the public demo?
  5. What is your roadmap for transitioning from this prototype to a commercial product?
  6. How do you plan to monetize this safety solution?

Back to contents

Investment/Partnership Verdict

Not evidenced. The description provides no information about funding rounds, valuations, or investment status beyond the hackathon submission. No evidence of commercial traction, revenue, or customer adoption is provided.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.