OpenAI 2026 hackathon

VibeProof — prove you own your AI-assisted code

Build with AI, prove you know why it works. VibeProof is an AI-allowed Ownership Challenge that records how a candidate debugs a live incident, then gives recruiters a transparent Proof Replay.

Solo project by Seb Lew · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #7,550 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

VibeProof is a self-reported AI-allowed engineering ownership challenge tool designed for technical hiring. The product allows candidates to debug a simulated incident in a controlled environment and generates a "Proof Replay" — a chronological, cited evidence trail of their actions that recruiters can review without needing raw logs.

What changed

The project was submitted as part of the OpenAI 2026 hackathon. It is described as a research-informed prototype with no revenue or customer data. The team built it using AI tools like GPT-5.6 and OpenAI Codex, and deployed it using Docker, FastAPI, Godot, Node.js, Python, Railway, Remotion, TypeScript, and Vitest.

Single most important open question

Is there evidence that the tool can be scaled beyond a hackathon prototype to deliver consistent, reliable, and interpretable ownership signals in real-world hiring workflows?

Back to contents

What The Product Actually Is

The description states:

  • VibeProof is an AI-allowed engineering Ownership Challenge.
  • A candidate is dropped into a controlled "Homepage Latency Spike" incident with metrics, logs, traces, source code, and an AI assistant.
  • They investigate, record, revise a hypothesis, verify a fix, and submit a decision.
  • Every observable action is logged, then scored by a deterministic rubric into a Proof Replay — a chronological, cited evidence trail a recruiter can read without touching raw logs.

Inference The tool appears to be a simulation-based assessment platform for evaluating engineering ownership in technical interviews. It uses AI-assisted debugging and logs candidate actions to generate a structured report.

Not evidenced

  • Whether the Proof Replay is actually generated or used in practice.
  • The exact rubric scoring mechanism or how it evaluates ownership.
  • How the tool integrates with existing hiring platforms or workflows.

Back to contents

Positioning & Claim Evolution

The description states:

  • VibeProof is positioned as a way to prove AI-assisted code ownership in technical hiring.
  • It aims to help hiring teams with little senior-engineer time gather evidence before full interviews.
  • The tool does not try to detect who typed each line; it asks whether the candidate can own the outcome.

Inference The positioning is centered on transparency and ownership verification, especially in an AI-augmented hiring landscape. It claims to shift focus from output to process and explanation.

Not evidenced

  • How this differs from existing tools or methodologies in technical hiring.
  • Whether the tool has been tested with real recruiters or hiring teams.
  • The evolution of its positioning since the hackathon submission.

Back to contents

Target Customer & ICP

The description states:

  • Hiring teams with little senior-engineer time.
  • Technical recruiters, particularly in Malaysia (mentioned as next step).
  • Candidates undergoing technical interviews for engineering roles.

Inference The primary customer is hiring managers or recruiters who want to assess ownership and reasoning skills of candidates without extensive manual review.

Not evidenced

  • Specific job titles or seniority levels targeted.
  • The size or type of organizations using the tool.
  • Whether the tool targets specific engineering roles (e.g., backend, frontend, DevOps).

Back to contents

Business Model & Pricing Evidence

The description states:

  • VibeProof is a research-informed prototype — final hiring decisions stay with people.
  • No pricing or monetization model is described.

Inference There is no evidence of a business model or pricing structure at this stage. The tool appears to be in early-stage development and not yet commercialized.

Not evidenced

  • Revenue streams, licensing fees, or subscription models.
  • Plans for monetization or scaling beyond the hackathon.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with Docker, FastAPI, GDScript, Godot, GPT-5.6, Node.js, OpenAI Codex, Python, Railway, Remotion, TypeScript, Vitest.
  • Candidate app: Godot Incident Room (GDScript), exported to Web and deployed on Railway.
  • Backend: Python + FastAPI — an append-only event log, versioned scenarios, deterministic rubric scoring, and a Proof Replay report with cited evidence.
  • Codex drove the majority of the build, including product framing, implementation plan, scaffold, scenario loader, Web export, and Railway deployment.

Inference The tool is built using modern development stacks and AI-assisted tools, suggesting an engineering team with technical depth. It uses a deterministic scoring system to ensure reproducibility.

Not evidenced

  • The scalability or performance of the backend in real-world use.
  • How the tool handles concurrent users or large-scale deployment.
  • Whether the Proof Replay is actually generated or used in practice.

Back to contents

Traction & Maturity Signals

The description states:

  • Submitted to the OpenAI 2026 hackathon.
  • VibeProof is a research-informed prototype — final hiring decisions stay with people.
  • Next steps include interviewing Malaysian technical recruiters, running pilots, and measuring reviewer agreement against the automated rubric.

Inference The tool is in early-stage development and has not yet been deployed or tested in real-world hiring environments. It is described as a prototype with no revenue or customer data.

Not evidenced

  • Any actual users or pilot results.
  • Metrics on reviewer agreement or tool effectiveness.
  • Customer feedback or adoption rates.

Back to contents

Competitive Context

The description states:

  • AI now writes polished resumes, take-homes, and working code.
  • The tool aims to gather evidence of ownership and reasoning, not just output.

Inference VibeProof positions itself as a response to the rise of AI-generated artifacts in technical hiring, aiming to assess deeper understanding and responsibility.

Not evidenced

  • Direct competitors or similar tools in the market.
  • How it compares to existing platforms for technical assessment or ownership verification.
  • Market size or demand for such tools.

Back to contents

Key Risks & Red Flags

The description states:

  • The tool is a research-informed prototype.
  • Final hiring decisions stay with people.
  • Challenges included keeping scoring evidence-first, reconciling concurrent Codex sessions, and verifying Web builds.

Inference

Key risks include:

  • Lack of real-world testing or validation.
  • Over-reliance on AI tools for development and assessment.
  • Unclear scalability or integration into existing hiring workflows.

Not evidenced

  • Any data on user experience or tool reliability.
  • Whether the deterministic rubric is robust or interpretable.
  • How the tool handles edge cases or unexpected inputs.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific ownership signals are you trying to detect, and how do they differ from traditional technical assessments?
  2. How does the deterministic rubric score candidates, and what evidence supports its validity?
  3. Have you tested the tool with real recruiters or hiring teams? If so, what were the results?
  4. What is your plan for scaling beyond a hackathon prototype to a production-ready solution?
  5. How do you plan to integrate VibeProof into existing hiring platforms or workflows?
  6. What are the key assumptions underlying the product’s design and scoring system?

Back to contents

Investment/Partnership Verdict

The description states:

  • VibeProof is a research-informed prototype — final hiring decisions stay with people.
  • It was built in a hackathon environment using AI tools like GPT-5.6 and Codex.

Inference At this stage, the tool is not ready for investment or partnership. It lacks traction, revenue, or customer validation. The idea has potential but requires significant development and testing before it can be considered viable.

Not evidenced

  • Any financials, user base, or market traction.
  • A clear path to monetization or product-market fit.
  • Evidence of a scalable business model or competitive advantage.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.