Archive position — measured, not model output
2 likes on Devpost
221 of the 7,856 archived projects have more likes, and 285 share exactly 2 — so this project's #428 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
Provenance Guard: The Accountability Engine is a GitHub app designed to detect AI-generated pull requests (PRs) that lack human understanding. It uses an AI-powered "Auditor" persona to interrogate PR authors in real time, asking follow-up questions based on their own answers to determine if they actually understand the code they submitted.
What changed
The project was built as a hackathon submission for the OpenAI 2026 hackathon. It is described as a proof-of-concept tool that integrates with GitHub and uses Codex CLI, AI models, and a live chat interface to challenge PR authors on their changes.
Single most important open question
Is there any evidence of traction or adoption beyond this single hackathon demo? The description states no revenue, customers, or usage data are available.
Note: This analysis is based entirely on the self-reported, unverified project description provided by the caller. All claims are attributed to that description and not independently verified.
What The Product Actually Is
The description states that Provenance Guard is a GitHub app that:
- Intercepts pull requests (PRs) as they are opened.
- Checks out the PR branch into a temporary local workspace.
- Runs Codex CLI in headless mode against the full repository, not just the diff.
- Flags lines that break existing conventions or introduce risk.
- Opens a live Verification Terminal for the PR author.
- Uses an AI “Auditor” persona to ask specific questions about flagged lines.
- Follow-up questions are dynamically generated based on the author's answers.
- Captures typing behavior and paste detection.
- Scores each PR based on multiple behavioral signals.
- Displays risk tiers in a maintainer dashboard with transcript replays.
Inference: The system appears to be a hybrid of AI code analysis and adversarial interaction design aimed at detecting shallow or un-understood AI-generated contributions. It is not a general-purpose linter or CI tool, but rather a specialized verification layer for PR authors.
Positioning & Claim Evolution
The description states:
- The product addresses a gap in current tools: “Nobody checks whether the person who opened the PR actually understands it.”
- It positions itself as an answer to the problem of AI-generated PRs that pass CI but are not understood by maintainers.
- It is framed as a solution to issues like:
- AI-generated PRs with thousands of lines (e.g., OCaml compiler).
- Demoralizing and draining volume of low-quality AI PRs.
- GitHub’s own response to the PR volume problem.
Claim: The product aims to improve code review efficiency by identifying un-understood or faked contributions before they reach maintainers.
Inference: It is positioned as a tool for open-source maintainers and teams managing high volumes of PRs, not as a general-purpose AI assistant or code quality tool.
Target Customer & ICP
The description states:
- The primary use case is for open-source project maintainers.
- It targets teams dealing with large volumes of AI-generated PRs.
- It is designed to help reviewers focus on PRs that actually need human attention.
Inference: The target customer segment appears to be open-source maintainers, infrastructure teams, and organizations managing high-volume code contributions.
Not evidenced: No specific customer names, size of target organizations, or use cases beyond the hackathon demo.
Business Model & Pricing Evidence
The description states:
- It is a GitHub app.
- The authors mention future plans to support self-hosted deployments, SSO, audit logs, and multi-org support.
- No pricing information, revenue model, or monetization strategy is provided.
Not evidenced: No evidence of a business model, pricing structure, or monetization approach.
Inference: If this evolves into a commercial product, it may be priced per organization or repository, but that is speculative.
Technical & Delivery Signals
The description states:
- Built with: aiml-api, codex-cli, github-api, github-app, javascript, nextjs, node.js, openai, react, typescript, webhook, websockets.
- Uses a GitHub App to listen for pull_request events via webhook.
- Runs Codex CLI in headless mode.
- Employs an ephemeral local checkout of the repository.
- Implements a “Challenge Engine” that uses AIML API for dynamic follow-ups.
- Includes a Verification Terminal with split-pane UI, keystroke timing, and paste detection.
- Uses a scoring engine that does not rely on any single signal alone.
Inference: The system is built with modern web and AI tooling, and shows some sophistication in its adversarial design.
Not evidenced: No evidence of production-grade infrastructure, scalability, or deployment details beyond the demo.
Traction & Maturity Signals
The description states:
- This was a hackathon submission.
- The team size is 4.
- No revenue, customers, or usage data are provided.
- The authors mention future plans for self-hosted options and multi-org support.
Not evidenced: No evidence of traction, adoption, or user engagement beyond the demo.
Inference: The product is at a very early stage — a prototype with no commercial traction.
Competitive Context
The description states:
- Linters, CI tools, and bots like CodeRabbit already check code quality.
- The gap this fills is not whether the code is good, but whether the author understands it.
- It is positioned as a tool to combat AI-generated slop PRs that pass automated checks.
Inference: It competes with existing code review tools by focusing on behavioral signals and author understanding rather than code correctness.
Not evidenced: No mention of direct competitors or market positioning against specific tools.
Key Risks & Red Flags
- The product is described as a hackathon demo, not a production-ready tool.
- It is unclear whether the system can scale to real-world usage without performance or reliability issues.
- The adversarial design relies on detecting behavioral signals — which may be unreliable or evaded by determined actors.
- No evidence of any commercial traction or user feedback beyond the authors' own claims.
- The system’s reliance on a live chat interface for verification may not be scalable or practical in large teams.
Inference: The product is experimental and unproven in real-world settings.
Not evidenced: No data on performance, scalability, or user experience outside of the demo.
Diligence Questions To Ask The Founders
- What was the actual outcome of the hackathon submission? Was it accepted or awarded?
- Are there any early adopters or pilot users beyond the team?
- How does the system handle edge cases like PRs with no code changes or very small diffs?
- What are the technical limitations of running Codex CLI in headless mode on a local checkout?
- How is the AI-generated follow-up logic tested to ensure it cannot be gamed by external tools?
- What is the plan for monetization and commercial deployment beyond the demo?
- Has the team considered privacy or compliance implications of storing PR transcripts?
Investment/Partnership Verdict
Not evidenced: No evidence of a business case, traction, or financials to support an investment or partnership decision.
Inference: This is a conceptually interesting idea with potential for a real-world application, but it is currently at the prototype stage. It lacks commercial viability or traction as of the provided description.
Confidence level: Low — based on self-reported evidence only and no independent verification.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.
