Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,786 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
Done Yet? is a self-reported tool that checks whether AI agents have actually completed tasks by observing system states after agent actions. It applies "design by contract" principles at the task closeout boundary, using deterministic postcondition checks against typed acceptance contracts.
What changed
The project description indicates this is a proof-of-concept submitted for the OpenAI 2026 hackathon. It includes a CLI, React judge console, and Codex plugin to validate agent behavior through adversarial testing and retry-stable verification.
Single most important open question
Is there any evidence of real-world adoption or integration beyond the hackathon demo? The description states no revenue, customers, or traction data exist beyond its own claims.
What The Product Actually Is
The description states that Done Yet? is a tool that:
- Turns intended outcomes into typed acceptance contracts.
- Observes pre-, post-, and retry states.
- Runs deterministic postcondition checks (PASS, FAIL, HOLD).
- Uses a verification engine across CLI, tests, fixtures, filesystem observer, and React judge console.
- Includes a Codex plugin to enforce task completion only when a passing report is generated.
It also claims:
- A real filesystem observer rejects confident claims if no edit landed or if protected configuration changes.
- A synthetic helpdesk provides adversarial test cases (false success, partial commit, etc.).
- GPT-5.6 translates natural-language intent into contracts and helps explain/repair failures.
- The system evaluates canonical state rather than agent tone or tool responses.
Inference The product appears to be a narrow, technical validation layer for AI agents — not a general-purpose dashboard or trust platform.
Positioning & Claim Evolution
The description states:
- Done Yet? began with the question: “after the agent acts, do the systems now satisfy the user's actual acceptance criteria?”
- It is described as applying "design by contract" to agent side effects.
- The tool is not about scoring agents but about closeout checks against the world the agent was supposed to change.
Inference Positioning has evolved from a hackathon idea into a framework for verifying task completion in AI workflows. It frames itself as a technical validation layer, not a trust or monitoring platform.
Target Customer & ICP
The description does not state any target customer or ideal customer profile (ICP). It only describes the tool’s function and its use in adversarial testing during development.
Not evidenced No information on who uses Done Yet? beyond the author's own demonstration.
Business Model & Pricing Evidence
The description states:
- No account, API key, database, or customer data is required.
- The tool is presented as a CLI and React console for developers to test agent behavior.
- It was built using Codex and submitted to a hackathon.
Not evidenced No pricing model, monetization strategy, or business model is described. The project appears to be a proof-of-concept with no commercial traction.
Technical & Delivery Signals
The description states:
- Built with: Cloudflare Pages, Codex, GitHub, GPT-5.6, JavaScript, Node.js, Playwright, React, Vite.
- Uses JSON-pointer paths for checks (exists, equals, count, relation, unchanged, retry-stability).
- Includes a Codex plugin that enforces a Stop hook to prevent task closure without passing verification.
- Has 19 automated tests covering verifier, observer, CLI exit codes, and contract lifecycle.
- Provides six reproducible adversarial fixtures for testing.
Inference The tool is built with modern developer tools and includes a strong test suite. It uses a modular architecture that supports both local and plugin-based execution.
Traction & Maturity Signals
The description states:
- Submitted to the OpenAI 2026 hackathon.
- Includes a demo run:
npm installandnpm run demo:repo. - Has a live judge console for testing.
- Contains 19 automated tests and six adversarial fixtures.
Not evidenced No evidence of revenue, customers, or adoption beyond the author’s own submission. No data on usage, retention, or product-market fit is provided.
Competitive Context
The description does not mention any competitors or direct market context. It positions Done Yet? as a tool for validating AI agent behavior, but does not describe how it compares to existing tools in this space.
Not evidenced No competitive analysis, market positioning, or comparison to other agent validation or monitoring platforms is provided.
Key Risks & Red Flags
- No commercial traction or adoption: The project is a hackathon submission with no evidence of real-world usage.
- Unverified claims: All descriptions are self-reported and unverified.
- Limited scope: The tool is described as a proof-of-concept, not a production-ready product.
- Unclear scalability: No indication of how it would scale beyond the current adversarial test cases or developer use case.
Inference The project lacks commercial viability or market readiness. It may be useful for developers but does not appear to have a path to monetization or growth.
Diligence Questions To Ask The Founders
- What is the actual use case you're solving for beyond the hackathon demo?
- Is there any plan to integrate Done Yet? into existing AI agent platforms or workflows?
- How would this tool scale beyond the current adversarial test cases and developer-focused interface?
- Are there any early adopters or pilot users of this tool?
- What are the technical limitations of the current implementation that would need to be addressed for production use?
Investment/Partnership Verdict
The description states that Done Yet? is a hackathon submission with no revenue, customers, or traction data. It is presented as a proof-of-concept for validating AI agent behavior through deterministic checks.
Not evidenced No evidence of commercial viability, product-market fit, or scalability beyond the current demo.
Verdict This project appears to be an early-stage idea or prototype with no demonstrated traction or business model. It is not ready for investment or partnership consideration based on the self-reported description alone.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.
