Archive position — measured, not model output
1 like on Devpost
506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #1,029 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
Evidence-to-Test is a self-reported tool that captures browser walkthroughs and converts them into Playwright tests using AI agents. The author states it aims to create a verifiable boundary between observing UI actions and turning them into reliable automation, focusing on integrity, context, and outcome verification.
What changed
The project was built during the OpenAI 2026 hackathon. It evolved from an existing recorder foundation into a standalone product with explicit integrity boundaries, deterministic pipelines, and tamper detection features.
Single most important open question
Is there any evidence of real-world usage or traction beyond the synthetic demo? The description contains no data on customers, revenue, adoption, or product-market fit beyond the author’s own account.
What The Product Actually Is
The description states that Evidence-to-Test:
- Captures a fresh Chromium walkthrough.
- Compiles it into a "promoted evidence contract".
- Every action includes page/popup context, before/after state, proved locator, required waits, correlated outcomes, integrity hashes, and explicit unknowns.
- Uses AI agents (specifically Codex) to generate Playwright scenarios from this bounded contract.
- Includes deterministic verification of the generated test via fresh replay.
- Demonstrates tamper rejection by rejecting modified artifacts while preserving original evidence.
The tool is built in TypeScript and Node.js using technologies such as Playwright, Chromium, Zod, and OpenAI Codex. It runs entirely locally without external credentials or APIs.
Inference This appears to be a developer tool focused on improving the reliability of end-to-end browser automation by enforcing integrity checks and deterministic workflows. It is not a general-purpose recorder but rather one that emphasizes correctness over speed or ease-of-use.
Positioning & Claim Evolution
The author claims:
- Browser recorders often fail to capture intent, especially in brittle legacy UIs.
- Passing noisy recordings to coding agents only moves the guessing downstream.
- Evidence-to-Test creates a verifiable boundary between observation and automation.
- It promotes deterministic software for fact establishment, AI synthesis within those facts, and final replay verification.
Inference The positioning is that of a specialized tool for developers who want reliable, auditable browser automation. It positions itself as solving problems in brittle UIs where traditional recorders fall short — not as a general-purpose automation solution.
Target Customer & ICP
The description does not name specific customer segments or personas. However, it implies:
- Developers working with legacy or complex UIs.
- Teams seeking reliable end-to-end testing workflows.
- Users who value deterministic verification and integrity in their test pipelines.
Inference Based on the technical stack (Playwright, Chromium) and focus on reliability, the ICP likely includes developers or QA engineers in enterprise or SaaS environments where stable automation is critical.
Business Model & Pricing Evidence
No information is provided about pricing, monetization strategy, or business model. The project is described as a hackathon submission with no mention of commercial plans or revenue streams.
Not evidenced
Technical & Delivery Signals
The author states:
- Built using TypeScript, Node.js, Playwright, Chromium, Zod.
- Pipeline includes: fresh capture → immutable snapshot → evidence analysis → safety and promotion gates → bounded contract → Playwright scenario → fresh replay.
- Uses SHA-256 hashes for integrity tracking.
- Supports deterministic verification of scenarios via fresh replay.
- Demonstrates tamper rejection with real OUTPUT_MISMATCH errors.
- Runs locally without external dependencies or credentials.
Inference The tool is technically sophisticated and designed around integrity, determinism, and isolation. It avoids reliance on cloud services or AI accounts for core functionality.
Traction & Maturity Signals
There is no evidence of traction beyond the author’s own account:
- No customers, users, or adoption metrics.
- No revenue data.
- No product usage statistics.
- No public product launch or marketing activity.
- The demo is limited to a synthetic portal and one developer.
Not evidenced
Competitive Context
The description does not reference competitors directly. However, it implies:
- A space of browser automation tools (e.g., Playwright, Selenium, Cypress).
- Tools that convert recordings into code or tests.
- Focus on reliability over ease-of-use.
Inference It competes with browser automation tools that lack strong integrity guarantees. It may be positioned against generic recorders or low-fidelity test generators, but no direct comparison is made.
Key Risks & Red Flags
Key risks and red flags based on the description:
- The tool is described as a hackathon project with no known traction or commercialization.
- No evidence of real-world usage or feedback from users.
- The demo is limited to synthetic profiles, not arbitrary websites.
- The author states that the P0 proves this workflow on one synthetic profile — implying limited scope.
- No mention of scalability, performance, or integration into CI/CD pipelines.
Inference The tool may be too narrow in scope for broad adoption. Its focus on deterministic integrity and local execution could limit its utility in larger teams or distributed environments.
Diligence Questions To Ask The Founders
- What is the intended use case beyond the synthetic demo?
- How does this differ from existing tools like Playwright's built-in recording or other recorder-based solutions?
- Are there any plans to support real-world applications or integrations with CI/CD systems?
- Has the tool been tested in environments outside of the synthetic portal?
- What is the long-term vision for monetization or product development beyond the hackathon?
- How does the tool handle edge cases like dynamic content, iframes, or complex JavaScript interactions?
- Is there any plan to open-source components or make them available as libraries?
Investment/Partnership Verdict
The project is described as a hackathon submission with no evidence of traction, revenue, or customer adoption. It is technically well-designed and solves a specific problem in browser automation reliability, but lacks commercial viability indicators.
Confidence level Low
Verdict Not ready for investment or partnership at this stage. The tool shows potential for niche use cases but requires further development, testing, and market validation before it can be considered a viable product.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.

