Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #6,158 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
ProxyBreak is a self-reported tool that uses GPT-5.6 and deterministic simulation to test whether incentive KPIs can be "gamed" in a way that produces high proxy scores but poor mission outcomes. The product is presented as a browser-based application built with Next.js, TypeScript, and Vercel, using OpenAI's Responses API for structured policy interpretation and strategy generation. It allows users to define a mission and proposed KPI, then uses GPT-5.6 to compile the KPI into a bounded policy, generate strategies, and simulate outcomes under both proxy and mission conditions.
The system claims to produce "verified counterexamples" — concrete examples where a strategy follows the rules, earns high scores, but fails to achieve the intended outcome. These are then used to propose repairs to the original policy. The demonstration focuses on synthetic customer support, showing that a strategy can score 84 points under a flawed KPI while resolving zero tickets.
Key commercial signals:
- No evidence of revenue, customers or adoption.
- Product is described as a prototype for a hackathon submission.
- Founders are two individuals (Jonathan Aguilar, Swati Saxena).
- The tool is presented as a unit-testing-like framework for incentives.
- It requires deterministic replay and cryptographic trace digests to validate findings.
The single most important open question: Is this product intended for commercial use or is it purely a proof-of-concept? The description does not clarify whether ProxyBreak has moved beyond the hackathon prototype stage, nor does it indicate any traction, pricing model, or target market beyond its demonstration context.
What The Product Actually Is
The description states that ProxyBreak is a browser-based application built with Next.js and TypeScript, deployed on Vercel. It uses GPT-5.6 via OpenAI's Responses API to interpret natural-language KPIs into bounded policies, generate strategies, and simulate outcomes under both proxy and mission conditions.
Key technical components:
- Uses OpenAI’s structured outputs for policy compilation
- Employs a single run_strategy tool with bounded declarative DSL
- Implements deterministic simulation using integer time, stable ordering, and canonical serialization
- Applies SHA-256 trace digests for verification
- Separates strategy execution from event settlement
The system claims to:
- Compile natural-language KPIs into strict, bounded policies
- Generate strategies via GPT-5.6 that are then validated by application code
- Simulate proxy scores and mission outcomes independently
- Produce verified counterexamples with trace digests
- Propose policy repairs based on evidence
The product is described as a "unit testing for incentives" — a framework to test whether KPIs can be gamed without actually deploying them in real-world settings.
Positioning & Claim Evolution
The description positions ProxyBreak as a tool that helps organizations avoid designing incentive systems that can be easily gamed. It frames the core problem as: “Can someone follow the written rules, win the score, and still lose the mission?”
Claims:
- The product enables testing of KPIs before deployment
- It uses GPT-5.6 to search for concrete counterexamples
- It separates proxy performance from actual mission outcomes
- It provides deterministic verification through replay and trace digests
The positioning evolved from a hackathon prototype to a tool that demonstrates how incentive design can be tested using AI and simulation, with the goal of preventing flawed metrics from being implemented in real systems.
Target Customer & ICP
Not evidenced. The description does not state who the intended users or customers are beyond its demonstration context. No evidence of target personas, use cases, or customer segments is provided.
Business Model & Pricing Evidence
Not evidenced. There is no mention of pricing, monetization strategy, or business model in the self-reported description.
Technical & Delivery Signals
The product is built with:
- Next.js App Router
- TypeScript
- Vercel deployment
- OpenAI Responses API
- React, Tailwind CSS, Zod schemas
- Playwright for end-to-end testing
- ESLint and strict TypeScript checking
Key technical features:
- Bounded declarative DSL for strategies
- Deterministic simulation with integer time
- SHA-256 trace digests for verification
- Strict schema validation for all inputs and outputs
- Independent replay verification
- Canonical serialization of events
- Structured Outputs for policy interpretation
- Function-calling tooling
The system separates:
- GPT-5.6's creative role in proposing strategies and policies
- Application code's role in execution, measurement, qualification, and proof
Traction & Maturity Signals
Not evidenced. No data on users, customers, revenue, or adoption is provided.
Competitive Context
Not evidenced. The description does not mention any competitors or market context beyond its own demonstration.
Key Risks & Red Flags
- The product is described as a hackathon submission with no evidence of commercial traction.
- GPT-5.6 is used for creative search but not for authoritative scoring or certification.
- No evidence of real-world application or integration into existing systems.
- The system's maturity is unclear — it appears to be a prototype rather than a production-ready tool.
- There is no indication of how the tool would scale beyond its synthetic customer-support demonstration.
Diligence Questions To Ask The Founders
- Is this product intended for commercial use or is it purely a proof-of-concept?
- What are the specific use cases and target industries for ProxyBreak?
- How does the system handle edge cases or unexpected inputs from users?
- Are there plans to expand beyond the synthetic customer-support arena?
- What is the roadmap for scaling the tool and integrating with real-world incentive systems?
- How do you plan to validate that the simulated outcomes reflect real-world behavior?
Investment/Partnership Verdict
Not evidenced. No information on valuation, funding rounds, or investment interest is available from the self-reported description. The product appears to be a hackathon prototype without evidence of commercial viability or traction.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.
