OpenAI 2026 hackathon

Visual Contract

Compile a production brief into a machine-checkable spec, then audit images against it. GPT-5.6 observes and reports typed evidence; deterministic code issues the verdict. Same inputs, same verdict.

Solo project by aomizuki0307 Tomoyuki · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #7,577 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Visual Contract is a self-reported tool that compiles production briefs into machine-checkable image specifications using GPT-5.6, then audits images against those specs with deterministic checks and visual model observations. It offers a CLI and small web UI for use in CI/CD pipelines.

What changed

The author states they built this to automate image review processes that were previously manual and error-prone, inspired by how testing works in code. The tool is described as a four-step pipeline: brief input → contract compilation → audit → verdict issuance.

Single most important open question

Is there any evidence of real-world usage or adoption beyond the author’s own demo? The description contains no claims about customers, revenue, or product-market fit beyond self-reported performance on test sets.

Back to contents

What The Product Actually Is

The description states that Visual Contract is a four-step pipeline with:

  1. A CLI and small web UI.
  2. Input of a production brief.
  3. Compilation into a YAML-based "Visual Contract" using GPT-5.6, containing up to 15 assertions across 10 types (e.g., exact text, object count, spatial relations).
  4. Editing and re-validation of the contract.
  5. An audit process where deterministic checks are done in Python and visual checks go to GPT-5.6 Vision.
  6. A rule engine issues a verdict per assertion: pass, fail, or uncertain.
  7. Output includes JSON, JUnit XML, HTML reports, and annotated images.
  8. Exit codes plug into CI/CD pipelines.

The system is described as fully replayable without API keys, using recorded fixtures to simulate model responses.

Inference This appears to be a proof-of-concept or prototype tool for automating image validation in creative workflows, likely targeting teams producing AI-generated assets such as banners, mockups, or game assets.

Back to contents

Positioning & Claim Evolution

The author claims the product is inspired by software testing practices, aiming to bring deterministic verification to image creation and review. It positions itself as a way to avoid human error ("squinting at a screen") and replace subjective feedback with structured checks.

Key Claims

  • "Tests solved this for code decades ago. I wanted the same thing for images."
  • "Same inputs, same verdict" — implying deterministic behavior.
  • "GPT-5.6 observes and reports typed evidence; deterministic code issues the verdict."

Inference The positioning is early-stage, focused on solving a niche problem in AI-generated content workflows. It does not claim to be a commercial product or platform but rather a tool for internal use or experimentation.

Back to contents

Target Customer & ICP

The description states that the author ships AI-generated images (e.g., ad banners, product mockups, game assets) and was frustrated by manual review processes.

Inference

The likely target customer is:

  • Creative teams using AI tools to generate visual content.
  • Product or design teams needing consistent output quality.
  • Developers or engineers working in CI/CD environments who want to automate image validation.

However, no explicit ICP or customer segmentation is provided beyond the author’s personal use case.

Back to contents

Business Model & Pricing Evidence

There is no evidence of a business model or pricing structure. The description does not mention:

  • Revenue streams
  • Subscription plans
  • Licensing fees
  • Paid features
  • Monetization strategy

Inference The tool appears to be non-commercial, possibly an open-source prototype or hackathon submission, with no indication of monetization.

Back to contents

Technical & Delivery Signals

The project is built using:

  • Codex for scaffolding and implementation.
  • GPT-5.6 via OpenAI Responses API with Structured Outputs.
  • FastAPI, Pydantic, OpenCV, Pillow, Typer, Jinja, pytest, GitHub Actions.
  • A CLI and small web UI.

The author reports:

  • Use of structured outputs to enforce type safety.
  • Replay mode for testing without API calls.
  • Security hardening through multiple rounds of review.
  • Prompt injection resistance tested via adversarial examples.
  • Holdout test sets to validate honesty of evaluation.

Inference The technical stack and delivery approach suggest a highly technical prototype, likely built by one person, with strong attention to correctness and reproducibility. It is not yet a scalable SaaS offering.

Back to contents

Traction & Maturity Signals

There is no evidence of:

  • Customers or users
  • Revenue or monetization
  • Product-market fit
  • Adoption beyond the author’s own use case
  • Any form of product release or distribution outside GitHub

The project is described as a hackathon submission, and the only “demo” is a GitHub repo with replayable fixtures.

Inference This is an early-stage prototype, likely in pre-product-market fit or pre-commercialization phase. No traction or adoption data are evident.

Back to contents

Competitive Context

The description does not mention any competitors or direct market comparisons.

Inference No competitive context is provided. The tool may be addressing a niche within AI-generated image validation, but there is no evidence of existing solutions in the space.

Back to contents

Key Risks & Red Flags

  • Single-person team: The project has only one contributor, which raises questions about scalability and long-term maintenance.
  • No commercial traction or customers: The tool is described as a personal solution, not a product for sale.
  • Self-reported performance data: All metrics (e.g., 43/43 assertions correct) are from internal testing, with no external validation.
  • No pricing or monetization strategy: No indication of how the tool would be sold or used commercially.
  • Unverified model outputs: The system relies heavily on GPT-5.6, which is not independently verified for consistency or reliability.

Back to contents

Diligence Questions To Ask The Founders

  1. What is your actual workflow with AI-generated images? Is this a personal tool or part of a larger team process?
  2. Have you tested the system with real-world inputs from multiple users or teams?
  3. How do you plan to scale beyond a single developer’s use case?
  4. Are there any plans for monetization or commercial deployment?
  5. What are the limitations of GPT-5.6 in this context, and how do you handle model hallucinations or inconsistencies?
  6. Is there any intention to open-source or release this as a product?

Back to contents

Investment/Partnership Verdict

Not evidenced.

There is no evidence of:

  • Revenue
  • Customers
  • Product-market fit
  • Commercial traction
  • Any form of investment interest or partnership opportunity

The project is described as a hackathon submission, built by one person, with no indication of commercial viability or scalability.

Inference This is a pre-product-market-fit prototype. It may be of interest for early-stage experimentation or strategic partnerships if the team intends to build out a product, but it does not currently meet criteria for investment or partnership consideration based on the evidence provided.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.