OpenAI 2026 hackathon

VetIOS ProofLoop

Turns verified clinical outcomes into executable AI evals, regression tests, and release gates.

Solo project by John Bruce · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #2,172 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

Company: VetIOS ProofLoop

Self-reported basis: The description is entirely self-reported and unverified, based on a Devpost submission for the OpenAI 2026 hackathon. No third-party corroboration or historical data available.

What it appears to be: A proof-of-concept tool that links AI model inferences to verified clinical outcomes using GPT-5.6 and Codex, enabling outcome-derived regression testing and release gates for AI systems in veterinary medicine.

What changed: The project was built as an extension to an existing VetIOS platform, with a new ProofLoop module added after the OpenAI Build Week submission period began.

Single most important open question: Does this system actually work end-to-end in real-world clinical settings, or is it limited to synthetic demonstrations?

Back to contents

What The Product Actually Is

The description states that ProofLoop:

  • Links AI model inferences to de-identified outcome evidence and human confirmation.
  • Produces a hash-addressed Outcome Receipt with source lineage and review state.
  • Uses GPT-5.6 through the Responses API to classify failures, identify affected slices, and emit structured evaluation specifications.
  • Invokes Codex in the target repository to create regression fixtures, run suites, and return auditable patches.
  • Blocks model promotion until outcome-derived gates pass.

Inference: The system is described as a hybrid AI/developer tooling platform that bridges AI observability with outcome verification. It is not a standalone observability platform but rather an extension or complement to existing systems.

Not evidenced: No details on how the system integrates with existing clinical workflows, model deployment pipelines, or evaluation platforms beyond its own documentation and demo.

Back to contents

Positioning & Claim Evolution

The description states that:

  • AI observability shows what a model said and which tools it called.
  • It usually cannot prove what happened in the real world afterward.
  • ProofLoop is designed to export tests to existing evaluation systems rather than replace them.
  • The system targets veterinary medicine as a first wedge due to fragmented evidence.

Inference: The positioning is that ProofLoop is a ground-truth verification layer for AI systems, particularly in clinical environments where outcomes matter. It is not a replacement for observability tools but a complementary outcome-driven testing mechanism.

Not evidenced: No claims about market traction, adoption, or competitive differentiation beyond its own self-description.

Back to contents

Target Customer & ICP

The description states:

  • The system targets veterinary medicine as the first wedge.
  • Evidence is fragmented across clinics, labs, imaging systems, and follow-ups.
  • Wrong or overconfident outputs can be consequential.

Inference: The initial customer segment is veterinary healthcare providers, particularly those using AI in diagnostic or treatment workflows. The ICP is likely AI developers or clinical teams in veterinary medicine who need outcome verification for AI models.

Not evidenced: No evidence of specific customers, use cases, or adoption beyond the demo and self-reporting.

Back to contents

Business Model & Pricing Evidence

The description states:

  • ProofLoop is designed to export tests to existing evaluation systems.
  • It does not replace horizontal observability platforms but supplies outcome-verified ground truth they do not own.

Inference: The business model appears to be platform integration or SaaS-as-a-service, where ProofLoop provides a verification layer that can be plugged into existing AI development and deployment pipelines.

Not evidenced: No pricing, revenue models, or monetization strategy are described. No evidence of paid customers or subscriptions.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with: codex, codex-sdk, gpt-5.6, next.js, openai-responses-api, programmatic-tool-calling, supabase, typescript.
  • Uses GPT-5.6 for reasoning across heterogeneous evidence.
  • Codex is used for repository-aware execution and test generation.
  • Programmatic Tool Calling enables bounded control flow.

Inference: The system uses a hybrid AI + developer tooling stack, combining LLMs with code generation and execution capabilities to automate outcome verification and regression testing.

Not evidenced: No evidence of production-grade infrastructure, scalability, or integration with real-world systems beyond the demo.

Back to contents

Traction & Maturity Signals

The description states:

  • VetIOS existed before Build Week.
  • ProofLoop is a new extension built after July 13, 2026.
  • The repository identifies pre-existing baseline and new commits.

Inference: This is a proof-of-concept or prototype, not a mature product. It was developed during a hackathon and has no evidence of real-world deployment or adoption.

Not evidenced: No revenue, customers, usage metrics, or product maturity beyond the demo.

Back to contents

Competitive Context

The description states:

  • Horizontal observability platforms capture traces.
  • ProofLoop supplies outcome-verified ground truth they do not own.

Inference: The competitive context is AI observability + outcome verification, with competitors likely being platforms like LangSmith, LlamaIndex, or similar AI monitoring tools. ProofLoop is positioned as a complementary tool to these systems.

Not evidenced: No mention of existing competitors, market positioning, or differentiation strategies beyond its own claims.

Back to contents

Key Risks & Red Flags

  • The system is described as a demo-only prototype, not a production-ready product.
  • It relies on GPT-5.6 and Codex, which are not publicly available or standardized tools.
  • No evidence of real-world clinical integration, safety compliance, or regulatory alignment.
  • The project is self-reported only with no third-party validation.

Inference: The biggest risk is that this system has no demonstrated real-world utility or scalability, and may be limited to synthetic or controlled environments.

Back to contents

Diligence Questions To Ask The Founders

  1. What are the actual clinical workflows where this would be applied?
  2. How does it handle edge cases in outcome data (e.g., missing, conflicting, or delayed data)?
  3. Is there any integration with real-world EHRs or lab systems?
  4. What is the plan for regulatory compliance and safety validation?
  5. How does it scale beyond a single veterinary clinic or use case?

Back to contents

Investment/Partnership Verdict

Not evidenced: No financials, traction, or investment history to evaluate.

Inference: This project is a preliminary prototype with potential in a niche market (veterinary AI). It lacks evidence of commercial viability, real-world adoption, or scalability. It may be a seed-stage idea worth exploring further if the team can demonstrate real-world utility and integration capabilities.

Confidence level: Low — based entirely on self-reported claims and no external validation.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.