OpenAI 2026 hackathon

TEKMERION: Proof Surface Compiler

TEKMERION reveals when “approval” and “success” logs cannot prove an agent executed what a human authorized, then compiles the missing evidence into a reviewable patch and deterministic check.

Solo project by Fer G · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #7,177 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

TEKMERION: Proof Surface Compiler is a developer tool that claims to compile structured evidence plans for agentic workflows using GPT-5.6 and Codex. It aims to detect when an action executed by an agent does not match what was human-approved, by generating deterministic proof surfaces from operational guarantees.

What changed

The project description presents a self-contained prototype built for the OpenAI 2026 hackathon. It includes a monorepo implementation with scenario generation, GPT-5.6-based claim compilation, and Codex-assisted patch materialization. The MVP validates one proof slice (action equivalence) but does not yet support broader operational claims or production deployment.

Single most important open question

Is there any evidence of traction, revenue, customer adoption, or real-world usage beyond the hackathon prototype?

Back to contents

What The Product Actually Is

The description states that TEKMERION is a Proof Surface Compiler for agentic workflows, designed to identify and compile evidence required to validate operational guarantees. It uses:

  • GPT-5.6 to generate structured candidate evidence plans;
  • Codex to materialize instrumentation and verification patches;
  • Deterministic software to evaluate the resulting evidence.

It operates on bounded workflows, simulates compliant and violating executions, and maps evidence requirements to repository symbols and execution boundaries.

Inference The product is a developer tool for auditing agentic systems in production-like environments. It does not appear to be a commercial SaaS offering or a general-purpose logging system.

Back to contents

Positioning & Claim Evolution

The author states that TEKMERION began with the thesis:

"Many operational guarantees are not only violated. They are untestable because the system never preserved the evidence required to evaluate them."

It positions itself as a solution to the problem of auditability in agentic systems, where traditional telemetry logs prove activity but not compliance.

The product claims to:

  • Not log everything;
  • Identify specific evidence needed for a given operational claim;
  • Use GPT-5.6 to compile structured plans;
  • Use Codex to materialize patches;
  • Evaluate evidence deterministically.

Inference The positioning is that of a developer tool for compliance and auditability, not a general-purpose observability platform or AI assistant.

Back to contents

Target Customer & ICP

The description states that TEKMERION targets agentic workflows in production environments, where agents are allowed to perform actions after human approval but lack sufficient evidence to prove compliance.

It is built for developers working with:

  • Infrastructure-as-code;
  • Deployment automation;
  • Agent-based systems;
  • Operational guarantees requiring auditability.

Inference The ICP likely includes developers or DevOps engineers in organizations using AI agents for production tasks, such as infrastructure deployment or content publishing.

Back to contents

Business Model & Pricing Evidence

Not evidenced. The description does not mention any pricing model, monetization strategy, or business model.

Back to contents

Technical & Delivery Signals

The project is described as a TypeScript monorepo with packages for:

  • Workflow simulation;
  • Telemetry normalization;
  • Deterministic proof predicates;
  • Scenario generation;
  • GPT-5.6 claim compilation;
  • Web-based judge replay.

It uses:

  • OpenAI Responses API with structured output;
  • Codex for patch materialization;
  • GPT-5.6 to map obligations to concrete repository symbols;
  • Deterministic software to validate and enforce policies.

Inference The delivery is a developer tool, likely intended for integration into CI/CD or deployment pipelines. It is not a hosted service.

Back to contents

Traction & Maturity Signals

Not evidenced. No mention of:

  • Customers;
  • Revenue;
  • Usage metrics;
  • Product adoption;
  • Market traction.

The project is described as an MVP built for a hackathon and does not claim to be in production or used by any organization.

Back to contents

Competitive Context

Not evidenced. The description does not reference competitors, market size, or existing solutions in the space of agentic workflow auditing or operational guarantees.

Back to contents

Key Risks & Red Flags

  • No traction or commercialization evidence: The project is a hackathon prototype with no signs of real-world usage.
  • Unproven scalability: The MVP only validates one proof slice; broader implementation remains future work.
  • Dependency on GPT-5.6 and Codex: These tools are not yet widely available or stable for production use.
  • Limited scope: The system is described as a prototype, with the full “proof surface” remaining unvalidated.
  • Self-reported maturity: No external validation or third-party evidence of product quality or performance.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific operational guarantees are you targeting in production environments?
  2. How do you plan to scale beyond the MVP’s single proof slice?
  3. Have you validated your approach with any real-world agents or workflows?
  4. What is the roadmap for moving from prototype to a commercial product?
  5. Are there any existing customers or pilot programs?
  6. How do you handle privacy and data governance in the evidence collection process?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The description presents a self-reported hackathon prototype with no evidence of traction, revenue, or customer adoption. The project is not yet a commercial product, nor does it appear to be in a position for investment or partnership at this stage.

Confidence Low. This analysis is based entirely on the self-reported project description and lacks any external validation or market signals.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.