OpenAI 2026 hackathon

NEMA Boundary Runtime

Turn conversational signals into deterministic, executable AI safety policies—with inspectable traces, response controls, and post- generation verification.

Solo project by setsuki kagio · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,508 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

NEMA Boundary Runtime is a self-reported project that claims to offer a deterministic runtime layer for AI safety policies in conversational systems. It separates conversational-state inference from policy execution, aiming to provide inspectable traces and control over AI responses through a policy-driven interface.

What changed

The description indicates this is a prototype submitted to the OpenAI 2026 hackathon. It does not describe any prior version or evolution beyond its initial development phase.

Single most important open question

Is there evidence of real-world integration, usage, or traction with actual customers or developers beyond the hackathon sandbox?

Back to contents

What The Product Actually Is

The description states that NEMA Boundary Runtime is a policy layer designed to separate conversational-state inference from policy execution. It uses a deterministic runtime, where validated ControlState enters and policies evaluate that state, emitting response-control directives.

It includes:

  • A browser interface showing inferred signals, fired policies, control directives, baseline vs controlled responses.
  • Token-level difference views.
  • Post-generation directive verification.
  • An exact policy execution path.
  • Temporary overrides with schema validation and size bounds.
  • Integration with GPT-5.6 via OpenRouter for inference.

The system is built using:

  • FastAPI
  • Pydantic
  • Docker
  • Playwright
  • Pytest

It was developed as part of a 40-case offline development contract and demonstrated in a 12-call integration proof using GPT-5.6 Sol through OpenRouter.

Inference The product is described as a runtime system for AI safety controls, not a model or service per se.

Back to contents

Positioning & Claim Evolution

The author states that NEMA aims to:

  • Turn conversational signals into deterministic, executable AI safety policies.
  • Provide inspectable traces and response controls.
  • Offer post-generation verification.
  • Allow temporary policy overrides without persisting changes.

It positions itself as a policy layer that operates independently of the model, rather than asking the model to explain its own behavior.

The claim is that:

  • The system makes runtime-generated explanations instead of relying on model introspection.
  • It supports replayable thresholds and traces.
  • It separates conformance testing from model performance evidence.

Inference This is a self-contained safety control layer, not a general-purpose AI tool or platform. The positioning emphasizes determinism, inspectability, and policy-driven control over LLM outputs.

Back to contents

Target Customer & ICP

The description does not name specific customers or target personas.

However, it implies:

  • Developers working with conversational AI systems.
  • Teams seeking to apply AI safety policies in controlled environments.
  • Users who want inspectable, replayable, and schema-validated control over AI-generated responses.

It also suggests a focus on developers or engineers who are building or integrating LLMs into applications, particularly those concerned with AI safety and compliance.

Inference The ICP likely includes developers or engineering teams in enterprise or research contexts where AI safety is a concern, but no explicit customer segments are defined.

Back to contents

Business Model & Pricing Evidence

There is no evidence of pricing, monetization, or business model in the description.

The project is presented as a hackathon prototype, not a commercial product.

Inference No business model or pricing is evident. The system appears to be built for demonstration and testing purposes only.

Back to contents

Technical & Delivery Signals

Key technical elements:

  • Uses FastAPI and Pydantic for strict validation.
  • Implements a priority-ordered deterministic policy runtime.
  • Includes a responsive browser interface.
  • Supports replayable thresholds, traces, and schema-validated overrides.
  • Integrates with GPT-5.6 via OpenRouter.
  • Built with Docker, Playwright, Pytest.

The system is described as:

  • Offline-first (40-case contract).
  • Public sandbox with no credentials required.
  • Designed for inspectability and reproducibility.

Inference The technical stack is consistent with a developer-focused prototype, likely built for internal or demonstration use. No evidence of production-grade infrastructure or scalability.

Back to contents

Traction & Maturity Signals

The project is described as:

  • A hackathon submission.
  • A 12-call integration proof using GPT-5.6.
  • A 40-case offline development contract.
  • A public sandbox with no credential requirement.

There is no mention of:

  • Customers
  • Revenue
  • Usage metrics
  • Product adoption
  • Production deployment

Inference No traction or maturity signals are evident beyond the prototype stage and hackathon demonstration.

Back to contents

Competitive Context

The description does not reference competitors or similar products.

It implies a niche in AI safety controls, particularly around:

  • Deterministic policy execution.
  • Post-generation verification.
  • Inspectable AI behavior.

Inference The space is likely related to AI governance, LLM safety layers, and explainability tools, but no direct competitive analysis is provided.

Back to contents

Key Risks & Red Flags

  • Prototype-only status: No evidence of real-world usage or integration beyond a hackathon.
  • No revenue or customer data: The system is not described as commercialized or in production.
  • Limited scope: The 40-case contract and 12-call proof are not representative of generalization or scalability.
  • Self-reported only: No independent verification of claims, performance, or safety guarantees.
  • No clear path to monetization or product-market fit.

Inference The project is at a very early stage. It lacks commercial traction, and its value proposition remains unproven in real-world settings.

Back to contents

Diligence Questions To Ask The Founders

  1. What are the specific use cases where this system would be deployed?
  2. How does it integrate with existing LLM platforms or APIs?
  3. Has it been tested beyond the 40-case offline contract and 12-call proof?
  4. Are there plans to move beyond prototype status, and if so, what’s the roadmap?
  5. What are the key assumptions about policy design and control that underpin this system?
  6. How does it handle edge cases or unexpected conversational inputs?
  7. Is there any internal testing or validation of safety claims beyond the sandbox?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The project is described as a hackathon prototype, with no evidence of:

  • Revenue
  • Customers
  • Product-market fit
  • Commercial traction
  • Scalable deployment

It is presented as a conceptual safety control layer, not a product or service.

Confidence Low. The description is self-reported and unverified, and there is no indication that the system has moved beyond experimental or demonstration status.

Inference This is an early-stage idea with potential but no demonstrated commercial viability or traction. It may be worth exploring further if the founders are planning to build out a product, but it does not yet constitute a viable investment or partnership opportunity based on the supplied information.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.