OpenAI 2026 hackathon

BOUNDARY: AI That Asks Before It Acts

A policy and approval layer for AI agents that turns plain-English company rules into enforceable actions—allow, redact, route privately, require approval, or block—with a complete audit trail.

Team of 2 · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #723 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

BOUNDARY is a self-reported policy and approval layer for AI agents that translates plain-English company rules into enforceable actions (allow, redact, route privately, require approval, or block), with an audit trail. It was built as a hackathon submission for the OpenAI 2026 hackathon.

What changed

The project is presented as a control layer between AI agent intent and action execution, using deterministic code to enforce decisions based on structured policy rules interpreted from natural language by GPT-5.6. The system includes adversarial testing capabilities and a simulated workflow for demonstration.

Single most important open question

Is there evidence of real-world adoption or traction beyond the hackathon demo? The description states no revenue, customers, or usage data exist outside of synthetic examples.

Back to contents

What The Product Actually Is

The description states that BOUNDARY is a policy and approval layer for AI agents. It converts plain-English company rules into structured enforcement decisions using deterministic code. Each proposed action receives one of five outcomes: Allow, Redact and allow, Route privately, Require approval, or Block.

It also includes an adversarial policy-testing workflow where GPT-5.6 proposes bypass attempts that are reviewed and tested against the policy engine.

The system is built with:

  • Next.js
  • TypeScript
  • Zod for schemas
  • OpenAI Responses API
  • Redis (Upstash)
  • React, Tailwind CSS, Vitest

All tools, refunds, emails, and actions in the demo are synthetic or simulated.

Inference The product appears to be a policy enforcement engine, not an AI agent itself. It functions as a middleware layer that evaluates AI-generated actions against pre-defined rules before allowing them to proceed.

Back to contents

Positioning & Claim Evolution

The description states that BOUNDARY was built to place a deterministic control layer between an AI agent’s intention and its real-world action. It positions itself as a solution to the problem of prompt-only restrictions, which it claims are insufficient for enforcing organizational policies.

It also claims that:

  • AI agents are moving beyond answering questions into performing actions like reading records, calling tools, updating systems, and communicating with customers.
  • Prompt-based controls alone are not enforceable.
  • The system allows humans to confirm interpreted policy while deterministic code enforces decisions.

Inference BOUNDARY positions itself as a policy enforcement middleware, aimed at organizations seeking to regulate AI agent behavior in enterprise settings. It evolves from the idea that AI safety cannot rely only on better prompts, but requires structured boundaries.

Back to contents

Target Customer & ICP

The description states that BOUNDARY was developed with a focus on customer-support workflows involving personal information, refunds, private transcripts, external emails, and destructive tool requests.

It also mentions that future work would include protecting:

  • Finance and refund agents
  • HR assistants
  • Healthcare workflows
  • Sales automation
  • Internal knowledge agents
  • Developer and deployment agents

However, there is no evidence of specific customer segments or personas identified beyond these use cases.

Inference The target ICP appears to be enterprise organizations using AI agents, particularly those in support, finance, HR, and internal operations, where compliance and control are critical.

Back to contents

Business Model & Pricing Evidence

There is no evidence in the description of a business model or pricing structure. The project is described as a hackathon submission with no mention of monetization, licensing, or sales channels.

Inference No commercial business model or pricing information was provided.

Back to contents

Technical & Delivery Signals

The system is built using:

  • Next.js
  • TypeScript
  • Zod for schema validation
  • OpenAI GPT-5.6 via official API
  • Redis (Upstash)
  • React, Tailwind CSS, Vitest
  • Codex as primary development environment

Key technical features include:

  • Deterministic enforcement engine
  • Human confirmation and approval workflows
  • Redaction and private routing transformations
  • Simulated tools and safe audit events
  • Adversarial policy testing
  • Session persistence using Upstash Redis
  • Request throttling, error handling, health checks
  • 89 automated tests

The system uses GPT-5.6 only for bounded language tasks:

  • Interpreting plain-English policy into structured proposals
  • Producing non-authoritative adversarial suggestions

GPT-5.6 cannot approve actions or invoke tools directly.

Inference The architecture suggests a modular, deterministic enforcement engine, with GPT used only for interpretation and adversarial testing, not for decision-making or execution.

Back to contents

Traction & Maturity Signals

The project is described as a hackathon submission. All tools, refunds, emails, and actions in the demo are synthetic or simulated.

There is no evidence of:

  • Revenue
  • Customers
  • Usage metrics
  • Product-market fit
  • Real-world deployment
  • Production systems

Inference The product exists only in a demo or prototype form, with no evidence of traction or real-world adoption.

Back to contents

Competitive Context

The description does not mention any competitors. It is unclear whether similar solutions exist in the market for AI agent policy enforcement or control layers.

Inference No competitive landscape was provided, and there is no evidence of existing products addressing this space.

Back to contents

Key Risks & Red Flags

  • No real-world adoption: The project is a hackathon demo with no evidence of traction.
  • Unverified claims: All descriptions are self-reported and unverified.
  • Limited scope: The system only demonstrates a customer-support workflow, with no indication of broader applicability or scalability.
  • Dependency on GPT-5.6: While the system uses GPT for interpretation, it does not rely on it for enforcement, which is good; however, this raises questions about how well it scales beyond the demo.
  • No business model: No indication of monetization strategy or customer acquisition plan.

Inference The project lacks commercial viability indicators and may be a conceptual prototype, not a product ready for market.

Back to contents

Diligence Questions To Ask The Founders

  1. What is your definition of “deterministic enforcement” in practice? How do you ensure consistency across different policy interpretations?
  2. Has the adversarial testing workflow been validated with real-world scenarios beyond the demo?
  3. Are there any plans to integrate with existing enterprise systems (e.g., CRM, ERP, identity providers)?
  4. What are the key assumptions about how organizations will adopt this product?
  5. How do you plan to scale beyond a single use case like customer support?
  6. Is there any internal or external testing of the system’s ability to detect bypass attempts in real-world conditions?

Back to contents

Investment/Partnership Verdict

The description states that BOUNDARY is a hackathon submission, and no evidence exists of revenue, customers, or traction beyond synthetic examples.

There is no indication of:

  • Commercial readiness
  • Product-market fit
  • Scalability
  • Go-to-market strategy
  • Competitiveness in the AI safety/control space

Inference This project is currently a conceptual prototype, not a viable investment or partnership opportunity. It may be a valuable idea for future development, but it lacks evidence of commercial viability or traction.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.