OpenAI 2026 hackathon

Before You Approve

A flight simulator for the human side of AI agents: practice Allow, Ask, or Block before the consequences are real.

Solo project by Tony Lee · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,904 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Before You Approve is an interactive supervision simulator designed to teach users how to make decisions about AI agent actions — Allow, Ask, or Block — before those actions are executed in real systems. It is presented as a flight simulator for human oversight of AI agents, allowing learners to practice decision-making without consequences.

What changed

The author states that they pivoted from an initial concept involving runtime firewalls to an educational tool focused on teaching the skill of human supervision over AI agents. This shift was informed by research using GPT-5.6 and a desire to avoid building another warning banner.

Single most important open question

Is there evidence that this project has moved beyond prototype stage, or whether it has been tested with educators or learners in real-world settings?

Back to contents

What The Product Actually Is

The description states that Before You Approve is an interactive supervision simulator. It presents learners with:

  • An original assignment
  • A proposed tool call (in MCP format)
  • Evidence needed to judge the action: literal effect, target, arguments, reversibility, and provenance

Learners must choose one of three responses:

  • Allow
  • Ask
  • Block

Feedback includes:

  • The decisive evidence used in the decision
  • The smallest fact that would change the decision
  • A replayable event list showing consequences (but no real action is executed)

The system uses:

  • A React and TypeScript frontend
  • JSON fixtures for scenarios
  • A Node.js CLI (bya-trace) to record tool calls via JSON-RPC
  • Deterministic scoring and validation logic

Not evidenced No mention of actual users, customer data, or production use cases. The product is described as a prototype.

Back to contents

Positioning & Claim Evolution

The author claims that the project is:

  • A flight simulator for the human side of AI agents
  • Designed to teach people how to prompt an AI and question its answers, but not how to decide whether an agent should be allowed to act
  • Not another warning banner, but a place to rehearse decisions

They also state that they pivoted from a concept involving runtime firewalls to an educational tool focused on human supervision skills.

Inference The project evolved from a security-focused idea into an educational one, driven by GPT-5.6 research and feedback.

Not evidenced No claims about market positioning, competitive differentiation, or adoption in real-world contexts.

Back to contents

Target Customer & ICP

The description states that the product is intended for:

  • Learners who need to practice AI agent supervision
  • People who want to understand how to supervise AI agents before they act

It is described as a tool for teaching one concrete skill — not abstract AI safety, but specific decision-making in AI agent interactions.

Not evidenced No explicit identification of target personas (e.g., students, educators, enterprise users), or any evidence of who has used it or how it would scale.

Back to contents

Business Model & Pricing Evidence

The description states that:

  • The hosted lesson is credential-free
  • It does not call external models, services, mailboxes, stores, or repositories
  • It never executes displayed actions

There is no mention of monetization, pricing tiers, or any business model beyond the prototype.

Not evidenced No evidence of revenue, customers, or commercialization plans.

Back to contents

Technical & Delivery Signals

The product is built with:

  • Frontend: React, TypeScript, Vite
  • Backend: Node.js CLI (bya-trace)
  • Data Format: JSON fixtures, newline-delimited JSON-RPC (MCP-style)
  • Validation: Strict validator for unknown fields, broken provenance, unsafe labels, etc.
  • Security Features:
    • Deterministic grading (not using the same model that helped draft scenarios)
    • Local policy to withhold or review requests
    • SHA-256-linked receipts
    • Tamper detection

The system is described as a consequence-free simulator, with no real execution of actions.

Not evidenced No evidence of scalability, performance metrics, or deployment in production environments.

Back to contents

Traction & Maturity Signals

The description states:

  • The project contains three four-action cases
  • It includes Allow, Ask, and Block decisions for each case
  • A completion view calculates scores, unsafe approvals, and unnecessary blocks
  • A Progress page is labeled as seeded demonstration data
  • The prototype does not claim measured learning gains

There is no evidence of:

  • Real users or learners
  • Customer feedback or usage metrics
  • Product adoption or retention data
  • Any form of validation beyond internal testing

Not evidenced No traction, user base, or performance data.

Back to contents

Competitive Context

The author states that GPT-5.6 research revealed that their first runtime-firewall concept overlapped with an existing product, which led to a pivot toward education.

Inference There may be competitors in the AI supervision or safety space, but no specific names or market positioning are mentioned.

Not evidenced No mention of direct competitors, market size, or competitive advantages.

Back to contents

Key Risks & Red Flags

  • Prototype-only: The project is described as a prototype with no validated curriculum or classroom testing.
  • No real-world validation: No evidence of educator co-design, pre/post studies, or learning outcome measurement.
  • Unproven scalability: The system is built for simulation and education, not production use.
  • Lack of commercialization plan: No evidence of monetization, distribution, or go-to-market strategy.
  • No external validation: The project does not claim to have been tested with real learners or educators.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific learning outcomes are you trying to measure in a classroom setting?
  2. Have you co-designed drills with educators, and if so, what were the key insights?
  3. How do you plan to validate that practice reduces unsafe approvals without producing blanket blocking?
  4. Are there any plans for localization or accessibility features?
  5. What is your roadmap for moving from prototype to a validated curriculum or product?
  6. Do you have any early adopters or pilot users in educational or enterprise settings?

Back to contents

Investment/Partnership Verdict

Not evidenced No financials, traction, or commercial viability data are provided.

Confidence Level Low This is a self-reported prototype with no evidence of product-market fit, user adoption, or revenue. The author pivoted from a security idea to an educational one, but there is no indication that it has moved beyond the experimental stage.

Inference If the project evolves into a validated educational tool with real-world testing and partnerships, it may have potential. However, as presented, it is not a product ready for investment or partnership.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.