OpenAI 2026 hackathon

Classroom Mirror

Classroom Mirror is a preflight simulator for lesson plans, using GPT-5.6 agents and evidence-grounded critique to catch hidden learning barriers before students encounter them.

Team of 2 · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,286 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Classroom Mirror is a preflight simulator for lesson plans, using GPT-5.6 agents and evidence-grounded critique to catch hidden learning barriers before students encounter them. The product is described as an instructional review tool that helps teachers identify potential issues in their lesson plans through AI-assisted analysis.

What changed

The project evolved from an initial concept involving simulated learner personas to a system based on bounded instructional-review lenses, skeptical verification, and evidence-led findings. This redesign was driven by concerns over credibility and the desire to avoid pretending that AI can accurately predict individual student behavior.

Single most important open question — the commercial due-diligence read

Is there a viable market need for an AI-powered preflight check for lesson plans, and does the current prototype demonstrate sufficient utility or traction to warrant further investment or partnership?

Back to contents

What The Product Actually Is

The description states that Classroom Mirror is a tool that allows teachers to submit a lesson plan and receive feedback from GPT-5.6 agents across four instructional lenses:

  1. Conceptual clarity
  2. Academic language
  3. Attention and feedback
  4. Transfer and challenge

Each finding must cite an exact passage from the submitted lesson. A skeptical verification stage challenges these findings, labeling them as verified, qualified, or rejected.

The system uses:

  • GPT-5.6 via OpenAI API
  • Multi-agent architecture for parallel discovery (one root agent and three subagents)
  • Skeptical verification via direct GPT-5.6 call
  • Strict JSON schema outputs
  • Substring validation to ensure evidence quotes are literal substrings of the lesson

It is built with:

  • Next.js, React, TypeScript, Node.js
  • Deterministic no-key demo for judges
  • Live GPT-5.6 analysis for local testing
  • Adversarial benchmark harness

Not evidenced: No mention of actual deployment, user base, or commercial usage beyond a hackathon submission.

Back to contents

Positioning & Claim Evolution

The description states that the project was inspired by the problem of teachers discovering ambiguity and misconception-producing examples only after a lesson is underway. The original idea involved simulated learner personas but was later redesigned to avoid pretending AI could predict individual learners.

Key claims:

  • "Classroom Mirror helps teachers find the barrier before a learner does."
  • "AI surfaces possibilities; teachers make the decisions."
  • "Trustworthy educational AI is less about generating impressive content and more about showing evidence, exposing uncertainty, and preserving teacher agency."

Evolution:

  • Started with simulated learner personas
  • Shifted to bounded instructional-review lenses
  • Introduced skeptical verification and grounding checks
  • Emphasized teacher control over final decisions

Inference: The evolution reflects a move from a generative AI tool toward an evaluative one, focused on transparency and accountability.

Back to contents

Target Customer & ICP

The description states that the target user is teachers who plan lessons. It does not specify grade levels or subject areas beyond general education.

ICP inferred:

  • Educators planning lesson plans
  • Teachers seeking feedback on instructional design
  • Likely in K–12 or higher education contexts

Not evidenced: No data on specific demographics, geographic scope, or institutional type (e.g., public vs. private schools).

Back to contents

Business Model & Pricing Evidence

The description does not provide any information about pricing, monetization strategy, or business model.

Not evidenced: No mention of revenue streams, subscription tiers, licensing models, or customer acquisition costs.

Back to contents

Technical & Delivery Signals

The system uses:

  • GPT-5.6 via OpenAI API
  • Multi-agent architecture with root and subagents
  • Skeptical verification process
  • Strict structured output (JSON schema)
  • Substring validation for evidence grounding
  • Built using Next.js, React, TypeScript, Node.js

Key technical features:

  • Deterministic no-key demo for judges
  • Live GPT-5.6 analysis for local testing
  • Adversarial benchmark harness
  • Contract, safety, accessibility, and production-build tests

Not evidenced: No information on scalability, infrastructure, or API usage limits.

Back to contents

Traction & Maturity Signals

The description states that the team:

  • Built a live GPT-5.6 Sol benchmark run
  • Detected all 12 planted instructional risks
  • Grounded all 28 findings in exact lesson quotes
  • Achieved 100% planted-risk recall
  • Achieved 100% exact-quote grounding

They also mention:

  • Six benchmark lessons with 12 deliberately planted failures
  • Evaluation harness refuses to score deterministic demo data
  • Repository includes benchmark cases and saved results for reproducibility

Not evidenced: No real-world usage, customer feedback, or adoption metrics beyond the hackathon submission.

Back to contents

Competitive Context

The description does not provide any information about competitors or competitive positioning.

Not evidenced: No mention of existing tools in the education AI space, nor how Classroom Mirror differentiates from them.

Back to contents

Key Risks & Red Flags

  • Unproven market demand: The product is described as a hackathon submission with no evidence of real-world traction or adoption.
  • Over-reliance on AI tooling: Heavy dependence on GPT-5.6 and OpenAI APIs may create dependency risks.
  • Limited scope: The system is currently limited to four instructional lenses; lack of expansion plans is concerning.
  • No commercial viability: No evidence of pricing, monetization, or business model.
  • Self-reported performance: All performance claims are based on internal benchmarks, not external validation.

Inference: Without real-world usage or feedback, the product may not meet actual teacher needs.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific pain points in lesson planning did you observe that led to this idea?
  2. How do you plan to validate whether teachers actually use and value the feedback provided?
  3. Are there any institutional partnerships or pilot programs with educators already underway?
  4. What are your plans for expanding beyond the current four instructional lenses?
  5. How will you ensure long-term sustainability given reliance on OpenAI APIs?
  6. Have you considered how to scale this beyond a hackathon prototype?

Back to contents

Investment/Partnership Verdict

Not evidenced: No financials, funding history, or commercial traction are provided.

The project is described as a hackathon submission with no evidence of revenue, customers, or adoption beyond internal testing and benchmarking. While the concept shows promise in addressing a real problem (lesson plan review), there is insufficient evidence to support investment or partnership interest at this stage.

Confidence level: Low

The description is self-reported and unverified; it does not contain any data on market demand, user engagement, or commercial viability. The product demonstrates technical capability but lacks proof of traction or scalability.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.