OpenAI 2026 hackathon

Blindspot

Catches correct math answers built on wrong reasoning.

Solo project by Mohamd Imran · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,967 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be: Blindspot is a self-reported educational tool built for teachers to detect hidden misconceptions in student work where the final answer is correct but the method is flawed. It uses AI (GPT-5.6) to analyze fraction simplification steps and identify invalid operations that may not surface until later problems.

What changed: The project was submitted as part of an OpenAI 2026 hackathon, indicating it is a prototype or early-stage product with no evidence of prior traction, revenue, or customer adoption.

Single most important open question: Is the author’s claim that the system can reliably detect and verify flawed reasoning in student math work testable, or does it rest on unverifiable assumptions about model behavior?

Back to contents

What The Product Actually Is

The description states that Blindspot is a Next.js application using GPT-5.6, built with TypeScript, styled with Tailwind CSS, and leveraging the OpenAI JavaScript SDK and Responses API.

It reviews student work on fraction simplification, identifying whether the method used is mathematically valid in general — not just if the final answer is correct.

It provides:

  • A diff-style review of invalid operations;
  • An explanation of valid rules in educator-facing language;
  • A same-concept problem with different numbers;
  • Deterministic replay of flawed methods to show failure.

The system includes two main routes:

  1. Review route: Uses structured outputs via Zod and system-level instructions.
  2. Stress-test route: Asks GPT-5.6 for adversarial candidates, then validates them using deterministic fraction arithmetic.

Not evidenced: No information on how the product is deployed, whether it has a web interface, or if it integrates with existing LMS platforms.

Back to contents

Positioning & Claim Evolution

The author states that Blindspot addresses a blind spot in grading — where correct answers are accepted without scrutiny of the reasoning behind them. It positions itself as an AI tool that goes beyond traditional grading to find misconceptions hidden by correct outcomes.

It claims:

  • Most education software looks for misconceptions after wrong answers.
  • Correct answers can conceal dangerous misunderstandings (e.g., digit cancellation in fractions).
  • The tool identifies these flaws before they become visible through incorrect answers.

The project’s positioning is framed around detecting flawed reasoning, not just checking correctness — a distinction the author emphasizes as central to its value proposition.

Inferred: This suggests a shift from reactive feedback (after wrong answer) to proactive detection of conceptual gaps. However, this evolution is based on self-reported claims and lacks evidence of prior use or validation.

Back to contents

Target Customer & ICP

The description states that Blindspot targets teachers, particularly those dealing with student math work — specifically in fraction simplification.

It mentions:

  • Teachers facing a grading burden.
  • The need for timely feedback that scales to an entire class.
  • A focus on helping educators identify hidden misconceptions.

Not evidenced: No mention of specific grade levels, school types (K12 vs. higher ed), or whether the tool is intended for individual or group use beyond “Class review.”

Inferred: The ICP likely centers around K12 math teachers who want to improve student understanding through deeper feedback than standard grading allows.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

The description does not include any information about:

  • Revenue model (e.g., SaaS, freemium, licensing);
  • Pricing structure;
  • Monetization strategy;
  • Customer acquisition plans;
  • Any commercial relationships or partnerships.

The project is presented as a hackathon submission and lacks any indication of business development or monetization efforts.

Back to contents

Technical & Delivery Signals

The system is built using:

  • Next.js App Router
  • TypeScript
  • Tailwind CSS
  • OpenAI JavaScript SDK
  • GPT-5.6 model
  • Zod-backed Structured Outputs

Key technical features include:

  • Use of system-level instructions.
  • Deterministic verification layer to validate AI-generated counterexamples.
  • Schema-safe validity results.
  • Stress-test route that generates and validates adversarial examples.

Not evidenced: No information on deployment infrastructure, scalability, or integration capabilities beyond the described tech stack.

Inferred: The tool appears to be a lightweight prototype built for demonstration purposes, not production-ready software.

Back to contents

Traction & Maturity Signals

Not evidenced.

There is no evidence of:

  • Revenue;
  • Customers;
  • User adoption;
  • Product usage metrics;
  • Any form of pilot or beta testing;
  • Prior versions or iterations;
  • Market validation.

The project is described as a hackathon submission, suggesting it is in an early stage and not yet mature for commercial use.

Back to contents

Competitive Context

Not evidenced.

No mention of:

  • Competitors in the education AI space;
  • Existing tools that address similar issues (e.g., misconception detection);
  • Market size or competitive positioning.

The description does not reference any prior art or market analysis, nor does it compare Blindspot to other educational platforms or tools.

Back to contents

Key Risks & Red Flags

  1. Unverifiable claims: The author states the system can detect flawed reasoning and verify its outputs, but there is no evidence of testing or validation.
  2. Model dependency without transparency: Reliance on GPT-5.6 implies potential unreliability if the model fails to behave as expected in real-world use.
  3. Limited scope: The tool currently focuses only on fraction simplification and is described as expanding — this raises questions about scalability and generalization.
  4. No commercial viability evidence: As a hackathon project, there’s no indication of product-market fit or monetization strategy.
  5. Verification gate may be insufficient: While the system includes deterministic checks, it's unclear how robust these are in practice.

Inferred: The tool may not yet be ready for classroom deployment due to lack of testing and real-world validation.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific validation or testing has been done on the AI-generated counterexamples?
  2. How does the deterministic verification layer handle edge cases or ambiguous student work?
  3. Has the tool been tested with actual teachers or students? If so, what were the results?
  4. Is there a plan to expand beyond fraction simplification into other domains (e.g., algebra)?
  5. What is the intended business model and go-to-market strategy?
  6. How does the system handle variations in how students write out their work (e.g., visual vs. symbolic notation)?

Back to contents

Investment/Partnership Verdict

Not evidenced.

There is no evidence of:

  • Funding rounds;
  • Revenue or ARR;
  • Customer base;
  • Team size beyond one person;
  • Any investment interest or partnership discussions.

The project is described as a single-person hackathon submission, indicating it is in an early stage and not yet suitable for investment or partnership consideration.

Inferred: Blindspot shows promise in addressing a real educational challenge, but lacks the traction, validation, or commercial readiness to warrant serious due diligence at this time.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.