Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,337 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be: Misconception Map is a teacher-facing diagnostic tool that uses AI to analyze student exam work and identify systematic misconceptions behind errors. The author states it is built as a local-first application using GPT-5.6, SQLite, and React, with no cloud accounts or data transmission.
What changed: The project evolved from an idea about how grading should be followed by diagnostic insight into a working prototype that processes handwritten exams through AI to find misconceptions, with engineering rigor around confidence thresholds and refusal to guess.
The single most important open question: Does the author's self-reported demonstration of 15 booklets, 393 questions diagnosed, and 59 uncertain items reflect actual product capability or a limited test case that may not scale?
What The Product Actually Is
The description states that Misconception Map is a "teacher-facing diagnostic workspace" that processes exam materials (PDF, photo, or typed) to extract student work, diagnose misconceptions, and generate follow-up teaching resources. It uses GPT-5.6 for transcription and diagnosis, with structured outputs and SQL triggers enforcing validation rules.
The product is described as a local-first application built with Next.js, SQLite, and React, where no student names are sent to an API and all processing happens on the teacher's machine.
Evidence: The author states that the tool extracts exam structure, transcribes student work step-by-step, diagnoses misconceptions with quoted evidence from the copy, and generates teaching briefs or practice sheets. It also includes a "Prediction Lab" that builds per-student models based on repeated errors.
Inference: The tool appears to be designed for classroom-level use by individual teachers rather than school-wide deployment.
Positioning & Claim Evolution
The author positions Misconception Map as an AI diagnostic tool that moves beyond simple grading to identify the root causes of student errors. It claims to find misconceptions, not just wrong answers, and to provide actionable teaching resources based on those insights.
The project evolved from a conceptual understanding of educational research (Brown & Burton, 1978; Sleeman, 1984) into an implementation that uses AI for diagnosis with specific constraints around evidence and confidence thresholds.
Evidence: The author claims that the tool finds systematic procedural bugs in student errors, distinguishes between slips and misconceptions, and provides targeted teaching resources. It also states that it enforces rules like "no AI score reaches the gradebook before explicit validation" and that "a wrong answer alone is never a misconception."
Inference: The positioning reflects an educational technology product aimed at improving teacher effectiveness through data-driven insights.
Target Customer & ICP
The target customer is described as teachers who need to understand why students are making errors on exams, particularly in mathematics. The tool is positioned for classroom-level use and appears to be designed for individual educators rather than institutions or large-scale deployment.
Evidence: The product is described as a "teacher-facing diagnostic workspace" with no accounts or cloud infrastructure, implying it's intended for single-user local deployment.
Inference: The ICP likely includes K-12 teachers, especially those in math instruction, who want to move beyond grading toward deeper pedagogical insight.
Business Model & Pricing Evidence
No explicit business model or pricing information is provided. The project is described as a hackathon submission and appears to be built for local use without any mention of monetization or customer acquisition strategies.
Evidence: The description states that the tool runs locally on a teacher’s machine, with no accounts or data transmission to an API. There is no mention of revenue streams, subscriptions, or pricing tiers.
Inference: It's unclear whether this is intended as a commercial product, a prototype for further development, or a proof-of-concept.
Technical & Delivery Signals
The tool uses GPT-5.6 with structured outputs and strict schema enforcement. The stack includes Next.js, SQLite, React, Docker, and better-sqlite3. It has an adversarial test suite covering failure modes like faint ink and PDF signature abuse.
Evidence: The author states that Codex wrote most of the code, but boundaries were set to prevent drift. The tool uses a strict root-object structured output for all GPT calls, records prompt/schema versions, token counts, and latency. It also includes SQL triggers to enforce validation rules.
Inference: The engineering approach shows attention to quality control and reproducibility, with explicit mechanisms to avoid incorrect diagnoses.
Traction & Maturity Signals
There is no evidence of revenue, customers, or adoption beyond the author’s own testing. The project was submitted as a hackathon entry and has not been demonstrated in real-world settings.
Evidence: The author reports running five exams through the pipeline with 15 booklets, 393 questions diagnosed, and 59 uncertain items. However, this is described as a limited test case ("real data is humbling") and not representative of broader usage.
Inference: The product appears to be in early development or prototype stage, with no evidence of market traction or user base.
Competitive Context
No direct competitors are mentioned in the description. The tool addresses a niche within educational technology focused on diagnostic tools for teachers, but there is no indication of existing solutions in this space.
Evidence: The author references prior research (Brown & Burton, Sleeman) and mentions that pedagogy constraints belong in database triggers, suggesting an approach distinct from generic AI grading tools.
Inference: This may be a novel or underserved area within edtech, but there is no evidence of competitive landscape analysis or market positioning against existing tools.
Key Risks & Red Flags
- Limited testing scope: The author notes that the demo was based on limited test cases (French brevet papers) and that the Prediction Lab showed nothing for that class.
- Unclear scalability: The tool is built for local use, which may limit its ability to scale or integrate with larger systems.
- No commercial viability: No evidence of a business model, pricing strategy, or customer base suggests the project may not be ready for market.
- Dependency on AI accuracy: Reliance on GPT-5.6 for transcription and diagnosis introduces risk if performance degrades or becomes costly.
Evidence: The author states that the tool was tested with only 15 booklets and that the Prediction Lab did not find predictive models, indicating limited real-world validation.
Diligence Questions To Ask The Founders
- What is the actual accuracy rate of the AI in diagnosing misconceptions versus false positives?
- How does the system handle edge cases like multiple students with similar errors or complex handwriting?
- Are there plans to expand beyond local-first use, and what would that look like technically?
- Has the tool been tested with real teachers in classrooms, and what feedback has it received?
- What are the long-term costs of running this system at scale, especially with GPT-5.6?
Investment/Partnership Verdict
Not evidenced.
The project is described as a hackathon submission with no evidence of traction, revenue, or customer adoption. While the technical approach shows rigor and attention to detail, there is insufficient information to assess its commercial viability or potential for investment or partnership.
Evidence: The author describes a working prototype but does not provide any data on usage, impact, or monetization.
Inference: This appears to be an early-stage idea with strong engineering execution, but lacks the commercial signals needed for investment or strategic interest.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.
