Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #4,371 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
The project described by the caller is a technical demo named GPT Error Archaeologist, built for the OpenAI Build Week Education category. It explores how an AI system might diagnose student math errors in a way that treats hypotheses as provisional, not final verdicts. The system uses GPT-5.6 Luna to analyze synthetic handwritten math solutions and propose candidate explanations, each linked to evidence from the work. It then generates a follow-up problem designed to distinguish between these candidates, with SymPy independently verifying the mathematical validity of those predictions before they are shown to users.
What changed
This is a working demo built during a hackathon; it does not represent a product in production or an operational service. The author states that this is a narrow exploration of one gap in math education—diagnosing why a student made a mistake—and not a full solution for classroom use. It includes no real-world data processing, authentication, uploads, or teacher confirmation features.
Single most important open question
Is there evidence that educators find value in the diagnostic workflow proposed here? The demo proves an end-to-end interaction but does not show traction, adoption, or effectiveness in improving learning outcomes.
What The Product Actually Is
The description states:
- GPT Error Archaeologist is a technical demo for the OpenAI Build Week Education category.
- It analyzes synthetic handwritten math solutions using GPT-5.6 Luna.
- It proposes two candidate explanations for an error, each connected to evidence from the student's work.
- It generates a differentiating follow-up problem and uses SymPy to verify that predictions differ between candidates.
- The system updates hypothesis support based on simulated student responses.
- It includes abstention as part of its behavior when input is ambiguous or invalid.
Inference The product is not a commercial offering but a prototype exploring a specific pedagogical interaction model. It is built as a modular monolith with React frontend, FastAPI backend, and Docker deployment on Google Cloud Run.
Positioning & Claim Evolution
The description states:
- The system treats hypotheses as candidates—not facts about the student’s mental state.
- It avoids turning uncertain evidence into confident labels.
- It proposes plausible explanations, shows where each comes from, and asks one small question to produce new evidence.
- Diagnosis is falsifiable; the AI does not get the final word—the student’s next answer can support, weaken, or fail to distinguish hypotheses.
Inference The positioning emphasizes honesty in diagnosis over certainty. It positions itself as a diagnostic tool that supports teacher decision-making rather than replacing it. The author explicitly rejects “mind-reading” language and aims for transparency in how hypotheses are formed and updated.
Target Customer & ICP
The description states:
- Initial users are teachers and tutors.
- The demo focuses on Grade 7–9 math education.
- It does not yet support real-student data processing or uploads.
- No mention of specific customer segments beyond educators involved in tutoring or intervention.
Inference The target ICP appears to be educators working with struggling students, particularly those in intervention or remedial settings. However, the demo is limited to synthetic samples and lacks integration with actual classroom workflows or real student data.
Business Model & Pricing Evidence
The description states:
- No pricing information is provided.
- There is no indication of monetization strategy.
- The demo does not include features like authentication, class aggregation, or teacher confirmation.
- Future steps include testing a paid pilot with intervention or tutoring providers, but this is speculative.
Inference There is no evidence of a business model or pricing structure. Any future commercialization would likely depend on validation through educator interviews and pilot studies.
Technical & Delivery Signals
The description states:
- Built with React, Vite, FastAPI, Python, Docker, OpenAI API, SymPy, SQLite, and Google Cloud Run.
- Uses GPT-5.6 Luna with structured outputs.
- Follow-up problems are generated using SymPy to ensure mathematical distinctiveness.
- Independent verification is done via SymPy before backend marks a probe as verified.
- The system separates model reasoning from deterministic verification.
- Tests include 15 backend tests, four frontend workflow tests, and a production build.
- A fake model adapter allows judges to reproduce behavior without API keys.
Inference The technical architecture is modular and reproducible. It uses a hybrid approach combining AI reasoning with symbolic algebra for validation. The separation of concerns between the language model and deterministic checker suggests an intentional design choice to maintain reliability.
Traction & Maturity Signals
The description states:
- This is a hackathon demo.
- It supports curated synthetic samples only.
- No real-world data processing, authentication, or uploads are included.
- The author notes that current demo proves an end-to-end technical interaction, not classroom efficacy.
- No mention of users, customers, revenue, or adoption.
Inference There is no evidence of traction, customer base, or market validation. The project exists at the prototype stage and has not moved beyond a proof-of-concept.
Competitive Context
The description states:
- There are existing tools for grading and tutoring.
- Grading shows that an answer is wrong; tutoring provides remediation.
- This tool explores a gap between those two approaches—diagnosing why a mistake happened.
- No direct competitors are named or described.
Inference The project addresses a niche in math education where diagnosis is underdeveloped. It does not appear to directly compete with existing AI grading or tutoring platforms but instead proposes an alternative diagnostic workflow.
Key Risks & Red Flags
The description states:
- The demo supports only curated synthetic samples.
- No real-world data processing, authentication, or uploads are included.
- The author acknowledges that public statistics do not prove product effectiveness.
- No evidence of educator demand or learning impact.
- Future steps include external evaluation sets and interviews with teachers, but these have not yet occurred.
Inference Key risks include lack of real-world validation, limited scope, and absence of any commercial or educational traction. The project is far from a product-ready solution and lacks clear evidence of market demand or user need beyond the author’s own hypothesis.
Diligence Questions To Ask The Founders
- What specific feedback have you received from educators about this diagnostic workflow?
- How do you plan to validate that the proposed hypotheses are accurate in real-world use cases?
- Have you conducted any pilot studies or usability tests with teachers or students?
- What is your roadmap for moving from a demo to a scalable product?
- Are there plans to integrate with existing educational platforms or systems?
- How will you handle privacy and data governance, especially if real student data is introduced?
Investment/Partnership Verdict
The description states:
- This is a hackathon demo.
- It does not represent a commercial product or service.
- No revenue, customers, or traction are evidenced.
- The author plans to validate the concept through external evaluation sets and educator interviews.
Inference At this stage, there is no basis for investment or partnership. The project is a technical exploration with no demonstrated market fit, user adoption, or business model. It may be of interest as a proof-of-concept or research direction, but it does not meet criteria for commercial due diligence.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.
