OpenAI 2026 hackathon

MindDiff, git diff for mental models

Diagnoses the false belief behind a student's failing code and proves it by running an experiment, not by guessing it. Built with Codex and GPT-5.6.

Solo project by Jugger Naut · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,310 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

Company: MindDiff — a self-reported educational tool that uses AI to diagnose student misconceptions in coding assignments by running experiments instead of guessing.

What Changed: The project description states the team built a system that diagnoses false beliefs behind student code errors using GPT-5.6 and Pyodide sandboxed execution, aiming to move beyond traditional autograders or AI tutors that only report failure or provide fixes.

Single Most Important Open Question: Is there evidence of real-world use or traction with students or instructors? The description contains no data on adoption, usage, or feedback from actual users.

Back to contents

What The Product Actually Is

The description states that MindDiff is a tool for diagnosing student misconceptions in coding assignments. It reads failing Python submissions, hypothesizes the false belief behind the error using GPT-5.6, and designs discriminating probes to test each hypothesis via execution in a Pyodide sandbox. The system determines whether the misconception is confirmed, ruled out, or unverifiable based on execution results.

Inference: The tool appears to be built for educational use, specifically in programming education, with an emphasis on Socratic questioning and diagnostic feedback over direct correction.

Back to contents

Positioning & Claim Evolution

The description states that MindDiff was inspired by the problem of AI-generated homework leading to "quietly disappearing understanding" because current tools (autograders and AI tutors) do not diagnose why a student is wrong, only that they are. The tool aims to prove its diagnosis through execution rather than assertion.

Inference: The positioning is that MindDiff is a diagnostic tool for education that moves beyond simple feedback to deeper understanding by using experimental verification.

Back to contents

Target Customer & ICP

The description states that the tool is intended for students and instructors in programming education, particularly those working with Python fundamentals. It mentions an instructor dashboard with heatmaps and clustering of misconceptions, suggesting it targets educators who want to understand class-wide patterns.

Inference: The primary customer segments are likely students learning programming and instructors teaching introductory coding courses.

Back to contents

Business Model & Pricing Evidence

The description does not state anything about pricing or a business model. It only describes the tool’s functionality and how it was built.

Not evidenced

Back to contents

Technical & Delivery Signals

The project was built using:

  • Stack: Next.js, TypeScript, Drizzle ORM over Turso/LibSQL, Zod validated structured outputs from OpenAI SDK, Pyodide for sandboxed Python execution.
  • Core Engine: Built in a single Codex CLI session on GPT-5.6 Terra.
  • Deployment: Vercel/Turso.

The system includes:

  • A Pyodide sandbox for executing probes
  • Structured output from GPT-5.6
  • Rate limiting and input validation

Inference: The tool is built with modern web stack and uses LLMs for reasoning, with sandboxed execution to avoid server-side code risks.

Back to contents

Traction & Maturity Signals

The description states that the system works end-to-end on real submissions, including edge cases. It mentions stress testing and fixing bugs in prompts and cost exploits, but does not provide any data on:

  • Number of users
  • Usage frequency
  • Feedback from educators or students
  • Classroom piloting or integration

Not evidenced

Back to contents

Competitive Context

The description does not mention any competitors or market context beyond the general problem of autograders and AI tutors.

Not evidenced

Back to contents

Key Risks & Red Flags

  1. No Traction: The tool is described as a hackathon project with no evidence of real-world adoption.
  2. Unverified Claims: The system claims to prove diagnoses through execution, but the description does not include validation or testing beyond internal stress tests.
  3. Single Developer: The team size is listed as one (Jugger Naut), which may limit scalability or product development speed.
  4. No Pricing or Monetization Strategy: No indication of how this would be monetized in a real product.

Back to contents

Diligence Questions To Ask The Founders

  1. Has the system been tested with actual students and instructors? What feedback have you received?
  2. How do you plan to scale beyond a single developer?
  3. Are there any plans for classroom piloting or integration with LMS platforms?
  4. What is your path to monetization, if any?
  5. Can you demonstrate how the system handles edge cases in real-world usage?

Back to contents

Investment/Partnership Verdict

Not evidenced

The description is a self-reported account of a hackathon project. It does not contain any evidence of revenue, customers, traction, or a clear business model. The tool’s functionality is described but not validated by external use or feedback.

Confidence: Low. The entire analysis is based on one author's account with no independent verification or data to support claims of impact or viability.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.