OpenAI 2026 hackathon

Model Jury

An evidence-first game for students where they learn content and use it to take a flawed AI to court. The students prosecute the AI by exposing unsupported claims, testing repairs, and eventually win.

Solo project by Aditya Chandran Arvind · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,364 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Model Jury is a self-reported educational tool built as a demo for the OpenAI 2026 hackathon. The author describes it as an "evidence-first game" for students that teaches content while improving AI literacy through simulated courtroom-style interactions with flawed AI models.

What changed

The project is presented as a prototype or demo lesson, not a commercial product. It was built using Codex and GPT-5.6 during OpenAI Build Week, with no indication of prior development or production use beyond this single submission.

Single most important open question — the commercial due-diligence read

Is there evidence that Model Jury has moved beyond a one-person hackathon demo into any form of scalable, repeatable educational product or service? The description provides no information on traction, revenue, customers, or adoption beyond its own self-report.

Back to contents

What The Product Actually Is

The description states that Model Jury is an “evidence-first game for students” where they learn content and use it to take a flawed AI to court. Students identify unsupported claims in AI responses, test repairs, and eventually engage in a courtroom-style battle. It includes:

  • A biology lesson using real photographs and source notes.
  • An AI model (BioGuide) that is intentionally flawed in one of five field notes.
  • A student journey from learning through practice cases to a final court-like trial.
  • A teacher studio for defining tasks, goals, and reviewing AI-generated lessons.

The product is described as a demo lesson built with Next.js 16, React 19, TypeScript, Cloudflare runtime, Zod, Zustand, Motion, Vitest, Testing Library, Playwright, and Codex/GPT-5.6 integration.

Inference The author claims this is a complete educational experience, but it is not evidenced to be anything more than a prototype or demo.

Back to contents

Positioning & Claim Evolution

The description states that the project was inspired by a desire to improve AI literacy beyond simple hallucination detection — aiming instead for students to learn how to evaluate mostly correct answers. It positions itself as a way to teach both content and critical thinking about AI use in an engaging, game-like format.

It also claims to be a “new way” for students to learn subjects while improving AI literacy in a fun, game-like way.

Inference The positioning is focused on education and AI literacy, not commercialization or scalability. There is no evidence of prior market positioning or branding beyond this single submission.

Back to contents

Target Customer & ICP

The description states that Model Jury targets students and teachers in an educational setting. Students engage with the content through lessons and courtroom-style interactions, while teachers use a separate “studio” to define tasks, goals, source packets, and review AI-generated lessons.

Inference The target customer is likely K-12 or higher education institutions, but there is no evidence of actual customers or institutional adoption. The ICP appears to be educators and students using the tool for learning and AI literacy training.

Back to contents

Business Model & Pricing Evidence

The description does not provide any information about pricing, monetization, or business model. It only describes a demo lesson built as part of a hackathon submission.

Inference No evidence exists of a business model or pricing structure beyond the author’s own self-reporting.

Back to contents

Technical & Delivery Signals

The project is described as built with:

  • Next.js 16, React 19, strict TypeScript
  • Cloudflare-compatible runtime
  • Zod, Zustand, Motion, CSS Modules, Vitest, Testing Library, Playwright
  • Deterministic no-key mode for judging
  • Server-only OpenAI provider boundary using official JavaScript SDK and Responses API
  • Structured Outputs with Zod validation
  • Bounded retries and timeouts
  • Trace IDs, redacted errors, store disabled

It also includes:

  • A local evaluation engine scoring evidence grounding, uncertainty calibration, misconception avoidance, and helpfulness.
  • A trial engine applying assignment-specific removal standards without certifying entire models.
  • Visual system built with Codex/GPT-5.6.

Inference The technical stack suggests a modern, scalable architecture, but there is no evidence of production deployment or usage beyond the demo.

Back to contents

Traction & Maturity Signals

The description states that this was built as a “demo lesson” during OpenAI Build Week and submitted to a hackathon. It includes:

  • A public app URL: https://model-jury.lucid-nightmare-og.chatgpt.site
  • Core lesson URL: https://model-jury.lucid-nightmare-og.chatgpt.site/demo
  • Courtroom showcase URL: https://model-jury.lucid-nightmare-og.chatgpt.site/trial/bioguide?dev=1
  • GitHub source code: https://github.com/lucid-nightmares/model-jury

There is no evidence of user base, revenue, or adoption beyond the author’s own account.

Inference The project has no demonstrated traction or maturity beyond a single-person hackathon demo.

Back to contents

Competitive Context

The description does not mention any competitors or existing products in the AI literacy or educational space. It only describes its own approach to teaching AI literacy through simulated courtrooms and evidence-based reasoning.

Inference No competitive landscape is described, nor is there any indication of similar tools or platforms in the market.

Back to contents

Key Risks & Red Flags

  • No commercial traction or revenue: The project is described as a demo lesson with no evidence of adoption or monetization.
  • Single-person team: Only one member listed (Aditya Chandran Arvind), raising questions about scalability and long-term development.
  • No institutional or user feedback: There is no mention of pilot programs, teacher feedback, or student engagement beyond the author’s own account.
  • Hackathon origin: The project was submitted to a hackathon, suggesting it may not be intended for commercial use or further development.

Inference These are red flags for commercial viability and scalability without additional evidence.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended path from this demo to a scalable educational product?
  2. Has there been any user testing with teachers or students beyond the author’s own playtesting?
  3. Are there plans to expand beyond biology or add more subjects?
  4. How does the team plan to monetize or distribute this tool if it were to evolve into a product?
  5. What are the technical and legal considerations around using AI models like GPT-5.6 in an educational setting?

Back to contents

Investment/Partnership Verdict

The description states that Model Jury is a demo lesson built for a hackathon, with no evidence of commercial traction or product-market fit beyond its own self-reporting.

Inference There is insufficient evidence to support investment or partnership interest at this time. The project appears to be an experimental prototype with no demonstrated path to scale or revenue generation.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.