OpenAI 2026 hackathon

Eleza

A transparent oral-defense engine: it examines essays, code, lab reports, and case analyses in a live voice viva, shows why it asks every question, and hands teachers evidence, never a verdict.

Solo project by Jeremiah Somoine · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,899 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be: Eleza is a self-reported AI-powered oral-defense engine designed to examine submitted academic work (essays, code, lab reports, case analyses) through a live voice viva. It parses submissions into structured claim graphs and uses GPT-5.6 to generate questions that target specific claims, with the examiner showing its reasoning for each question in real time. The system is built to be transparent, auditable, and non-judgmental, producing dossiers of evidence rather than verdicts.

What changed: The project was built solo by one developer over seven days as part of an OpenAI hackathon submission. It is described as a proof-of-concept with no revenue or customer data, but includes detailed technical documentation and a build log.

The single most important open question: Is there any evidence that Eleza has been used in real-world academic settings or tested with actual students and educators?

Back to contents

What The Product Actually Is

  • The description states that Eleza is a transparent oral-defense engine for submitted academic work.
  • It parses submissions into a graph of claims, each anchored to exact character offsets.
  • A live viva follows, conducted via voice or typed answers, where the examiner decides questions and rationale streams in real time.
  • The system uses GPT-5.6 across multiple roles: claim graph generation, live examiner, and divergence analysis.
  • It produces a dossier, not a verdict — linking transcript timestamps to document spans with suggested follow-up questions.
  • The engine is domain-parameterized, meaning profiles define vocabulary and semantics while invariants remain universal.

Note: No evidence of actual product usage or deployment beyond the developer's own build log and demo.

Back to contents

Positioning & Claim Evolution

  • The description states that Eleza was built to address two problems:
    1. AI making student work unfalsifiable.
    2. Existing AI oral-assessment tools evaluating behind closed doors without transparency.
  • It positions itself as a transparent alternative that scales the conversation and makes the examiner show its work.
  • The author claims it started as an answer to the failure of current AI tools to be more auditable than the work they examine.
  • Eleza is described as domain-parameterized, supporting essays, code, lab reports, and case analyses.

Inference: This is a self-reported positioning shift from general-purpose AI assessment to transparent, auditable academic defense. No evidence of prior versions or market traction.

Back to contents

Target Customer & ICP

  • The description states that Eleza targets teachers and academic integrity offices, particularly those dealing with large classes where traditional oral exams are unscalable.
  • It is designed for use by educators who need to assess student understanding in a scalable yet transparent way.
  • The system is described as being built for teacher-configurable vivas, suggesting it's intended for institutional adoption.

Not evidenced: No specific customer names, usage data, or institutional partnerships are mentioned. The target audience is inferred from the stated use case.

Back to contents

Business Model & Pricing Evidence

  • The description does not mention any pricing model or business model.
  • It states that the hosted demo caps sessions at about three minutes as a public-cost control.
  • The system is described as domain-parameterized, with profiles defining node vocabulary and edge semantics, but no indication of how these would be monetized.

Not evidenced: No revenue streams, pricing tiers, or commercialization plans are provided.

Back to contents

Technical & Delivery Signals

  • Built solo in seven days during a full-time internship.
  • Uses Codex, GPT-5.6, Next.js, Supabase, TypeScript, Vercel.
  • The build process is documented in an append-only BUILD_LOG.md, with 36 entries including failures and corrections.
  • Three distinct GPT-5.6 models are used: gpt-5.6-sol for claim graphs, gpt-5.6-terra for live examiner.
  • The system enforces architectural invariants:
    • Voice model talks; examiner decides.
    • Append-only decision log.
    • Schema-level rationale receipts.
    • No register comparison or verdicts.
  • Audio retention was declined due to privacy concerns; transcripts carry evidence without biometric residue.

Inference: Strong technical rigor is claimed, but no independent verification of performance, scalability, or reliability in real-world use.

Back to contents

Traction & Maturity Signals

  • The project is described as a self-contained hackathon submission.
  • No revenue, customers, or traction data are provided.
  • The system includes acceptance tests, fixtures with documented weak spots, and expected findings.
  • The build log shows iterative development and testing phases.

Not evidenced: No evidence of real-world deployment, user feedback, or adoption metrics.

Back to contents

Competitive Context

  • The description states that existing AI oral-assessment tools evaluate behind closed doors and return grades or flags.
  • It positions Eleza as a transparent alternative to such tools.
  • It is implied that the product addresses a gap in current AI-based academic integrity solutions.

Not evidenced: No mention of competitors, market size, or competitive positioning beyond self-description.

Back to contents

Key Risks & Red Flags

  • The system is described as a proof-of-concept, not a production-ready tool.
  • No evidence of real-world testing or user feedback.
  • The entire build was done solo in seven days; no team or external validation is mentioned.
  • The product relies heavily on GPT-5.6, which may introduce risks related to hallucination, consistency, and control.
  • The lack of any commercialization plan or pricing model raises questions about scalability.

Inference: High risk due to unproven market fit, lack of real-world testing, and reliance on a single developer.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the current status of Eleza beyond the hackathon demo?
  2. Has it been tested with actual students or educators?
  3. Are there any plans for commercialization or institutional partnerships?
  4. How does the system handle edge cases in claim graphing or examiner reasoning?
  5. What are the limitations of the GPT-5.6 models used, and how are they mitigated?

Back to contents

Investment/Partnership Verdict

  • The description states that Eleza is a self-reported proof-of-concept built for an OpenAI hackathon.
  • No evidence of revenue, customers, or traction exists beyond the author’s own account.
  • It is described as a domain-parameterized tool with strong technical design but no commercialization strategy.
  • The system is positioned to address a real problem in academic integrity but lacks any demonstration of adoption or impact.

Verdict: Not ready for investment or partnership. The project shows strong technical execution and clear intent, but there is no evidence of traction, market validation, or scalability beyond the author's solo build.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.