OpenAI 2026 hackathon

Evidence-Led Litigation Review

Turn mixed evidence and repeated Japanese legal-drafting revisions into a verifiable, local-first Codex workflow.

Solo project by KAZUKI Tozawa · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,995 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

The project described by the caller is a local-first Codex plugin for reviewing Japanese litigation filings. It is presented as a tool that converts evidence into reusable text dossiers, records source and verification states, draws two-sided chronologies, and lists deterministic inconsistency candidates between filings, claims, and exhibits. The system uses a two-layer architecture: deterministic processing (Python-based) for structural checks and GPT-5.6 in Codex for reasoning.

What changed

The author describes building a self-contained, privacy-conscious tool for legal document review that separates deterministic logic from model judgment. It is not a commercial product but a hackathon submission with no real-world deployment or customer data.

Single most important open question

Is there any evidence of traction, revenue, or adoption beyond the author’s solo development and fictional demo?

Back to contents

What The Product Actually Is

The description states that Evidence-Led Litigation Review is a local-first Codex plugin for reviewing Japanese litigation filings. It performs:

  • Conversion of supported local evidence into reusable text dossiers
  • Recording of source and verification states
  • Generation of two-sided chronology (for self, opponent, and neutral material)
  • Listing of deterministic inconsistency candidates between filings, claims, and exhibits
  • Production of proofreading lists for mistypes, registered misconversions, duplicate input, punctuation defects, unknown exhibit references, and exhibit-number repairs

The system includes a seven-pass review process:

  1. Evidence inventory
  2. Issue chain
  3. Evidence-claim matrix
  4. Causation and damages
  5. Strongest-opponent review
  6. Expression
  7. Filing-day verification

It also supports:

  • TXT, Markdown, DOCX, JSON inputs
  • Optional local PDF/image extraction via Poppler and Tesseract
  • Outputs in Markdown, JSON, CSV, HTML, SVG

The system is built using Codex-plugins, codex-skills, Python, GPT-5.6, and other tools.

Not evidenced No evidence of actual commercial use, real customer data, or live deployment beyond the author’s solo development and fictional demo.

Back to contents

Positioning & Claim Evolution

The description states that the project was built to address a problem in legal drafting: repeated revisions can improve prose but weaken evidentiary structure. The tool aims to:

  • Begin with the record
  • Preserve connection between claims and evidence
  • Test the strongest opposing explanation
  • Make every important revision easier to verify

It positions itself as a local-first, privacy-conscious solution that separates deterministic checks from model reasoning.

The author also notes that it avoids two extremes:

  1. Generic grammar/regex checking (which should not decide semantic/legal questions)
  2. Systems that appear to decide legal merit (to maintain human responsibility)

Inferred The tool is positioned as a compliance and quality-assurance aid, not a decision-making engine or legal advice provider.

Back to contents

Target Customer & ICP

The description states that the product is intended for reviewing Japanese litigation filings, particularly in contexts where:

  • Evidence must be preserved
  • Claims and exhibits must be linked
  • Structural errors (e.g., citation drift, damage period overlap) are common

It is built for legal professionals or paralegals who work with complex, multi-party disputes.

Not evidenced No evidence of actual customers, use cases beyond the fictional demo, or target market segmentation.

Back to contents

Business Model & Pricing Evidence

The description does not mention any pricing model, revenue streams, or commercialization strategy. It is presented as a hackathon submission, with no indication of monetization.

Not evidenced No evidence of business model, pricing, or customer acquisition.

Back to contents

Technical & Delivery Signals

The system uses:

  • A two-layer architecture: deterministic Python tools (no network requests) and GPT-5.6 in Codex for reasoning
  • Local processing only, with optional OCR via Poppler and Tesseract
  • Supports TXT, Markdown, DOCX, JSON inputs
  • Outputs in Markdown, JSON, CSV, HTML, SVG
  • Built using Codex-plugins, codex-skills, Python, GPT-5.6, and other tools

The author states:

  • The deterministic core is dependency-free
  • It includes automated testing (23 unit/integration tests)
  • A cross-platform one-command demo requires no API key
  • Privacy scans are included to prevent re-identification

Not evidenced No evidence of scalability, performance benchmarks, or production deployment.

Back to contents

Traction & Maturity Signals

The project is described as a hackathon submission, with:

  • No real-world use cases
  • A fictional public sample (no real party, filing, address, case number, medical record, or exhibit image)
  • No customer data, revenue, or adoption metrics

Not evidenced No evidence of traction, usage, or product-market fit beyond the author’s solo development.

Back to contents

Competitive Context

The description does not mention any competitors. It is presented as a novel approach to legal document review, combining deterministic checks with model reasoning in a local-first environment.

Not evidenced No evidence of existing solutions or competitive landscape.

Back to contents

Key Risks & Red Flags

  • The project is a solo-built hackathon submission, not a commercial product
  • No evidence of real-world adoption, customers, or revenue
  • The fictional demo may misrepresent the tool’s utility in actual legal practice
  • The separation between deterministic and model layers is described as intentional, but it's unclear how this would scale or be validated in practice
  • The use of GPT-5.6 implies reliance on a proprietary, non-open model, which could pose risks for long-term deployment

Inferred The tool may not yet be ready for commercial use or integration into legal workflows.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual legal context in which this would be used? Is it intended for a specific jurisdiction or type of case?
  2. How does the deterministic layer handle edge cases that are not explicitly programmed?
  3. Has the model reasoning been validated by legal professionals or practitioners?
  4. What are the limitations of the current architecture in terms of scalability and performance?
  5. Are there plans to integrate with existing legal document management systems?
  6. Is there any plan for monetization, or is this purely a proof-of-concept?

Back to contents

Investment/Partnership Verdict

The project is described as a solo-built hackathon submission with no evidence of traction, revenue, or customer adoption.

Not evidenced No evidence of commercial viability, product-market fit, or scalability beyond the author’s development environment.

Inference This is likely an early-stage idea or prototype, not a product ready for investment or partnership. It may be a useful proof-of-concept but lacks the commercial signals needed to assess its potential for growth or impact.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.