OpenAI 2026 hackathon

Falsify

An open-source adversarial evidence engine that stress-tests research, public claims, and strategic narratives against primary sources, contradictions, and logical consistency.

Solo project by Takahiro Tsuchiya · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #4,045 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Falsify is an open-source adversarial evidence engine designed to stress-test research, public claims, and strategic narratives against primary sources, contradictions, and logical consistency. It decomposes documents into testable claims, searches for supporting and contradictory evidence, and evaluates whether conclusions are logically sound or insufficiently evidenced.

What changed

The project was submitted as part of the OpenAI 2026 hackathon. The author describes it as a single Next.js and TypeScript application built with GPT-5.6, Codex, and Gemini, deployed via Vercel. It is publicly available under an MIT License on GitHub.

Single most important open question

Is there any evidence of real-world usage or adoption beyond the author’s own development and demo?

Back to contents

What The Product Actually Is

The description states that Falsify is an open-source adversarial evidence verification engine. It is described as a tool that:

  • Decomposes documents and public statements into testable claims.
  • Identifies what evidence each claim would require.
  • Searches for supporting and contradictory evidence.
  • Checks whether citations actually support the claims attached to them.
  • Highlights logical or causal leaps that exceed available evidence.

It also distinguishes between supported, partially supported, contradicted, insufficiently evidenced, outdated, selectively framed, and logically overextended claims.

The tool is built as a single Next.js and TypeScript application, using GPT-5.6 and structured outputs for claim decomposition and evidence synthesis.

Evidence

  • The author states that Falsify is an adversarial evidence engine.
  • It is described as a single Next.js + TypeScript app.
  • It uses Codex, GPT-5.6, and Gemini for development and deployment.
  • It is publicly available under MIT License on GitHub.

Inference It appears to be a proof-of-concept or prototype built for a hackathon, not yet a commercial product.

Back to contents

Positioning & Claim Evolution

The author states that Falsify was inspired by the idea of asking AI to try to prove an argument wrong, rather than make it more convincing. It is positioned as a tool that challenges claims and exposes evidentiary gaps or logical inconsistencies.

It is described as being useful for:

  • Strategic narratives (e.g., Japan’s defense spending).
  • Academic research.
  • Policy reports.
  • Journalism.
  • Corporate claims.

Evidence

  • The author states the inspiration was to “ask AI to try to prove the argument wrong.”
  • It is positioned as a tool that distinguishes between supported and contradicted claims.
  • It is described as useful for multiple domains, including policy, journalism, and research.

Inference The positioning suggests a focus on adversarial verification in information environments where trust and accuracy are critical. However, the lack of real-world use cases or feedback implies this is an early-stage concept.

Back to contents

Target Customer & ICP

The description states that Falsify can be applied to:

  • Strategic narratives.
  • Academic research.
  • Policy reports.
  • Journalism.
  • Corporate claims.

It is also described as useful for analyzing public statements and official narratives, such as those from Japan’s government.

Evidence

  • The author mentions applications in policy analysis, journalism, and strategic narratives.
  • It is described as useful for analyzing public claims and official statements.

Inference The target customer likely includes researchers, journalists, policy analysts, and public affairs professionals. However, no specific customer segments or personas are defined.

Back to contents

Business Model & Pricing Evidence

There is no evidence of a business model or pricing strategy in the description.

Evidence

  • The project is described as open-source.
  • It is publicly available under MIT License.
  • No mention of monetization, subscriptions, or paid features.

Inference It appears to be a prototype or open-source tool with no commercial revenue model at this time.

Back to contents

Technical & Delivery Signals

The author states that Falsify is built as a single Next.js and TypeScript application, using:

  • GPT-5.6.
  • Structured outputs.
  • Codex for development.
  • Gemini 3.1 Flash-Lite integration (for demo purposes).
  • Vercel for deployment.

It also includes:

  • Typed API responses.
  • Evidence Map UI.
  • Provenance safeguards.
  • Deterministic audits.
  • Testing and security hardening.

Evidence

  • The tool is built with Next.js, TypeScript, GPT-5.6, and Codex.
  • It uses structured outputs for claim decomposition.
  • It includes UI elements like an “Evidence Map.”
  • It is deployed on Vercel.
  • It is open-source under MIT License.

Inference It is a technical prototype built with modern AI tooling, but lacks evidence of production-grade infrastructure or scalability.

Back to contents

Traction & Maturity Signals

There is no evidence of traction, customers, or adoption beyond the author’s own development and demo.

Evidence

  • The project was submitted to a hackathon.
  • It is described as a single-person effort.
  • No mention of users, feedback, or real-world usage.
  • No data on engagement, downloads, or community adoption.

Inference The tool appears to be in early development and has not yet demonstrated any measurable traction or user base.

Back to contents

Competitive Context

There is no evidence of competitors or competitive positioning in the description.

Evidence

  • The author does not mention any existing tools or platforms with similar functionality.
  • No comparison to other adversarial verification, fact-checking, or AI-assisted research tools.

Inference It’s unclear whether Falsify is unique or if there are existing tools addressing similar needs. This is a gap in the description.

Back to contents

Key Risks & Red Flags

  • No evidence of traction or adoption: The tool appears to be a prototype with no real-world usage.
  • Single-person team: No indication of team size beyond one person, which may limit development and scalability.
  • Limited commercialization strategy: It is open-source and lacks any monetization model.
  • Unclear competitive positioning: No mention of existing tools or how it differentiates from them.
  • Demo-only deployment: The public demo uses a free quota, suggesting limited production readiness.

Evidence

  • No revenue, customers, or usage data.
  • Only one team member is mentioned.
  • It’s open-source and not monetized.
  • Deployment is via Vercel with free quotas.
  • No mention of competitors or differentiation.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended user base for Falsify, and how are you planning to reach them?
  2. Are there any real-world use cases or early adopters beyond the demo?
  3. How do you plan to monetize or scale this tool if it’s open-source?
  4. What are the limitations of the current implementation, especially in terms of accuracy and scalability?
  5. How does Falsify handle edge cases or ambiguous claims?
  6. Are there any plans for a more robust backend or API layer beyond the demo?

Back to contents

Investment/Partnership Verdict

Not evidenced.

There is no evidence of revenue, customers, traction, or financials to support an investment or partnership decision.

The project appears to be an early-stage prototype submitted for a hackathon, built by a single developer and released as open-source. It has not demonstrated any commercial viability or market traction.

Confidence Low. The description is self-reported and unverified, with no evidence of real-world usage, adoption, or business model.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.