OpenAI 2026 hackathon

Synthesis Bench

A long horizon, deterministic medical evidence benchmark

Solo project by Cyrus Nouroozi · 1 likes · 0 comments

Archive position — measured, not model output

1 like on Devpost

506 of the 7,856 archived projects have more likes, and 1,758 share exactly 1 — so this project's #2,022 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Synthesis Bench is a self-reported open-source benchmark for evaluating AI agents on medical evidence synthesis tasks. The project was built by one individual (Cyrus Nouroozi) over six days using Codex and Python, with the goal of measuring how well AI systems can extract data from full-text papers and judge risk-of-bias in line with expert decisions.

What changed

The author states that this is a new benchmark designed to evaluate deterministic performance in medical evidence synthesis. It uses public datasets and deterministic grading logic, without relying on human judges or reward models. The project was submitted to the OpenAI 2026 hackathon.

Single most important open question — commercial due-diligence read

Is there any indication that this project has traction, adoption, or a path toward monetization? There is no evidence of revenue, customers, or product-market fit beyond the author's own description. The benchmark is described as a research tool with no clear commercial application.

Back to contents

What The Product Actually Is

The description states that Synthesis Bench is a benchmark for evaluating AI agents on medical evidence synthesis tasks. It measures whether an AI can extract arm-level results (means, SDs, Ns, events) and judge risk of bias from full-text papers, using deterministic code to grade against expert decisions.

It includes:

  • A dataset covering 13 reviews / 40 papers / 194 extraction facts
  • 10 reviews / 35 papers / 178 risk-of-bias judgments
  • Gold data pinned to upstream release SHAs
  • Deterministic grading logic with recall-only and forgiveness-asymmetric scoring

The tool was built using Codex and Python, and the author reports that it is fully reproducible and open-source.

Evidence

  • The description states: “synthesis-bench measures whether an AI agent can do medical evidence synthesis...”
  • It uses “deterministic code against the published expert decisions.”
  • It includes a dataset built on public MetaPsy databases.
  • It was built in one continuous six-day Codex session.

Inference This is a research-grade benchmark, not a commercial product or SaaS offering.

Back to contents

Positioning & Claim Evolution

The author positions Synthesis Bench as:

  • A benchmark for AI agents doing medical evidence synthesis
  • A deterministic, verifiable RL environment
  • A tool to evaluate how well frontier models perform on tasks that are “some of the highest-stakes, most labor-intensive knowledge work that exists”

It is framed as a research contribution and not a product or service.

Evidence

  • The tagline: “A long horizon, deterministic medical evidence benchmark”
  • The author states: “This is exactly the shape a verifiable-reward RL environment needs”
  • It was submitted to a hackathon (OpenAI 2026)

Inference The project is positioned as a research tool for AI evaluation in medicine, not a commercial offering.

Back to contents

Target Customer & ICP

Not evidenced. The description does not state who the intended users or customers of Synthesis Bench are beyond its own authors and potential researchers or developers evaluating AI systems.

Evidence

  • No mention of target users or customer segments
  • No indication of whether it is meant for clinicians, researchers, or AI developers

Back to contents

Business Model & Pricing Evidence

Not evidenced. There is no mention of pricing, monetization, or business model in the description.

Evidence

  • No revenue streams, pricing tiers, or commercial use cases described
  • The project is open-source and self-reported as a benchmark for research

Back to contents

Technical & Delivery Signals

The author reports:

  • Built with Codex and Python
  • One continuous six-day session using 154 asks, 7,221 commands, 881 patches, 385 subagents
  • The system resolved DOIs, downloaded PDFs, and OCR'd 753 papers with zero failures
  • Gold data was verified by researchers themselves
  • Grading is recall-only and forgiveness-asymmetric
  • Deterministic grading logic without judge models

Evidence

  • “One continuous six-day Codex session”
  • “Codex built the acquisition pipeline that resolved DOIs, downloaded PDFs, and OCR'd 753 papers with zero failures”
  • “Grading is recall-only and forgiveness-asymmetric by design”

Inference The technical architecture is described as a hybrid of deterministic scripts and LLM-driven execution, with strong emphasis on reproducibility and determinism.

Back to contents

Traction & Maturity Signals

Not evidenced. There is no mention of adoption, usage, or traction beyond the author’s own development and submission to a hackathon.

Evidence

  • No customers, users, or product adoption reported
  • No revenue or monetization data
  • No evidence of product-market fit or commercial interest

Back to contents

Competitive Context

Not evidenced. The description does not mention competitors or similar tools in the space of AI-driven medical evidence synthesis or benchmarking.

Evidence

  • No reference to existing benchmarks, tools, or platforms in this domain
  • No comparison with other research efforts or open-source projects

Back to contents

Key Risks & Red Flags

  • No commercial application or monetization path: The project is described as a research tool, not a product.
  • Unproven adoption or traction: There is no evidence of usage beyond the author’s own development.
  • Highly specialized domain: Medical evidence synthesis is niche; it's unclear if there is broader market demand.
  • Single-person team: The entire project was built by one individual, which raises questions about scalability and long-term maintenance.

Inference The project lacks commercial viability or traction. It is a research prototype with no clear path to monetization or product-market fit.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the intended use case for Synthesis Bench beyond research?
  2. Are there any plans to expand this into a product or service?
  3. How does this benchmark compare to existing tools in the field of medical evidence synthesis?
  4. Is there any interest from academic institutions, pharmaceutical companies, or health organizations in using or adopting this tool?
  5. What are the long-term goals for this project beyond the hackathon submission?

Back to contents

Investment/Partnership Verdict

Not evidenced. There is no indication that Synthesis Bench has attracted investment, partnerships, or commercial interest.

Evidence

  • No funding rounds, investors, or partnership announcements
  • No evidence of product-market fit or traction
  • The project is described as a research tool submitted to a hackathon

Inference At this stage, the project appears to be a prototype with no clear path to investment or commercialization. It may have academic or research value but lacks signs of a viable business model or market demand.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.