OpenAI 2026 hackathon

Biomedical-Conjecture: Scientific Generator of Next Target

It builds causal hypotheses (biomedical studies) from millions of pre-cutoff literature and scores each against what was discovered after, an earned hit rate you can trust for its next, untested one.

Solo project by Ahmed Hassoon · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,940 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be: A self-reported scientific hypothesis generator for biomedical research that uses a machine learning pipeline to extract, analyze, and score causal hypotheses from literature before a cutoff date, then evaluates them against later evidence. The system claims to implement predictive backtesting with temporal firewalls and model training cutoffs to avoid data leakage.

What changed: The project description is a self-reported submission for the OpenAI 2026 hackathon. It describes an experimental research tool built by one person (Ahmed Hassoon) using AI tools like ChatGPT, Claude, and Codex. No evidence of commercial traction, funding, or product use exists in the description.

Single most important open question: Is this a working prototype or a conceptual framework? The description states it is "a system that scores its own output against reality", but there are no verifiable claims about actual performance metrics, real-world usage, or validation with external data.

Back to contents

What The Product Actually Is

The description states the product is Biomedical-Conjecture: Scientific Generator of Next Target, which builds causal hypotheses from biomedical literature and scores them based on how well they align with subsequent findings. It implements a 11-stage pipeline including search, retrieval, extraction, synthesis, detection, generation, and evaluation.

It claims to:

  • Extract structured findings with provenance (effect estimates, populations, doses, adjustments, verbatim quotes)
  • Detect anomalies in the literature
  • Generate typed causal hypotheses that resolve tensions
  • Backtest these hypotheses using a temporal firewall between pre- and post-cutoff literature
  • Score each hypothesis against later evidence with calibration metrics

The system is described as running offline, deterministically, and for free under mocked retrievers.

Evidence: The author's own write-up.

Inference: This is an experimental research tool built by one developer using AI frameworks. No evidence of commercial product or real-world deployment exists.

Back to contents

Positioning & Claim Evolution

The description states that the tool aims to make "the synthesis" of biomedical science systematic, scoreable, and honest about its own limits. It positions itself as a way to automate the process of identifying anomalies in literature and generating testable hypotheses — an attempt to institutionalize what is currently a rare human feat.

It claims:

  • The system can detect when evidence doesn't add up
  • It generates hypotheses with mechanisms, scopes, rival explanations, and falsifying experiments
  • It scores hypotheses against future evidence using predictive backtesting
  • It distinguishes between "grounded" (computed from source-verified numbers) and "conjectural" (model-proposed) hypotheses

Evidence: The author's own write-up.

Inference: This is a tool for hypothesis generation in scientific research, not clinical guidance or commercial product development. Its positioning is as a research assistant rather than a commercial offering.

Back to contents

Target Customer & ICP

The description does not name specific customers or target users beyond biomedical scientists and researchers who work with literature and hypothesis formation.

It implies:

  • Researchers working on biomedical studies
  • Scientists seeking to identify promising next targets for experimentation
  • Labs or institutions interested in automating parts of the scientific discovery pipeline

Evidence: The author's own write-up, which describes the tool as helping scientists "chase hypotheses" and "change medicine".

Inference: The ICP is likely academic or research-oriented users who are looking to improve their hypothesis generation process. No evidence of commercial customers or user base.

Back to contents

Business Model & Pricing Evidence

There is no mention of pricing, monetization, or business model in the description.

The system is described as:

  • Running offline and for free under mocked retrievers
  • Deterministic and transactional
  • With enforced budgets and caching

Evidence: The author's own write-up.

Inference: No commercial model is evident. It appears to be a research prototype or hackathon project, not a product with a defined revenue path.

Back to contents

Technical & Delivery Signals

The description states:

  • The system uses a fixed 11-stage pipeline
  • Each stage commits a persisted RunState transactionally
  • Runs are resumable, idempotent, and inspectable
  • It implements a temporal firewall between pre- and post-cutoff literature
  • Model training cutoffs are audited to prevent leakage
  • Extraction is verifiable with controlled vocabularies and full-text papers
  • The system reports Brier score, ECE, and Wilson confidence intervals

Evidence: The author's own write-up.

Inference: These are technical claims about the architecture and validation methods. No evidence of actual deployment or performance data.

Back to contents

Traction & Maturity Signals

There is no evidence of traction, revenue, customers, or adoption in the description.

The project:

  • Was submitted to a hackathon (OpenAI 2026)
  • Is described as a single-person effort
  • Has no mention of users, usage metrics, or product releases beyond the prototype

Evidence: The author's own write-up.

Inference: No signs of commercial maturity or traction. It is likely an experimental prototype or proof-of-concept.

Back to contents

Competitive Context

The description does not reference competitors or existing tools in the space.

It implies:

  • The tool is designed to improve upon human hypothesis generation
  • It aims to be more systematic and scoreable than current methods

Evidence: The author's own write-up.

Inference: No competitive landscape is described. The tool appears to be positioned as a novel approach to scientific discovery, not a replacement for existing tools.

Back to contents

Key Risks & Red Flags

Key risks and red flags include:

  • Unverified claims: All claims are self-reported and unverified
  • No commercial traction: No evidence of users, customers, or revenue
  • Single-person development: The entire project is attributed to one individual
  • Prototype nature: Described as a hackathon submission with no indication of production readiness
  • Limited validation: No real-world testing or benchmarking data provided
  • Technical complexity without demonstration: Claims about temporal firewalls and model audits are made but not demonstrated

Evidence: The author's own write-up.

Inference: This is a conceptual or experimental system, not a proven product. Risks include lack of validation, scalability, and commercial viability.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual performance of the system on real-world biomedical datasets?
  2. How does it handle edge cases in literature extraction and anomaly detection?
  3. Has it been tested against expert annotations or gold-standard datasets?
  4. Is there any plan to move beyond the hackathon prototype into a usable product?
  5. What are the limitations of its current approach, especially regarding full-text access and paywalled papers?
  6. How does it ensure reproducibility and avoid hallucinations in hypothesis generation?
  7. Are there any known issues with model training data leakage that have not been addressed?

Back to contents

Investment/Partnership Verdict

Not evidenced: No evidence of commercial traction, funding, or product use exists in the description.

The project is described as a hackathon submission by one developer and lacks any indication of a viable business model or market readiness. It is positioned as an experimental research tool with no clear path to monetization or customer adoption.

Confidence level: Low. The description is self-reported, unverified, and lacks any commercial or technical validation.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.