OpenAI 2026 hackathon

Science Twin

A compiler for supported physiological models: GPT-5.6 plans a user-approved experiment, while deterministic code runs and verifies the evidence.

Solo project by Andrei Vince · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #6,572 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Science Twin is a developer tool that compiles supported physiological models into reproducible simulations using GPT-5.6 for experiment planning and deterministic code for execution. It aims to improve scientific reproducibility by generating traceable, user-approved experiments with evidence bundles.

What changed

The project description indicates a proof-of-concept build week effort focused on a toy cellular mechanism. It demonstrates how GPT-5.6 can interpret plain English questions and generate executable code for simulation within a sandboxed environment, while maintaining an immutable evidence graph.

Single most important open question

Is there sufficient scientific or technical validation to support the claim that this tool can reliably reproduce complex physiological models beyond a toy example?

Back to contents

What The Product Actually Is

The description states that Science Twin is a developer tool for supported physiological models, where users describe a bounded experiment in plain English. GPT-5.6 interprets the request through constrained tools and prepares an immutable draft, which is then executed deterministically by code.

Key components include:

  • A CellML source importer
  • Rust-based simulation pipeline
  • Docker sandboxed execution
  • Evidence comparison against numerical references
  • Five-axis qualification report generation

The system uses GPT-5.6 via the OpenAI Agents SDK for planning and diagnostics, but not for authoring equations or choosing parameters.

Not evidenced: whether this is a standalone product, SaaS offering, or internal tool; no mention of UI/UX beyond browser workflow.

Back to contents

Positioning & Claim Evolution

The description states that Science Twin turns “reproducing a scientific result” into a developer workflow. It positions itself as a tool for computational model developers who want to ensure their simulations are traceable and rerunnable.

It claims to:

  • Turn plain English questions into inspectable experiments
  • Provide user-approved, deterministic execution
  • Generate reproducibility packages

The author also notes that the tool separates implementation fidelity from biological validation or regulatory context — implying a focus on technical correctness over clinical utility.

Inferred: The tool is positioned as a reproducibility engine, not a full modeling platform or clinical decision support system.

Not evidenced: No claims about market positioning, competitive differentiation, or target adoption rate.

Back to contents

Target Customer & ICP

The description states that Science Twin targets computational model developers who need to know whether their implementation can be traced, rerun, and checked against independent references.

It is implied that these users are likely:

  • Researchers working with physiological models
  • Developers in computational biology or biomedical engineering
  • Teams needing reproducible simulation workflows

Not evidenced: No explicit customer segments, personas, or use cases beyond the toy example. No indication of whether the tool is aimed at academic, industry, or regulatory users.

Back to contents

Business Model & Pricing Evidence

The description does not provide any information about:

  • Revenue streams
  • Pricing model
  • Monetization strategy
  • Customer acquisition plans

Not evidenced: No evidence of a business model or pricing structure.

Back to contents

Technical & Delivery Signals

The system uses:

  • Rust for simulation pipeline
  • CellML and libCellML for model import
  • Docker for sandboxed execution
  • GPT-5.6 via OpenAI Agents SDK for planning and diagnostics
  • React, TypeScript, Vite, Bun, SQLite, OpenAPI, Zod for frontend and backend

The build week proof used a toy cellular gate-recovery mechanism with:

  • 21 of 21 comparison points matching a closed-form reference
  • 35 immutable artifacts
  • Evidence graph tied to exact completed runs

Inferred: The tool is built on a reproducible, deterministic architecture, with strong emphasis on traceability and immutability.

Not evidenced: No details about scalability, deployment infrastructure, or long-term maintainability beyond the proof-of-concept.

Back to contents

Traction & Maturity Signals

The description states that this is a Build Week proof from an OpenAI hackathon. It includes:

  • A working prototype
  • A toy example with numerical fidelity
  • 35 immutable artifacts
  • Evidence graph tied to execution

Not evidenced: No customer data, revenue, or usage metrics. No indication of product-market fit or user feedback.

Back to contents

Competitive Context

The description does not mention any competitors or market context.

Not evidenced: No competitive landscape, pricing, or positioning relative to other tools in computational modeling or reproducibility.

Back to contents

Key Risks & Red Flags

  • Unverified claims: The tool is described as proving numerical implementation fidelity for one bounded cellular mechanism, but not biological validation or clinical utility.
  • Limited scope: Only a toy example was tested; no indication of broader applicability.
  • Dependency on GPT-5.6: While constrained, the use of an AI model raises questions about consistency and control over outputs.
  • No commercialization plan: No mention of monetization, customer acquisition, or product roadmap beyond the proof-of-concept.

Inferred: The tool is in a very early stage, with no clear path to market traction or commercial viability.

Back to contents

Diligence Questions To Ask The Founders

  1. What are the specific constraints and limitations of GPT-5.6 integration? Can it be replaced or extended?
  2. How does Science Twin handle model versioning and updates in a production environment?
  3. Are there any plans to expand beyond toy examples into more complex physiological models?
  4. What is the long-term vision for the tool — is it intended as a standalone product or part of a larger ecosystem?
  5. Has the team considered regulatory or compliance implications of using AI-assisted scientific tools?

Back to contents

Investment/Partnership Verdict

The description indicates that Science Twin is an early-stage proof-of-concept built during a hackathon. It shows technical capability in a limited domain but lacks evidence of traction, revenue, or commercial viability.

Confidence level: Low — based on self-reported, unverified information only.

Not evidenced: No indication of market demand, customer validation, or scalability beyond the toy example.

Inferred: This is a research-grade prototype, not a product ready for investment or partnership.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.