OpenAI 2026 hackathon

Benchwork

ChatGPT tells you about biology. Benchwork does the biology, and shows its work.

Team of 2 · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #2,910 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Benchwork is a self-reported tool that allows users to ask biology questions in plain English and receive answers grounded in live biological databases. The system claims to trace every claim back to its source, distinguishing between directly measured facts (T1), inferred facts (T2), and gaps with proposed experiments (T3). It uses a structured knowledge graph built from APIs like Open Targets, ChEMBL, STRING, PubMed, and AlphaFold.

What changed

The project description indicates this is a hackathon submission (Devpost entry for OpenAI 2026 hackathon) and does not show evidence of commercial traction or product maturity beyond the prototype stage. It is described as a proof-of-concept with no revenue, customers, or deployment history.

The single most important open question

Is Benchwork's grounding mechanism reliable enough to be trusted in real research workflows? The description states that citation precision reaches 100% on live queries, but there are no independent validations or metrics about how often the system actually fails to ground claims properly.

Note

This analysis is based entirely on the self-reported project description provided by the caller. No external verification or historical data is available. All findings are drawn from that single source and labeled as such.

Back to contents

What The Product Actually Is

The description states that Benchwork:

  • Accepts biology questions in plain English
  • Plans queries, extracts intent, resolves terms to real database IDs
  • Fans out to live APIs (Open Targets, ChEMBL, STRING, PubMed, AlphaFold) in parallel with retries and graceful degradation
  • Verifies every extracted claim against raw JSON sources, dropping unsupported claims entirely
  • Builds a typed knowledge graph instead of disconnected results
  • Tiers claims as T1 (directly retrieved), T2 (inferred path stitched from T1 edges), or T3 (gap with proposed experiment)
  • Renders sourced briefs with clickable evidence network, bioactivity chart, and 3D structure viewer
  • Includes a "high mode" that runs path search algorithms over the graph to find mechanistic routes between proteins

The system is described as using FastAPI + asyncio backend, with a DAG orchestrator for concurrent tool calls. It uses SQLite for caching, NetworkX for graph representation, and React/TypeScript frontend with Cytoscape.js, Vega-Lite, and 3Dmol.js.

Claim

Benchwork is a biology research assistant that provides grounded answers from live databases.

Evidence Author's own write-up

Confidence Low — this is self-reported functionality without independent verification

Back to contents

Positioning & Claim Evolution

The description states that the project was inspired by frustration with manual cross-referencing of six or seven biology databases. The core positioning is:

  • ChatGPT tells you about biology, Benchwork does the biology and shows its work
  • Every fact comes with a receipt (database record)
  • System is honest about what's measured vs inferred vs unknown
  • Tier 3 (gap) includes concrete proposed experiments instead of vague disclaimers

The evolution of the claim appears to be:

  1. Start with frustration over manual database cross-referencing
  2. Build a system that answers questions from live databases
  3. Add grounding layers to distinguish T1, T2, and T3 claims
  4. Include visualization tools for evidence networks and 3D structures

Claim

Benchwork positions itself as a grounded alternative to general chatbots in biology.

Evidence Author's own write-up

Confidence Low — this is self-positioning, not market validation or customer feedback

Back to contents

Target Customer & ICP

The description states that the target user is:

  • Researchers doing early-stage drug target research
  • Who manually cross-reference multiple databases per question
  • Looking for answers that are grounded in actual biological data rather than memory-based responses

The ICP appears to be researchers who:

  • Work with biology databases (Open Targets, ChEMBL, STRING, PubMed, AlphaFold)
  • Need to avoid incorrect gene names or directions that could derail months of research
  • Want to see evidence for every claim made

Claim

Benchwork targets early-stage drug target researchers.

Evidence Author's own write-up

Confidence Low — no customer data or user interviews cited

Back to contents

Business Model & Pricing Evidence

Not evidenced.

Finding

No information provided about business model, pricing strategy, monetization approach, or revenue streams.

Back to contents

Technical & Delivery Signals

The system is described as:

  • Built with FastAPI + asyncio backend
  • Uses DAG orchestrator for parallel tool calls
  • Implements cache wrapping through SQLite
  • Uses NetworkX MultiDiGraph with provenance attached to edges
  • Verifies claims against raw JSON sources before tiering
  • Runs bounded k shortest paths search over induced subgraphs
  • Uses React/TypeScript/Vite frontend with Cytoscape.js, Vega-Lite, 3Dmol.js
  • Delivers real-time updates via SSE

Claim

Benchwork has a technical architecture designed for grounding and verification.

Evidence Author's own write-up

Confidence Low — this is self-reported architecture without independent validation

Back to contents

Traction & Maturity Signals

Not evidenced.

Finding

No evidence of customers, users, revenue, or product adoption beyond the hackathon submission. The project is described as a prototype with no commercial traction.

Back to contents

Competitive Context

The description does not provide any information about competitors or competitive landscape.

Finding

No competitive analysis or market positioning relative to existing tools in biology research or AI-assisted scientific discovery.

Back to contents

Key Risks & Red Flags

  • Grounding reliability: The system claims 100% citation precision on live queries, but the authors note that bugs only surfaced during end-to-end testing against real data, not mocked tests.
  • Technical debt: Several critical issues (druggability filter deleting verified evidence, STRING ID scheme mismatch) were found only after running real queries.
  • Model dependency: The system's trustworthiness depends heavily on model choice, which is described as an architectural decision rather than a minor implementation detail.
  • Limited scope: The project is described as a hackathon submission with no indication of long-term product development or commercial viability.

Inference The system may not be reliable enough for production use due to unverified claims about grounding accuracy and past technical issues discovered only in real-world testing.

Back to contents

Diligence Questions To Ask The Founders

  1. How does Benchwork handle conflicts between different databases when they disagree on a claim?
  2. What is the actual performance of the system under load? How long does it take to process complex queries?
  3. Can you provide examples of how often the grounding mechanism fails in practice, and what happens when it does?
  4. Are there any plans for integrating additional biological databases beyond those mentioned (e.g., UniProt, Reactome)?
  5. What are the limitations of high mode, and how do you plan to scale it beyond one protein pair?
  6. How do you ensure consistency across different LLMs used in various parts of the pipeline?
  7. Has there been any user feedback from researchers who have tried using Benchwork in real research workflows?

Note

These questions are based on the self-reported description and aim to probe deeper into unverified claims.

Back to contents

Investment/Partnership Verdict

Not evidenced.

Finding

No information provided about valuation, funding rounds, or investment interest. The project is described as a hackathon submission with no commercial traction or financial data available.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.