OpenAI 2026 hackathon

ResearchOS

From a question to a verified, connected, replayable research record — AI proposes, you verify, and every claim shows its source.

Hackathon project · 2 likes · 0 comments

Archive position — measured, not model output

2 likes on Devpost

221 of the 7,856 archived projects have more likes, and 285 share exactly 2 — so this project's #442 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

ResearchOS, as described by its author, is a structured research tool that uses AI to process scientific literature, extract claims, and build verifiable knowledge graphs. It enforces human verification of AI-generated outputs and supports hypothesis formation, experiment planning, and research replay. The system is built around GPT-5.6 with domain-specific constraints and integrates with physics simulation tools like MACE.

What changed

The author describes a product that moves away from general-purpose AI chat to a structured, human-in-the-loop workflow where AI proposes outputs (amber), and humans verify them (green). It introduces a formalized research process grounded in provenance and evidence-based reasoning.

Single most important open question

Is there any evidence of real-world usage or adoption by researchers? The description is entirely self-reported and lacks any indication of traction, revenue, or customer data.

Back to contents

What The Product Actually Is

The description states that ResearchOS is a system that turns a research question into a structured, verifiable record across five surfaces:

  1. Literature — GPT-5.6 reads uploaded papers and extracts claims, each pinned to an exact quoted span from the source.
  2. Knowledge Graph — verified claims become an evidence graph of materials and properties; numeric disagreements surface as red dashed contradiction edges.
  3. Hypotheses — drafted only from verified evidence, and required to be measurable by an experiment the system can actually run.
  4. Experiment Planner — runs real physics: an ASE + MACE-MP-0 substitution-energy screen on CPU, not a mock.
  5. Research Replay — because every action is an event, the entire project replays from the first question, with provenance behind each step.

The core rule is that AI proposes (amber), humans verify (green), and only verified evidence counts. There is no free-text chat window; the model produces bounded, structured proposals.

Evidence

  • The author states this is how the product works.
  • It uses GPT-5.6 across multiple tiers: gpt-5.6-terra for extraction, gpt-5.6-sol for hypothesis drafting and plan-filling, gpt-5.6-luna for entailment checks.

Inference This is a research workflow automation tool with strong emphasis on verifiability and reproducibility through structured AI outputs and human validation.

Back to contents

Positioning & Claim Evolution

The author positions ResearchOS as a system that moves beyond AI fluency to focus on provenance and verification, where trust comes from traceability, not just language generation. The product is described as:

  • Not a chatbot or general-purpose AI assistant.
  • A tool for structured research workflows with human-in-the-loop validation.
  • Designed to support scientific rigor in hypothesis formation, experiment planning, and result interpretation.

Evidence

  • “Trust in an AI research tool comes from provenance and verification, not fluency.”
  • “Amber-to-green with a human in the loop is the whole product.”

Inference The positioning has evolved from a generic AI-powered research assistant to a structured, evidence-based research platform that emphasizes reproducibility and scientific rigor.

Back to contents

Target Customer & ICP

The description does not explicitly state who the target customer or ideal customer profile (ICP) is. It implies use by researchers working in domains like materials science, particularly hydrogen storage, but no names, roles, or organizational affiliations are mentioned.

Evidence

  • The demo uses a hydrogen-storage domain pack.
  • The system supports real physics simulations and conflict detection in scientific domains.

Inference The likely ICP includes researchers or research teams working in materials science or related fields, especially those needing structured workflows for hypothesis testing and reproducibility.

Back to contents

Business Model & Pricing Evidence

No evidence of a business model or pricing structure is provided. The description does not mention monetization, licensing, subscriptions, or any commercial arrangement.

Evidence

  • No mention of revenue, customers, or pricing.
  • The project was submitted to a hackathon and is described as self-built by one person.

Inference There is no evidence of a business model or pricing strategy. This is likely an early-stage prototype or proof-of-concept.

Back to contents

Technical & Delivery Signals

The system is built using:

  • AI stack: GPT-5.6 (with multiple specialized models), OpenAI API, Codex.
  • Backend: FastAPI with async SQLAlchemy, PostgreSQL 16 + pgvector, Redis with arq for task queuing.
  • Frontend: Next.js with sigma.js graph and SSE event stream.
  • Compute: MACE-MP-0 substitution-energy screen on CPU using ASE.
  • Domain knowledge: Encoded in a “hydrogen-storage domain pack” that includes property registry, unit aliases, comparability transforms, and conflict rules.

Evidence

  • The author describes the tech stack and how it was used to build the product.
  • Specific tools like Docker, React, PyTorch, Python, TypeScript are mentioned.

Inference The system is built with a modern, scalable stack for AI + data processing and scientific simulation. It shows technical maturity in handling structured outputs, grounding, and domain-specific logic.

Back to contents

Traction & Maturity Signals

There is no evidence of traction or adoption beyond the author’s own development. The team size is listed as 0, and there are no mentions of customers, users, revenue, or usage metrics.

Evidence

  • Team size: 0.
  • No mention of customers, users, or revenue.
  • Submitted to a hackathon — not an indication of commercial traction.

Inference This is an early-stage prototype or proof-of-concept. There is no evidence of product-market fit, user feedback, or real-world deployment.

Back to contents

Competitive Context

The description does not mention any competitors or direct market comparisons. It focuses on the unique aspects of ResearchOS (human-in-the-loop, grounding, conflict detection) but does not situate itself in a competitive landscape.

Evidence

  • No mention of existing tools or platforms in scientific research or AI-assisted R&D.
  • The author does not reference competitors or similar products.

Inference There is no evidence of competitive positioning. It’s unclear whether this product addresses an unmet need or competes with existing tools in the space.

Back to contents

Key Risks & Red Flags

  1. No traction or adoption: The project is described as self-built by one person, with no team or users.
  2. Unproven business model: No evidence of monetization or revenue streams.
  3. High technical complexity without real-world validation: The system uses advanced AI and scientific simulation but lacks real-world testing or feedback.
  4. Self-reported only: All claims are unverified; there is no third-party corroboration.
  5. No domain-specific commercial use case: While it includes a hydrogen-storage pack, there’s no indication of broader application or market demand.

Evidence

  • Team size: 0.
  • No revenue, customers, or usage data.
  • Submitted to a hackathon — not indicative of commercial viability.

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual research domain you're targeting beyond hydrogen storage?
  2. Have you tested this with real researchers or labs? What feedback did you get?
  3. How do you plan to scale the human verification process for larger datasets?
  4. Are there any plans for monetization or commercial partnerships?
  5. What are the key technical challenges that remain unresolved in production use?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The description is entirely self-reported and lacks any evidence of traction, revenue, customers, or a clear business model. The project appears to be an early-stage prototype or hackathon submission with no indication of commercial viability or market readiness.

Confidence Low. This analysis is based solely on the author’s own account — there is no independent verification, no third-party data, and no evidence of adoption or revenue.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.