OpenAI 2026 hackathon

ObserveOS — The Self-Improving Clinic Operating System

Evidence before eloquence: a human-governed AI loop that keeps reports, observations, inferences, and unknowns separate—and blocks stale analysis from the reviewed record.

Solo project by paul800901 Huang · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,633 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

ObserveOS is a self-reported, human-governed AI loop system designed for longitudinal professional work in clinical settings. It aims to separate client reports, practitioner observations, AI-generated inferences, and unknowns while preventing stale analysis from influencing future decisions. The system uses append-only event ledgers, tamper-evident hash chains, and formal save gates to govern what information enters evidence.

What changed

During OpenAI Build Week, the project evolved from a private, evolving whole-practice workflow into a public synthetic-only implementation using Codex and GPT-5.6. The author states that this version demonstrates governance patterns without exposing proprietary clinical rules or private data.

The single most important open question

Is there evidence of real-world adoption or traction beyond the synthetic demo? The description makes no claims about revenue, customers, or usage beyond the self-contained demonstration.

Back to contents

What The Product Actually Is

The description states that ObserveOS is:

  • A local browser application and standard-library Python service
  • An append-only case event ledger with a tamper-evident hash chain
  • An evidence-only projection that excludes AI questions
  • A system that tracks practitioner-answer provenance, stale-analysis invalidation, and formal-save gates
  • A deterministic replay mechanism with no model, account, API key, or package-install dependency
  • A system that permits at most one bounded reflection question citing existing evidence events
  • A system that keeps supported findings, bounded inference, and explicit unknowns separate
  • A system that marks previous analysis stale when new evidence arrives

The product is described as a "self-improving" system but explicitly states this is human-governed, not autonomous self-modification.

Back to contents

Positioning & Claim Evolution

The description states that ObserveOS:

  • Optimizes for longitudinal professional work rather than the next answer
  • Separates client reports from practitioner observations and AI inferences
  • Maintains a clear distinction between evidence and analysis
  • Prevents "stale analysis" from influencing future conclusions
  • Uses an "evidence-governance core"
  • Is positioned as a "human-governed AI loop that keeps reports, observations, inferences, and unknowns separate"

The positioning evolved from a private, evolving whole-practice workflow to a public synthetic-only implementation during OpenAI Build Week.

Back to contents

Target Customer & ICP

The description states:

  • The system is designed for longitudinal professional work
  • It was inspired by "longitudinal professional work" that has a harder problem: information arrives over time from sources with different authority
  • The system handles "case review, intake, governed transcription, source-separated knowledge, operations readback, websites, campaigns, and content workflows"
  • The author mentions "practitioner-answer provenance" and "clinical decision logic base"
  • It is described as a "CaseAgent workflow" for clinical settings

The target customer appears to be practitioners in clinical or professional services who need to manage longitudinal information flows with different authority levels.

Back to contents

Business Model & Pricing Evidence

Not evidenced. The description makes no claims about pricing, revenue streams, or business model.

Back to contents

Technical & Delivery Signals

The description states:

  • Built with: codex, event-sourcing, gpt-5.6, human-in-the-loop, python
  • A local browser application and standard-library Python service
  • Append-only case event ledger with tamper-evident hash chain
  • Evidence-only projection excluding AI questions
  • Practitioner-answer provenance, stale-analysis invalidation, formal-save gates
  • Deterministic Replay with no model, account, API key, or package-install dependency
  • Optional live GPT-5.6 review through existing Codex ChatGPT sign-in
  • 47 automated contract tests, four-round gold replay, three-case synthetic governance corpus, privacy audit, JavaScript syntax verification
  • Uses event-sourcing with JSONL events, sequence numbers, idempotency keys, previous hashes, and current hashes
  • Evidence projection includes governed source events and practitioner answers while excluding AI questions
  • Supports bounded reflection questions that cite existing evidence event IDs
  • New evidence invalidates prior analysis
  • Save gate stores current normalized analysis without second model call
  • Live analysis launches codex exec through existing ChatGPT-authenticated Codex session
  • Replay implements same demonstrated evidence-state transitions without requiring model entitlement

Back to contents

Traction & Maturity Signals

Not evidenced. The description makes no claims about revenue, customers, or adoption beyond the synthetic demo.

Back to contents

Competitive Context

Not evidenced. The description does not mention competitors or market positioning beyond stating that most AI products optimize for the next answer rather than longitudinal professional work.

Back to contents

Key Risks & Red Flags

  • The system is described as a "synthetic-only and independently runnable public implementation" but the private operational systems are its lineage, not hidden dependencies required by the judge
  • The description states that "the public CaseAgent Reflection Loop is the runnable Build Week project. The private operational systems are its lineage and product context, not hidden dependencies required by the judge"
  • No claims about real-world adoption or traction beyond the synthetic demo
  • The system appears to be a prototype/demo rather than a production-ready solution
  • The author states that "No OpenAI API key was created for the application. Live mode reuses Codex ChatGPT sign-in; Replay makes no model call"
  • The description mentions "Known limits" including citation validation confirming referenced evidence event IDs exist but semantic support remains human-reviewed

Back to contents

Diligence Questions To Ask The Founders

  1. What is the actual business model for ObserveOS?
  2. How does the system handle real-world edge cases that aren't covered in the synthetic demo?
  3. What are the specific clinical or professional use cases where this system would be applied?
  4. How does the human-in-the-loop mechanism scale to multiple practitioners?
  5. What is the validation process for the clinical rules and decision logic that aren't published in the public version?
  6. How does the system handle integration with existing clinical workflows and systems?
  7. What are the actual requirements for practitioners to use this system?
  8. How does the system ensure data privacy and security in real-world deployments?

Back to contents

Investment/Partnership Verdict

Not evidenced. The description makes no claims about funding rounds, valuations, or investment status. The project appears to be a prototype/demo submitted to a hackathon rather than a commercial venture with demonstrated traction or revenue.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.