Archive position — measured, not model output
0 likes on Devpost
2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #5,722 place in the like-ranked listing is a tie-break inside that group, not a ranking.
Projects (log scale)
Likes on Devpost. ▲ marks this project's group.
Show the figures
| Likes | Projects | Share of archive |
|---|---|---|
| 0 | 5,592 | 71.2% |
| 1 | 1,758 | 22.4% |
| 2 | 285 | 3.6% |
| 3–4 | 132 | 1.7% |
| 5–9 | 75 | 1.0% |
| 10+ | 14 | 0.2% |
Executive Summary
What the company appears to be
Operandi is a self-reported AI-powered system that claims to learn real business workflows from messy artifacts (e.g., spreadsheets, email threads), verify those workflows against historical data, and generate runnable automation tools — all without relying on model-generated figures for decision-making. It uses structured outputs from an LLM (GPT-5.6) to extract a workflow graph, then applies deterministic code to validate or conflict with stated rules.
What changed
The project description reflects a shift from a general AI pitch toward a more specific focus on proving workflows through data rather than trusting models. This is evidenced in its core claim: “can an AI learn how a business already works from its own messy artifacts, prove what it learned against that business's real history, and hand back a working automation — all in one sitting?”
Single most important open question
Does Operandi actually deliver on its promise of deterministic verification of business rules via structured LLM outputs and code-based validation? The author states this is the core innovation, but there is no evidence of external testing, customer feedback, or real-world deployment.
What The Product Actually Is
The description states that Operandi:
- Learns workflows from real files (e.g., spreadsheets + email threads)
- Extracts a structured Operational Graph tagging nodes as DECLARED vs OBSERVED
- Verifies rules against historical data using deterministic code
- Generates runnable OpenAI function-tools with auto-generated tests
- Recommends automation candidates based on effort × impact
It is built using:
- Codex for scaffolding and test mass
- GPT-5.6 (OpenAI Responses API) as reasoning layer
- TypeScript monorepo (17 packages, Node 24)
- Docker, Fastify, Next.js, PostgreSQL, Supabase, OpenAI APIs
The architecture enforces a strict rule: the LLM proposes, deterministic code proves.
Inference The product appears to be a proof-of-concept or prototype built for a hackathon. It is not described as having any live clients or production systems in use.
Positioning & Claim Evolution
The author states:
- "We wanted to invert the usual AI pitch" — instead of trusting future AI, they ask: can an AI learn how a business already works from its own messy artifacts?
- The key word is “prove” — not plausible summaries but numbers that trace back to business data.
- They emphasize that “evidence is the product.”
Inference The positioning has evolved from a generic AI tool to one focused on verifiable workflow learning and automation, with an emphasis on data-driven trust over model trust.
There is no evidence of prior versions or market positioning beyond this self-reported narrative.
Target Customer & ICP
The description does not name specific customer types or personas. However, it implies:
- Businesses that run on systems nobody fully understands
- Teams with messy artifacts (spreadsheets, email threads) describing processes
- Organizations looking to automate workflows but lacking clear documentation
Inference The target is likely small-to-medium-sized businesses or internal teams within larger enterprises who want to understand and automate their existing operations.
No evidence of customer segments, personas, or buyer journeys.
Business Model & Pricing Evidence
There is no mention of pricing, monetization strategy, or business model in the description. The project is presented as a hackathon submission with no indication of commercial intent beyond the idea itself.
Inference No evidence of a defined business model or pricing structure.
Technical & Delivery Signals
The system uses:
- Codex for scaffolding and test generation
- GPT-5.6 (OpenAI Responses API) for structured output extraction
- TypeScript monorepo with 17 packages
- Docker, Fastify, Next.js, Node.js, PostgreSQL, Supabase
- Deterministic verification via code-based engines
- Skill sandboxing with typed schemas and auto-generated tests
Key technical features:
- Operational Graph tagging (DECLARED vs OBSERVED)
- Rule IR engine for replaying rules against real data
- Independent verifier to recompute answers
- Action Gateway enforcing approval before execution
- Provenance tracking, pinned oracles, and typed blockers
Inference The architecture is designed with strong guardrails around trust boundaries — the LLM does not produce actionable figures; only code can.
Traction & Maturity Signals
The description states:
- It was built for a hackathon (OpenAI 2026)
- Runs fully offline in demo mode
- Has no live clients or real-world deployments
- No revenue, customers, or adoption data provided
Inference There is no evidence of traction, maturity, or commercial use beyond the author’s own demonstration.
Competitive Context
The description does not reference any competitors. It focuses on its unique approach to proving workflows rather than comparing itself to existing tools in the space.
Inference No competitive analysis or positioning against other workflow automation or AI tools is evident.
Key Risks & Red Flags
- Unproven architecture: The claim that "the LLM proposes, deterministic code proves" has not been validated outside of a demo.
- No external validation: There is no evidence of real-world testing, customer feedback, or live deployment.
- Hackathon prototype: The project was submitted to a hackathon — suggesting it may be early-stage and untested in production.
- Limited scope: The author explicitly rejects breadth for depth, which could limit scalability or commercial viability.
- Dependency on proprietary APIs: Heavy reliance on OpenAI, Composio, Supabase, etc., may pose risks if those services change or become unavailable.
Diligence Questions To Ask The Founders
- How does the system handle edge cases where data is incomplete or contradictory?
- What happens when a rule cannot be verified due to lack of data? Are there failure states?
- Has the deterministic verification been tested with real business datasets beyond the demo?
- Is there any plan for integrating with actual enterprise systems (e.g., ERP, CRM)?
- How does the system scale across multiple workflows or large organizations?
- What is the expected time investment from users to set up and run a workflow?
- Are there plans to support different data formats beyond spreadsheets and email threads?
Investment/Partnership Verdict
Not evidenced.
The description presents Operandi as a hackathon prototype with strong technical design but no evidence of traction, revenue, or commercial viability. The author claims the system proves workflows through deterministic code, but there is no independent verification or demonstration of this in practice.
Confidence level Low — based entirely on self-reported narrative and no external validation.
Verdict Early-stage concept with promising architecture; not ready for investment or partnership without further evidence of real-world testing, customer engagement, or product-market fit.
Source
Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.
The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.

