OpenAI 2026 hackathon

Sherlock

Most AI tries to find the right answer. Sherlock eliminates the wrong ones.

Solo project by aiwhisperer11 Linares · 4 likes · 0 comments

Archive position — measured, not model output

4 likes on Devpost

89 of the 7,856 archived projects have more likes, and 39 share exactly 4 — so this project's #120 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Sherlock is a self-reported AI investigation framework built for problem-solving by ruling out incorrect explanations rather than generating plausible answers. The author describes it as an end-to-end product that uses GPT-5.6 Structured Outputs and OpenAI Codex, with no hand-written code.

What changed

The project evolved through four implementation blocks, from a canonical schema to iterative reasoning, using structured outputs and semantic evaluation. It was submitted to the OpenAI 2026 hackathon.

Single most important open question

Is Sherlock a working investigation framework or an unvalidated idea? The description states it is built with GPT-5.6 and Codex but provides no evidence of real-world use, performance metrics, or customer feedback.

Back to contents

What The Product Actually Is

The description states that Sherlock is an "investigation framework" designed to make large language models (LLMs) think like investigators instead of assistants. It builds structured investigations by:

  • Classifying observations as expected or unexpected.
  • Maintaining competing hypotheses simultaneously.
  • Tracking rejected explanations in a "Hypothesis Graveyard."
  • Ranking missing evidence by value.
  • Recommending the next test to distinguish between leading hypotheses.
  • Supporting iterative learning where new evidence updates existing investigations.

The demo case involves a software incident (HTTP 500 errors after deployment), where Sherlock identifies that a TLS certificate renewal log should have appeared but did not, instead of blaming the deployment.

Evidence The author describes how it works in detail, including its components and reasoning process. However, there is no evidence of actual implementation beyond the self-reported narrative.

Inference Based on the description, Sherlock appears to be a structured AI tool for hypothesis falsification and evidence-based reasoning, not a general-purpose assistant or chatbot.

Back to contents

Positioning & Claim Evolution

The author positions Sherlock as an alternative to traditional AI tools that "try to find the right answer" by instead eliminating wrong ones. It draws analogies from Sherlock Holmes, differential diagnosis, and Karl Popper’s falsification theory.

Sherlock is described not as another prompting technique but as a "reasoning workflow" that constrains how models investigate problems. The goal is not to generate better answers but to make the model "earn the answer."

Evidence The author explicitly contrasts Sherlock with typical AI tools and positions it within philosophical frameworks of reasoning and falsification.

Inference This suggests a niche positioning for complex problem-solving, particularly in domains where missing evidence is critical — such as debugging, compliance, or forensic analysis.

Back to contents

Target Customer & ICP

The description does not name specific customers or personas. It mentions that the demo case is one many engineers recognize (a software incident), implying an engineering audience might be relevant.

However, the author states that Sherlock is intentionally domain-independent and can be applied beyond software incidents.

Evidence No explicit customer segments or ideal customer profiles are mentioned.

Inference The target market may include technical professionals who need structured reasoning for debugging, compliance audits, or investigative tasks. But there's no evidence of actual users or buyer personas.

Back to contents

Business Model & Pricing Evidence

There is no mention of pricing, monetization strategy, or business model in the description.

Evidence Not evidenced.

Inference Since this is a hackathon project submitted by one person and not described as commercialized, it likely has no revenue model at present. Any future business model would be speculative without further information.

Back to contents

Technical & Delivery Signals

The system was built using:

  • GPT-5.6 Structured Outputs
  • OpenAI Codex for implementation
  • Next.js, React, TypeScript, Tailwind CSS, Vercel
  • JSON Schema validation via AJV
  • A versioned investigation workflow with explicit rules

Development followed disciplined blocks with acceptance criteria and governance.

Evidence The author lists technologies used and describes development stages.

Inference The technical stack suggests a modern web-based AI application. The use of structured outputs and schema validation indicates attention to reasoning quality and auditability.

Back to contents

Traction & Maturity Signals

There is no evidence of traction, revenue, customers, or adoption beyond the author’s own account.

The project was submitted as part of a hackathon and includes a demo case but lacks any data on usage, performance, or impact.

Evidence Not evidenced.

Inference This is a prototype or proof-of-concept, not a mature product with real-world deployment or user engagement.

Back to contents

Competitive Context

The description does not reference competitors directly. However, it implies that current AI tools focus on generating explanations rather than ruling out incorrect ones.

It positions itself against general-purpose LLM assistants and suggests a unique approach to reasoning and hypothesis testing.

Evidence No mention of existing products or competitive landscape.

Inference If Sherlock gains traction, it could compete with AI-powered debugging tools, incident response platforms, or investigative software. But no such competitors are named or described.

Back to contents

Key Risks & Red Flags

  • Unvalidated assumptions: The system relies heavily on the author’s reasoning framework and model behavior without empirical validation.
  • Limited scope: It is described as a hackathon project with no evidence of scalability or real-world application.
  • No performance metrics: There are no benchmarks, accuracy claims, or evaluation results provided.
  • Single contributor: The team size is listed as one person, raising questions about long-term development and maintenance.
  • Unclear commercial viability: No indication of monetization strategy or path to market.

Evidence These risks stem from the lack of external validation, real-world usage, and business planning in the description.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific problems have you observed where traditional AI tools failed to help?
  2. How do you validate that Sherlock’s reasoning is sound and not just a clever narrative?
  3. Have you tested Sherlock with real users or teams? If so, what feedback did they give?
  4. What are the limitations of GPT-5.6 Structured Outputs in this context?
  5. Is there any plan to move beyond the hackathon prototype into a scalable product?
  6. How does Sherlock handle ambiguity when evidence is incomplete or contradictory?
  7. What kind of domain-specific customization will be needed for different use cases?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The description provides no data on financials, traction, or market opportunity that would support an investment or partnership decision.

This appears to be a hackathon submission with a compelling idea but no demonstrated product-market fit, revenue, or customer base.

It may represent a promising concept for further development, but there is insufficient evidence to assess its commercial potential or risk profile.

Confidence level Low. The project is described as a prototype with no external validation or performance data.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.