OpenAI 2026 hackathon

Evidence Workbench

Offline, evidence-first search that turns fuzzy technical questions into exact PDF pages, tables, and diagrams.

Team of 2 · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,990 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

Evidence Workbench is a self-reported technical search tool designed for internal engineering teams to find exact content within large PDF libraries. It claims to support offline, evidence-first search that answers ambiguous questions with precise page-level results and direct links to source documents.

What changed

The project was submitted as part of the OpenAI 2026 hackathon. The description indicates it is a proof-of-concept built in a short timeframe by two team members using AI-assisted coding tools like Codex and GPT-5.6, with no evidence of prior traction or commercial deployment.

Single most important open question

Is there any evidence that this tool has been validated in an actual engineering environment beyond the hackathon demo? The description states it is a prototype built for a hackathon and does not indicate whether it has moved past the experimental stage or been adopted by users.

Back to contents

What The Product Actually Is

The description states that Evidence Workbench is an offline, evidence-first search experience for technical PDF libraries. It allows users to describe a model, symptom, parameter, table, or diagram without knowing the source document and returns results with exact physical page numbers, inspectable crops, and direct links to the source PDF.

It supports:

  • Scanning existing PDF trees at startup
  • Manual full-library refreshes that process only new or changed PDFs
  • A bilingual English/Chinese UI
  • A local feedback journal with an explicit human-review workflow

The system is built using:

  • BGE-M3 dense+sparse retrieval
  • Reranking
  • Qdrant
  • PostgreSQL
  • Docker
  • FastAPI
  • pdfplumber
  • Ubuntu

It uses a zero-egress model loading approach, meaning models are not sent outside the local server.

Inference The tool appears to be a search engine tailored for technical documentation stored in PDF format, with a focus on precision and traceability of results. It is not described as a general-purpose AI assistant or chatbot.

Back to contents

Positioning & Claim Evolution

The description states that Evidence Workbench was inspired by a real-world engineering problem: the inefficiency of searching large confidential libraries for specific technical content. The authors claim it solves issues with existing RAG workflows being too sensitive to prompt wording and failing on fuzzy questions.

Key claims:

  • It turns fuzzy technical questions into exact PDF pages, tables, and diagrams
  • It provides scoped answers with physical page numbers
  • It returns inspectable crops of content
  • It supports ambiguous queries by offering one necessary clarification instead of a guess

These claims are framed as improvements over current RAG systems, which the authors describe as failing on vague or contextual questions.

Inference The positioning is that of a specialized, precision-focused search tool for technical documentation, not a general-purpose AI assistant. It positions itself as solving a gap in how engineers access and retrieve knowledge from large internal libraries.

Back to contents

Target Customer & ICP

The description states the target user is an engineering team working with large technical-PDF libraries. The system is designed to be deployed on an approved internal server, suggesting it targets enterprise or internal engineering environments where confidentiality and control are paramount.

It is built for:

  • Application Engineers
  • Teams supporting AI server bring-up, deployments, and issue triage

The authors note that the tool is intended for a one-time deployment on an internal server, implying a small to medium-sized team or department, not a broad enterprise rollout.

Inference The ICP appears to be internal engineering teams in enterprises with large technical libraries, particularly those needing precise search and retrieval of PDF-based documentation.

Back to contents

Business Model & Pricing Evidence

Not evidenced.

The description does not contain any information about pricing, licensing, monetization, or business model.

Back to contents

Technical & Delivery Signals

The system is described as:

  • Offline-first
  • Built with zero-egress model loading
  • Uses BGE-M3, bge-reranker-v2-m3, Codex, GPT-5.6, Qdrant, PostgreSQL, Docker, FastAPI, pdfplumber, Ubuntu
  • Supports manual full-library refreshes that process only new or changed PDFs
  • Includes a bilingual English/Chinese UI
  • Has a local feedback journal with human review workflow

The authors also mention:

  • A public demo using SHA-256-pinned public PDFs and lightweight deterministic retrieval
  • 115 passing automated tests in the public release
  • Use of Playwright journeys, privacy checks, and human review

Inference The tool is built with a strong emphasis on security, traceability, and reproducibility, using open-source or public tools where possible. It is not described as cloud-based or SaaS.

Back to contents

Traction & Maturity Signals

Not evidenced.

There is no mention of:

  • Revenue
  • Customers
  • Adoption
  • Usage metrics
  • Product-market fit
  • Prior deployments or internal use beyond the hackathon

The project is explicitly described as a hackathon submission, and no evidence suggests it has moved beyond prototype status.

Back to contents

Competitive Context

Not evidenced.

There is no mention of:

  • Competitors
  • Market analysis
  • Differentiation from existing tools
  • Industry benchmarks

The description does not reference any existing search or RAG tools in the market, nor does it compare Evidence Workbench to them.

Back to contents

Key Risks & Red Flags

  1. No traction or commercial validation: The tool is described as a hackathon project with no evidence of adoption or revenue.
  2. Limited scope: It is built for internal use only and not described as a SaaS or cloud-based solution.
  3. Unproven scalability: The system is designed for one-time deployment on an internal server, suggesting it may not be scalable to larger teams or organizations.
  4. AI-assisted development dependency: Heavy reliance on Codex and GPT-5.6 in development raises questions about whether the tool can be maintained without these tools or if it’s overly dependent on AI-generated code.
  5. No privacy or compliance evidence: While the authors mention privacy boundaries, there is no evidence of formal compliance or audit processes.

Back to contents

Diligence Questions To Ask The Founders

  1. Has this tool been tested in a real engineering environment beyond the hackathon?
  2. What are the actual use cases and workflows it supports in practice?
  3. Is there any plan to move beyond internal, one-time deployments to broader enterprise adoption?
  4. How is the system maintained or updated without relying on AI-assisted coding tools?
  5. Are there any plans for monetization or commercialization?
  6. What are the limitations of the current architecture in terms of scalability and performance?

Back to contents

Investment/Partnership Verdict

Not evidenced.

The description does not provide sufficient evidence to assess whether this project is ready for investment or partnership. It is a self-reported hackathon prototype with no demonstrated traction, revenue, or market validation. The tool’s focus on internal engineering teams and offline deployment suggests it may be too niche or early-stage for general investment interest.

The authors describe the system as built in a short timeframe using AI-assisted tools, but there is no indication of long-term viability, scalability, or commercial potential beyond its initial use case.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.