OpenAI 2026 hackathon

evidence-wiki

Verifiable research workspaces where autonomous agents cite every claim—or clearly request the evidence they still need.

Solo project by Denis Sivagin · 0 likes · 0 comments

Archive position — measured, not model output

0 likes on Devpost

2,264 of the 7,856 archived projects have more likes, and 5,592 share exactly 0 — so this project's #3,997 place in the like-ranked listing is a tie-break inside that group, not a ranking.

Projects (log scale)

1
10
100
1k
10k
05,592
11,758
2285
3–4132
5–975
10+14

Likes on Devpost. ▲ marks this project's group.

Show the figures
LikesProjectsShare of archive
05,59271.2%
11,75822.4%
22853.6%
3–41321.7%
5–9751.0%
10+140.2%
Devpost like counts for all 7,856 archived projects, captured when this archive was built.

Executive Summary

What the company appears to be

The author describes EvidenceWiki as a research workspace for autonomous agents that cites every claim or clearly requests missing evidence. It is built to support evidence-based decision-making in software development and other domains, using structured knowledge and validation rules.

What changed

This project evolved from an earlier tool (agent-wiki-cli / python-wiki-llm) focused on codebase documentation into a general-purpose research agent with capabilities for automated workflow coordination, testing, quality assessment, and decision-making. The author states this is their first project where nearly the entire development lifecycle was carried out with active LLM involvement.

The single most important open question

Is there any evidence of traction, revenue, or customer adoption beyond the author's own account? The description does not indicate whether EvidenceWiki has been used by others, tested in real-world settings, or integrated into existing systems.

Note: This analysis is based solely on the self-reported project description provided by the author. No external verification, archived data, or third-party sources were used. All claims are attributed to the author's own account and should be treated as unverified.

Back to contents

What The Product Actually Is

The description states that EvidenceWiki is a research workspace for autonomous agents designed to support evidence-based workflows. It turns codebases into structured context that agents can understand, and it supports research in any domain where answers need to be grounded in traceable evidence—such as science, policy, legal guidance, product decisions, and software development.

It aims to:

  • Evaluate available evidence
  • Record provenance
  • Estimate confidence
  • Return results in a structured, machine-readable format
  • Explicitly block questions when evidence is insufficient and describe what information is still needed

The system treats external content as untrusted data and applies fail-closed acquisition and validation rules.

Claim: EvidenceWiki is an autonomous research agent with structured knowledge base capabilities.

Evidence: The author's own write-up.

Back to contents

Positioning & Claim Evolution

The author positions EvidenceWiki as a tool for autonomous agents that can perform research, validate outputs, and make decisions based on verifiable evidence. It evolved from earlier tools focused on codebase documentation to a more general-purpose system supporting automated workflows across domains.

Key claims include:

  • The system supports both software development and broader research tasks.
  • It enables agents to investigate codebases or external sources, compare evidence, and pass reliable findings.
  • It is not meant to replace human judgment but to provide a more reliable way to research, validate, document, and make decisions using inspectable evidence.

Claim: EvidenceWiki supports autonomous workflows in multiple domains.

Evidence: The author's own write-up.

Back to contents

Target Customer & ICP

The description does not clearly define target customers or ideal customer profiles (ICP). However, it implies that the system is useful for:

  • Software development teams using LLMs
  • Researchers working with evidence-based decision-making
  • Organizations seeking to automate workflows involving documentation and validation

It also suggests potential users include agents that interact with structured knowledge bases, as well as human reviewers who want to inspect and reuse findings.

Claim: The system targets software developers, researchers, and automated systems needing reliable evidence.

Evidence: Inferred from the author's description; not explicitly stated.

Back to contents

Business Model & Pricing Evidence

There is no mention of pricing models, monetization strategies, or business model in the description. The project appears to be a personal initiative by one developer (Denis Sivagin), with no indication of commercialization plans or revenue streams.

Claim: No evidence of business model or pricing.

Evidence: Not evidenced.

Back to contents

Technical & Delivery Signals

The author reports building EvidenceWiki using:

  • Codex
  • Python
  • VSCode

It was developed as a local, general-purpose research agent that uses wiki pages and collected sources as a structured knowledge base.

Key technical features include:

  • Structured processing of evidence
  • Provenance tracking
  • Confidence scoring
  • Handling incomplete or changing evidence
  • Coordination of multiple agents safely through deterministic workflows
  • Fail-closed acquisition and validation rules

Claim: EvidenceWiki uses Python, Codex, and VSCode; supports structured knowledge and agent coordination.

Evidence: The author's own write-up.

Back to contents

Traction & Maturity Signals

There is no evidence of traction, customers, or adoption beyond the author’s own account. The project was submitted to a hackathon (OpenAI 2026), but there are no indications of usage by others, integration into products, or measurable impact.

Claim: No evidence of traction or user base.

Evidence: Not evidenced.

Back to contents

Competitive Context

The description does not reference competitors or existing solutions in the space. It focuses on the novelty and functionality of EvidenceWiki rather than its place within a competitive landscape.

Claim: No competitive context provided.

Evidence: Not evidenced.

Back to contents

Key Risks & Red Flags

  • Lack of traction or adoption: The project is described only as a personal initiative with no evidence of real-world use.
  • Unproven scalability: While the author mentions coordination of autonomous agents, there’s no indication of how this scales beyond a single developer environment.
  • Limited validation: No mention of testing in production environments or feedback from users.
  • Unclear commercial viability: No business model or monetization strategy is described.

Inference: These risks arise from the lack of external validation and evidence of real-world application.

Evidence: Not evidenced.

Back to contents

Diligence Questions To Ask The Founders

  1. What specific use cases have you tested EvidenceWiki in, and how did those tests go?
  2. Have you attempted to integrate EvidenceWiki into any existing workflows or systems?
  3. How do you plan to scale the system beyond a single developer environment?
  4. Are there any early adopters or partners interested in using this tool?
  5. What are your thoughts on building a sustainable business around this concept?

Note: These questions are intended to probe for evidence that may not be present in the current description.

Back to contents

Investment/Partnership Verdict

There is insufficient evidence to assess whether EvidenceWiki represents a viable investment or partnership opportunity at this stage. The project appears to be an experimental tool developed by one individual, with no demonstrated traction, revenue, or customer base. While the concept shows promise in aligning with current trends in LLMs and autonomous systems, there is no indication that it has moved beyond prototype status.

Inference: This project lacks commercial readiness based on the provided description.

Evidence: Not evidenced.

Back to contents

Source

Submitted to the OpenAI 2026 hackathon on Devpost. Project home on DevPost.

The analysis above was generated by a language model from the project's own one-line description. It is not independent research and contains no verified traction, revenue or customer data.